Machine learning method for small-angle X-ray scattering measurement data analysis model
The method addresses the complexity of SAXS model selection by generating a training dataset that simulates realistic experimental conditions, enhancing the accuracy of SAXS data analysis for nanostructure characterization.
Patent Information
- Application Number
- JP2025515744
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-09-14
- Filing Date
- 2023-09-13
- Publication Date
- 2025-09-04
AI Technical Summary
The challenge in small-angle X-ray scattering (SAXS) analysis lies in the complexity of selecting an appropriate model for nanostructure analysis due to the wide range of scattering models and poorly characterized data, leading to inaccuracies in determining sample parameters.
A computer-implemented method for generating a training dataset using machine learning, which involves calculating and smearing SAXS intensity profiles based on experimental settings, adding noise, and normalizing data to improve model fitting accuracy.
Enhances the accuracy of SAXS data analysis by simulating realistic experimental conditions, allowing for precise determination of structural parameters through improved model fitting and prediction.
Smart Images

Figure 2025529473000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to the analysis of material nanostructures using small-angle X-ray scattering (SAXS), and more particularly to the training and inference of SAXS measurement data analysis models based on machine learning. [Background technology]
[0002] Small-angle X-ray scattering (SAXS) is a powerful analytical technique for studying the structure of bulk disordered samples, such as nanoparticles, colloids, nanocomposites, polymers in solution, and biomaterials.
[0003] SAXS measures the intensity of X-rays scattered by a sample as a function of the scattering angle. Figure 1 illustrates how SAXS works. In a typical experiment, a highly collimated monochromatic X-ray beam 1 is transmitted through a sample 2 approximately 1 mm thick. The scattered X-rays are collected on a two-dimensional area detector 3, which produces an azimuthal intensity distribution at scattering angles offset from the directly transmitted beam. The scattering angle (conventionally defined as 2θ) defines a "probe length" expressed as D = 2π / q (q = 4π sin(θ) / λ, where λ is the wavelength of the X-rays (typically 0.154 nm for a copper X-ray source). Measurements are performed over very small angles, typically in the range of 0.1° to 10°, and sample parameters can be determined from analysis of the scattered intensity profile of the sample as a function of angle.
[0004] Materials scientists have addressed SAXS data by fitting experimental scattering curves with theoretical scattering models for nanostructures in a medium (solvent liquid for dispersions, air for powders). The difficulty with SAXS analysis lies in the fact that a wide range of scattering models have been developed to describe various shape factors (related to particle shape) and combinations of shape factors, which may be combined with structure factors to describe sample interactions. Furthermore, different models (and / or related structural parameters) may be equally well-fit to experimental data. The large number of models and the fact that SAXS data are poorly characterized complicates the selection of an appropriate model for the user.
[0005] Against this background, machine learning (ML) methods have recently been used to assist users in selecting the best model for analyzing SAXS data. For example, a trained ML classification algorithm can automatically provide the fit probability of each model from a large set of theoretical scattering models to experimental SAXS measurement data. Summary of the Invention
[0006] An object of the present invention is to provide a SAXS analysis-assisted machine learning method with improved accuracy. To this end, the present invention provides a computer-implemented method for generating a training dataset and performing machine learning of a first small-angle X-ray scattering (SAXS) measurement data analysis model using the training dataset, wherein the generation of the training dataset includes: - calculating one-dimensional SAXS intensity profiles corresponding respectively to the variation parameters of at least one structural model of the sample in the medium; - estimating a first experimental one-dimensional SAXS intensity profile from each of the calculated one-dimensional SAXS intensity profiles based on a first set of SAXS measurement experimental settings including a first X-ray beam intensity profile at a two-dimensional detector.
[0007] More particularly, the estimation involves smearing each of the calculated one-dimensional SAXS intensity profiles by convolution with a point spread function associated with the first X-ray beam intensity profile.
[0008] A preferred, non-limiting embodiment of the method is as follows.
[0009] The first set of SAXS measurement experimental settings further includes a signal-to-noise ratio setting including a size of the two-dimensional detector and a beam position on the two-dimensional detector, and the estimation further includes adding noise to each of the calculated SAXS intensity profiles based on the signal-to-noise ratio setting.
[0010] The method further includes determining an experimental one-dimensional SAXS intensity profile of the blank solution based on the first set of SAXS measurement experimental settings and the theoretical one-dimensional SAXS intensity profile of the blank solution, and subtracting the experimental one-dimensional SAXS intensity profile of the blank solution from each of the experimental one-dimensional SAXS intensity profiles.
[0011] Calculation of a one-dimensional SAXS intensity profile corresponding to variation parameters of at least one structural model of the sample in the medium is performed for a plurality of structural models, and the first SAXS measurement data analysis model is a structural model classifier trained to predict the fit between the experimental SAXS measurement data and each structural model of the plurality of structural models.
[0012] The calculation of a one-dimensional SAXS intensity profile corresponding to a variation parameter of at least one structural model of the sample in the medium is performed for a single structural model, and the first SAXS measurement data analysis model is a regression model trained to predict the value of at least one of the variation parameters based on experimental SAXS measurement data.
[0013] In the method, generating the training data set further comprises normalizing, converting to a logarithmic scale and standardizing the determined experimental one-dimensional SAXS intensity profiles.
[0014] The first set of SAXS measurement experimental settings further includes the pixel size of the detector and the distance from the sample to the detector.
[0015] The first set of SAXS measurement experimental settings further includes a measurement time.
[0016] The method further includes determining a second experimental one-dimensional SAXS intensity profile for each of the calculated one-dimensional SAXS intensity profiles based on a second set of SAXS measurement experimental settings including a second X-ray beam intensity profile at the two-dimensional detector, wherein the first experimental one-dimensional SAXS intensity profile is used to perform machine learning of the first SAXS measurement data analysis model, and the second experimental one-dimensional SAXS intensity profile is used to perform machine learning of a second SAXS measurement data analysis model associated with the second set of SAXS measurement experimental settings.
[0017] The present invention also provides acquiring SAXS data corresponding to experimental measurements on a SAXS instrument; and processing the acquired SAXS data using the first SAXS measurement data analysis model trained according to the present invention.
[0018] The present invention also provides acquiring SAXS data corresponding to experimental measurements on a SAXS instrument; processing the acquired SAXS data using a structural model classifier trained according to the present invention; selecting the structural model that best fits the experimental SAXS measurement data; and processing the acquired SAXS data using the regression model trained according to the present invention, with the selected structural model as a single structural model, to determine the value of at least one of the variation parameters.
[0019] The method further includes determining parameters of the sample model by fitting the acquired SAXS data to the sample model, the determination being initiated based on the determined values and the selected structural model.
[0020] The present invention also provides acquiring SAXS data corresponding to experimental measurements on a SAXS instrument having predetermined experimental settings; selecting a first or second SAXS measurement data analysis model trained according to the present invention based on a predetermined experimental setting; and processing the acquired SAXS data using a selected one of the first or second SAXS measurement data analyses.
[0021] The present invention also provides acquiring SAXS data corresponding to experimental measurements with a SAXS instrument having an experimental setup; According to the present invention, performing machine learning of a first SAXS measurement data analysis model using the above experimental setup; and processing the acquired SAXS data using the learned first SAXS measurement data analysis model.
[0022] Other aspects, objects, advantages and features of the present invention will become more apparent from a reading of the following detailed description of preferred embodiments, given by way of non-limiting example and made with reference to the accompanying drawings, in which: [Brief explanation of the drawings]
[0023] [Figure 1] FIG. 1 is a diagram illustrating the principle of SAXS measurement. [Figure 2] FIG. 1 illustrates an example of an embodiment of a method according to the present invention for generating a training data set and using the training data set to perform machine learning of a SAXS measurement data analysis model. [Figure 3] FIG. 1 illustrates an example of the use of a SAXS measurement data analysis model trained according to the present invention. [Figure 4] FIG. 1 shows an example of live training of a SAXS measurement data analysis model according to the present invention. [Figure 5] 1 shows a theoretical SAXS intensity profile and the corresponding experimental SAXS intensity profile estimated according to the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0024] The present invention relates to a computer-implemented method for generating a training data set and using the training data set to perform machine learning of a SAXS measurement data analysis model associated with one or more sets of SAXS measurement experimental setups.
[0025] In one embodiment, the SAXS measurement data analysis model is a classification model trained to predict the fit between the experimental SAXS measurement data and each of a plurality of structural models.
[0026] In another embodiment, the SAXS measurement data analysis model is a regression model trained to predict the value of at least one parameter of the structural model based on experimental SAXS measurement data.
[0027] A structural model is an analytical shape factor that describes the scattering sample (scatterers) in a medium or buffer (usually a solvent in the case of dispersions, or air in the case of powders) with variable parameters. The variable parameters include variable geometric parameters that describe the size of the structures (and possibly the associated polydispersity). The variable parameters may also include electronic parameters. As known from the prior art, a structural model may be the product of a shape factor and a structure factor that describes the interactions between the scatterers.
[0028] As an example, consider six structural models associated with the following shape factors: sphere, cylinder, ellipsoid, core-shell sphere, core-shell cylinder, and core-shell ellipsoid. Geometric parameters include radius, length, shell thickness depending on the object, and polydispersity. Electronic parameters include the scattering length density of the sample in solution (obtained as the difference between the scattering length densities of the scatterer and the buffer), or, in the case of a core-shell model, the scattering length density of the core and the scattering length densities of the shell. The scattering length density defines the scattering power of the sample and can be estimated based on the composition and mass density of the sample and buffer.
[0029] 2, the method according to the present invention comprises a step E1 of calculating one-dimensional SAXS intensity profiles corresponding to the variation parameters of at least one structural model of the sample in the medium, in order to generate a training data set. When training a classification model, this step is carried out for multiple structural models. When training a regression model, this step is usually carried out for one structural model. The calculated one-dimensional SAXS intensity profiles can be stored in a database 4.
[0030] Each calculated one-dimensional SAXS intensity profile is a theoretical one-dimensional scattering curve associated with a shape factor and associated geometric and electronic parameters. Thus, calculated one-dimensional SAXS intensity profiles corresponding to the same structural model's variation parameters form a set of theoretical one-dimensional scattering curves associated with the same shape factor but with different values of the associated geometric and electronic parameters. When calculating such theoretical one-dimensional scattering curves associated with the same shape factor, random sampling may be applied to at least one of the associated geometric and electronic parameters.
[0031] In an exemplary embodiment, 10,000 curves I(q) are calculated for each shape factor, resulting in a total of 60,000 simulated curves. The curves are simulated over a q range defined by a lower limit qmin, an upper limit qmax, and a number of q points. As an example, for a core-shell cylinder, for each simulated curve, the radius value is randomly varied between 50 and 1000 Å, the length value is randomly varied between 1 and 30 times the radius, the shell thickness is randomly varied between 50 and 1000 Å, and the size polydispersity (expressed as a percentage of the full width at half maximum of a Gaussian function centered on the mean value) is randomly varied between 0 and 30%. The electronic parameters are calculated using a scattering length density of 9.10 for the solution. -6 Å -2 , the scattering length density of the core is 121.10 -6 Å -2 , the shell scattering length density is 18.10 -6 Å -2 It may be a fixed parameter such as:
[0032] 2, the method according to the present invention further includes step E2 of estimating an experimental one-dimensional SAXS intensity profile from each calculated one-dimensional SAXS intensity profile based on a set of SAXS measurement experimental settings. This step E2 estimates a one-dimensional experimental curve that would be obtained using a SAXS experimental setup with the set of SAXS measurement experimental settings. In other words, the effect of these settings on the SAXS measurement is simulated by converting the one-dimensional theoretical curve to a one-dimensional experimental curve. As described below, this conversion can model the effect of certain sample characteristics, instrument resolution parameters, and instrument signal-to-noise ratio parameters.
[0033] According to the present invention, each set of SAXS measurement experimental settings includes an instrument resolution setting that includes an X-ray beam intensity profile at a two-dimensional detector, and estimating an experimental SAXS intensity profile from each of the calculated SAXS intensity profiles includes smearing each of the calculated SAXS intensity profiles by convolution with a point spread function associated with the X-ray beam intensity profile. In this regard, Figure 5 shows a calculated SAXS intensity profile It and a corresponding experimental intensity profile Ie that has been smeared according to the present invention.
[0034] In one embodiment, the beam size and profile are given based on calculations that take into account the direct beam divergence and the size of the collimating slit, while in another embodiment, the actual direct beam size profile is used.
[0035] The point spread function can be defined by a smearing parameter that describes the beam size profile at the detector. This parameter can be the FWHM (full width at half maximum) of the intensity beam profile at the detector, estimated from the measurement configuration (e.g., from the optical design of the X-ray scattering beamline), or it can be obtained by experimental characterization of the beam profile at the detector and Gaussian fitting to determine this smearing parameter (e.g., FWHM). The experimental SAXS intensity profile is obtained by applying a convolution kernel of finite size (FWHM) to the theoretical intensity profile obtained in step E1.
[0036] In a preferred embodiment of the present invention, the complete one-dimensional intensity beam profile at the detector (e.g., Q of 9 or more orders of magnitude) is used to convolve the theoretical intensity profile. In practice, beams are not perfectly Gaussian, so approximating them with a Gaussian function can degrade the accuracy of the model. The direct beam profile can be determined experimentally and can therefore be used directly to smear the beam. This has been shown to improve accuracy with real data.
[0037] The resolution settings of the device may further include the pixel size of the detector and / or the distance from the sample to the detector.
[0038] The set of SAXS measurement experimental settings may further include sample property settings that characterize the scattering power of the sample, e.g., selected from the chemical composition of the sample or its scattering length density, the chemical composition of the buffer solution or its scattering length density, the thickness of the sample, and the volume fraction of the sample relative to the solution.
[0039] The set of SAXS measurement experimental settings may further include measurement condition settings such as measurement time.
[0040] The set of SAXS measurement experimental settings may further include a signal-to-noise ratio setting selected from, for example, beam intensity, detector area, and beam position on the detector. In this case, estimating the experimental curve from the theoretical curve may further include adding noise to each calculated SAXS intensity profile based on the signal-to-noise ratio setting. For example, Poisson statistics may be added to consider realistic data in terms of signal-to-noise ratio.
[0041] Unlike other methods used to train AI (artificial intelligence) models applied to SAXS, the present invention simulates scattering data using realistic statistics for a given model (form factor) and a given set of experimental settings, taking into account, among other things, the effect of the beam profile at the detector on data smearing. Applicants have discovered that this has limited impact in synchrotron applications, but is particularly relevant for experimental SAXS applications.
[0042] Additionally, the impact of measurement and detection geometry on the signal-to-noise ratio can be considered. Some models may apply simple Poisson statistics based on the scattering power of the sample (an inherent property resulting from the electron contrast of the material relative to the buffer or, in the case of powders, air), beam intensity, and measurement time. Here, the measurement and detection scheme may be considered in addition to the beam intensity, measurement time, and sample scattering properties. For example, the detection surface of the detector influences the number of pixels used to collect the signal and azimuthally average the 2D data pattern into a 1D scattering curve. This therefore affects the signal-to-noise ratio of the 1D scattering curve, especially near the end of the 1D scattering curve (Qmax).
[0043] In a preferred embodiment, step E2 further comprises simulating the scattering profile of a buffer (e.g., a solution in which the object to be characterized, e.g., nanoparticles, are dispersed). Such estimation is performed taking into account the experimental configuration (i.e., a set of experimental settings), e.g., in terms of both signal-to-noise ratio and resolution.
[0044] In this embodiment, the method includes estimating an experimental SAXS intensity profile of a blank solution (i.e., a buffer containing no sample to be characterized) based on a first set of SAXS measurement experimental settings and a theoretical SAXS intensity profile of the blank solution, and subtracting the experimental SAXS intensity profile of the blank solution from each of the estimated experimental curves (intensity profiles of the sample dispersed in the solution).
[0045] In a possible embodiment, the input is a shape factor model representing the sample (i.e., ideal theoretical intensity) and the buffer, respectively, to simulate the effects of the experimental setup. The output is a predicted expected measured intensity and signal-to-noise ratio at each q value available in the experiment. The number of photons scattered by the sample and collected by each pixel of the detector is estimated for both (a) the sample in the buffer and (b) the buffer only. Pixel-independent noise based on the signal-to-noise ratio setup may be added to each pixel in scenarios (a) and (b), for example, based on a Poisson distribution. The two-dimensional data matrix (pixel values at the detector) is reduced to a one-dimensional vector by averaging equivalent orientations and fully accounting for the effects of geometric distortion. The result may be normalized with respect to the number of photons scattered in (a) and (b), respectively. The sample-only signal can then be obtained by subtracting the previously obtained one-dimensional vector.
[0046] The estimated experimental SAXS intensity profiles form a training data set, which can be divided into a learning data set 41 (e.g., including 80% of the experimental profiles) and a test data set 42 (e.g., including 20% of the experimental profiles). Furthermore, the method includes a step E3 of performing machine learning of a SAXS measurement data analysis model DAM using the training data set. This step E3 is divided into a first step E31 of obtaining a model using an AI algorithm and the learning data set 41, and a second step E32 of evaluating the obtained model using the test data set 42 to test its prediction accuracy.
[0047] In step E31, the experimental profile may be pre-processed before being passed to the AI algorithm, which may include normalizing the experimental profile, converting it to a logarithmic scale, and standardizing it.
[0048] Normalization involves normalizing to the intensity I at the first q points, or normalizing to an invariant (dividing the curve by the area, making q a power of 2).
[0049] Conversion to a logarithmic scale is 10 -12 or 10 -6 ~10 -15 This involves applying a threshold to negative points that is set to a random value between .
[0050] Normalization can be performed using a standard scaler that subtracts the mean value of the experimental profile at a given q value from the experimental profile and divides the result by the standard deviation of the experimental profile at the q value.
[0051] If the model being trained is a regression model for predicting particle size (e.g., radius if the shape factor is spherical) and polydispersity, the training features are all I(q) values for each q-point associated with each experimental profile, and the targets are two scalars: size and polydispersity. Several AI regression algorithms are considered for predicting size and polydispersity. The best performance on the test dataset was achieved using classical machine learning random forests and one-dimensional convolutional neural networks. After training, the model takes only the I(q) curve as input, and after optional preprocessing, outputs two predicted values: size and polydispersity. In a regression model trained based on the random forest algorithm, the model's prediction can be the average of the predictions made by the N estimators (trees) that make up the forest. The model can also provide a standard deviation parameter associated with the predictions of the N estimators.
[0052] Many different types of laboratory SAXS instruments are commercially available. Some have somewhat fixed experimental settings, while others have a much wider range of settings. In both cases, at least some of the instrument and signal-to-noise configurations are based on user-defined settings, such as beam size and corresponding intensities. This requires creating different models for a given classifier or different models for a given regression algorithm (e.g., different analysis modules corresponding to the set of experimental parameters selected for a given algorithm trained for sphere sizing or ellipsoid sizing). For example, three models corresponding to three different beam size settings and associated intensities are trained. A typical example of such settings and application values is 0.0018 Å when applying the model to samples with scatterer sizes (or diameters) between approximately 1 nm and 200 nm. -1 , 0.0024Å -1 , and 0.0028 Å -1 and a point spread function of 1×10 6 ph / s, 1×10 7 ph / s, and 6 × 10 7 Three different models can be constructed based on the beam size settings and associated intensities, each corresponding to a beam intensity in ph / s.
[0053] Therefore, the present invention provides determining a first experimental SAXS intensity profile for each of the calculated SAXS intensity profiles based on a first set of SAXS measurement experimental settings including a first X-ray beam intensity profile at the two-dimensional detector; and determining a second experimental SAXS intensity profile for each of the calculated SAXS intensity profiles based on a second set of SAXS measurement experimental settings including a second X-ray beam intensity profile at the two-dimensional detector. the first experimental SAXS intensity profile is used to perform machine learning of a first SAXS measurement data analysis model associated with a first set of SAXS measurement experimental setup settings; The second experimental SAXS intensity profile is used to perform machine learning of a second SAXS measurement data analysis model associated with the second SAXS measurement experimental setup.
[0054] Of course, two or more SAXS measurement data analysis models may be trained using the above procedure.
[0055] The first and / or second SAXS measurement data analysis model can be labeled with a set of associated first or second SAXS measurement experimental settings, which can include associating a minimum scattering wave vector from the experimental SAXS intensity profile.
[0056] The present invention also relates to the use of a trained SAXS measurement data analysis model for performing SAXS measurement data analysis. Referring to Figure 2, such use typically involves, in step E4, acquiring SAXS data corresponding to experimental measurements by a SAXS apparatus, and processing the acquired SAXS data using the trained SAXS measurement data analysis model DAM to obtain predicted values PRED.
[0057] In a possible embodiment, a plurality of SAXS measurement data analysis models are trained, each associated with a predetermined set of experimental settings. These SAXS measurement data analysis models may consist of or include the first and second SAXS measurement data analysis models. In such a case, the SAXS measurement data analysis method may include the steps of acquiring SAXS data corresponding to experimental measurements using a SAXS apparatus having predetermined experimental settings, selecting one of the trained SAXS measurement data analysis models based on the predetermined experimental settings, and processing the acquired SAXS data using the selected one of the trained SAXS measurement data analysis models.
[0058] In a possible embodiment, the method further includes determining a usage boundary of the first SAXS measurement data analysis model after the first SAXS measurement data analysis model has been trained using the first set of SAXS measurement experimental settings. As exemplified below, determining the usage boundary may include evaluating a prediction error of the first SAXS measurement data analysis model when processing SAXS data corresponding to experimental measurements with a SAXS apparatus having a test experimental setting different from the first set of SAXS measurement experimental settings. If the prediction error is less than a threshold, the test experimental setting may be included within the usage boundary of the first SAXS measurement data analysis model.
[0059] Such bounds can be determined by, for example, assessing the robustness of the model to sample and experimental setup variability as defined by prediction error, and identifying settings that result in outlier errors greater than a threshold.
[0060] For example, sample-related and experimental settings include chemical composition, volume fraction, acquisition time, beam smearing, and intensity. Chemical compositions related to scattering power of measured samples vary from Au in water and SiO2 in water to polystyrene in toluene. Experimental measurements of dilute solutions of SiO2 nanoparticles in water under known conditions yield reference training values: a volume fraction of 0.005 vol / vol and an exposure time of 30 minutes. Test datasets obtained with variable parameters (specifically, acquisition times ranging from a few seconds to 1-2 hours, and volume fractions 3-4 orders of magnitude smaller than the reference values) are used to evaluate the robustness of the model to sample variability and experimental conditions. Outliers can be defined, for example, by analyzing the evolution of prediction errors. A two-fold increase in the root-mean-square error may be considered a criterion for outlier exclusion. In this way, a model created under the above test conditions (0.005 vol / vol, 30 min exposure time) can be used for X-ray scattering data of the same sample with a measurement time longer than 10 min and a volume fraction higher than 0.0005 vol / vol.
[0061] Furthermore, after the step of acquiring SAXS data corresponding to experimental measurement by the SAXS apparatus, the minimum wave vector Qmin of the acquired SAXS data exp a control step of checking the compatibility of the experimental settings with the usage boundaries of the first or second SAXS measurement data analysis model; and exp is the simulated Qmin sim (i.e., the minimum scattering wave vector determined from the experimental SAXS intensity profile). Indeed, for the ML algorithm, Qmin exp was used to create the analysis model for the first or second SAXS measurement data. sim It is important to be as close as possible to Qmin exp was used to create the analysis model for the first or second SAXS measurement data. sim If Qmin is greater than , the corresponding model will not be used to infer the input SAXS data. exp Qmin is usually determined by subtracting the background scattering signal of the solution sample and the buffer to extract the signal from the scatterers. The criterion for determining a good Qmin with little error is to consider the absolute scattering intensity ("absolute intensity") and consider the timing at which the scattering difference crosses the background scattering signal of the buffer. This allows for a consistent Qmin between experiments using previous measurements of the buffer or between experiments using a data standard reference value of the same buffer in the same experimental setting. exp The decision can be made.
[0062] Importantly, the control step involves checking whether the beam resolution parameters are compatible (i.e., within the usage boundaries) with the values used to create the data analysis model for the first or second SAXS measurement. Obtaining the beam resolution parameters during the experiment can be performed by extracting this value from a measurement without a sample or with a buffer sample immediately before the experiment. The beam resolution parameters (e.g., FWHM or full beam profile calculated from Gaussian fitting) are then compared with the beam resolution parameters used in step E2.
[0063] In a possible embodiment, the selection of an appropriate SAXS measurement data analysis model (i.e., one trained with appropriate experimental settings) can be performed automatically using a pre-selection tool configured to extract the experimental settings from a configuration file that is updated during the measurement process or from a data header that is filled in at the end of the measurement process. The pre-selection tool may also be configured to collect the experimental settings from user-entered data (e.g., chemical formulas, concentrations, etc.).
[0064] As mentioned above, the ability to user-define multiple experimental settings over multiple value ranges suggests the creation of multiple analytical models. The applicant has discovered that some instrument settings are more sensitive to deviations in the exact measurement location than others. In possible embodiments of the present invention, tolerances can be applied to certain resolution, signal-to-noise ratio, or sample characteristic parameters so that the experimental settings used to infer measurements deviate from (i.e., fall within) the boundaries of use from) the settings used to train the model. For example, tolerances of a few percent to up to 10 percent can be applied to the sample-to-detector distance or pixel size between the experimental settings used for measurements and those used to train the model. In this case, preliminary data processing can be performed (well-known binning or segmentation processes can be applied) to resample the experimental data on the same Q-grid.
[0065] In another possible embodiment, the SAXS measurement data analysis model is trained using predetermined sample properties (e.g., scattering power, concentration) and measurement settings (e.g., measurement time). The trained model can then be applied to experimental data acquired on samples with different properties using different measurement settings. Such experimental data may fall within the usage boundaries.
[0066] In a possible embodiment shown in Figure 3, the present invention relates to a method for analyzing SAXS measurement data, comprising step F1 of acquiring SAXS data corresponding to experimental measurements using a SAXS apparatus and step F2 of processing the acquired SAXS data using a classification model trained according to the present invention. Step F2 predicts the fit between the acquired SAXS data and each of a plurality of structural models (denoted as "Form Factor and Prediction" in Figure 3). The method further comprises step F3 of processing the acquired SAXS data using a regression model trained according to the present invention and associated with the structural model that best fits the acquired SAXS data, as predicted by the classification model in step F2. Step F3 makes it possible to determine the value of at least one variation parameter (in the case of a random forest regression model, unknown structural parameters such as size, polydispersity, and their standard deviations) from the acquired SAXS data. The method further comprises fitting the acquired SAXS data to the model to determine the parameters of the sample model, the determination being initiated based on the determined values and the selected structural model.
[0067] In another embodiment, ML training can be performed in real time, for example, when the number of user-defined parameter values is too large and the performance of the ML algorithm is too sensitive to such values. For this purpose, a pre-selection tool may be used to extract actual measurement conditions relevant to the experimental setup. The ML algorithm is then trained by creating training and test data sets relevant to the experimental setup. According to this embodiment, SAXS measurement data analysis is performed by: acquiring SAXS data corresponding to experimental measurements with a SAXS instrument having an experimental setup; performing machine learning of a first SAXS measurement data analysis model using the experimental settings as a first set of SAXS measurement experimental settings; and processing the acquired SAXS data using the learned first SAXS measurement data analysis model.
[0068] In a further embodiment shown in Figure 4, the present invention relates to a method for analyzing SAXS measurement data, comprising a step G1 of acquiring SAXS data corresponding to an experimental measurement with a SAXS apparatus having a predetermined set of experimental settings. In step G2, a compatibility check is performed between this experimental measurement (with respect to the sample and experimental settings, i.e., the scattering power of the sample, the resolution, the sample-to-detector distance, and the acquisition time) and a pre-trained model. The compatibility check involves verifying whether the experimental measurement is within the usage boundaries of the pre-trained model. If such a compatibility check is performed, step G3 is performed, in which one of the pre-trained models is selected based on the predetermined set of experimental settings, followed by step G4, in which the selected model is used to analyze the measured SAXS data. If such a compatibility check is not performed, step G2 is followed by steps G4, G5, G6, in which a model is trained based on the predetermined set of experimental settings (real-time training, as described above), and G6, in which the trained model is used to analyze the measured SAXS data.
[0069] The present invention is not limited to the above-described methods, but further relates to a computer program product including instructions that, when executed by a computer, cause the computer to perform any one of the above-described methods, a data processing unit including a processor configured to perform any one of the above-described methods, and a small-angle X-ray scattering apparatus including such a data processing unit. While the present invention has been described above with reference to preferred embodiments thereof, it should be understood by those skilled in the art that various changes in form and detail can be made without departing from the spirit and scope of the present disclosure as defined in the appended claims. In particular, those skilled in the art will understand that the training and learning datasets of the data analysis model (DAM) can be associated with a set of multiple SAXS measurement experimental configurations, each associated with a different given X-ray beam intensity profile at the two-dimensional detector, which are used to estimate the experimental SAXS intensity profile in step E2.
Claims
1. 1. A computer-implemented method for generating (E1, E2) training datasets (41, 42) and performing machine learning of a first small-angle X-ray scattering (SAXS) measurement data analysis model (DAM) using (E3, E31, E32) the training datasets, comprising: The generation of the training data set comprises: Calculating (E1) one-dimensional SAXS intensity profiles corresponding respectively to the variation parameters of at least one structural model of the sample in the medium; and (E2) estimating a first experimental one-dimensional SAXS intensity profile from each of the calculated one-dimensional SAXS intensity profiles based on a first set of SAXS measurement experimental settings including a first X-ray beam intensity profile at a two-dimensional detector; 2. The computer-implemented method of claim 1, wherein the estimating (E2) comprises smearing each of the calculated one-dimensional SAXS intensity profiles by convolution with a point spread function associated with the first X-ray beam intensity profile.
2. The computer-implemented method of claim 1 , further comprising determining a usage boundary for the first SAXS measurement data analysis model.
3. 3. The computer-implemented method of claim 2, wherein determining the usage boundaries includes evaluating a prediction error of the first SAXS measurement data analysis model when processing SAXS data corresponding to experimental measurements with a SAXS instrument having a test experimental setting different from the first set of SAXS measurement experimental settings.
4. 4. The computer-implemented method of claim 1, wherein the first set of SAXS measurement experimental settings further includes a signal-to-noise ratio setting including a size of the two-dimensional detector and a beam position on the two-dimensional detector, and the estimation further includes adding noise to each of the calculated SAXS intensity profiles based on the signal-to-noise ratio setting.
5. 5. The computer-implemented method of claim 1, further comprising: determining an experimental one-dimensional SAXS intensity profile of the blank solution based on the first set of SAXS measurement experimental settings and a theoretical one-dimensional SAXS intensity profile of a blank solution; and subtracting the experimental one-dimensional SAXS intensity profile of the blank solution from each of the experimental one-dimensional SAXS intensity profiles.
6. 6. The computer-implemented method of claim 1, wherein the calculation of the one-dimensional SAXS intensity profile corresponding to the variation parameters of at least one structural model of the sample in the medium is performed for a plurality of structural models, and the first SAXS measurement data analysis model is a structural model classifier trained to predict the fit between the experimental SAXS measurement data and each of the plurality of structural models.
7. 6. The computer-implemented method of claim 1, wherein the calculation of the one-dimensional SAXS intensity profile corresponding to the variation parameters of at least one structural model of the sample in the medium is performed for a single structural model, and the first SAXS measurement data analysis model is a regression model trained to predict the value of at least one of the variation parameters based on experimental SAXS measurement data.
8. 8. The computer-implemented method of claim 1, wherein generating the training data set further comprises normalizing, converting to a logarithmic scale, and standardizing the determined experimental one-dimensional SAXS intensity profiles.
9. The computer-implemented method of any one of claims 1 to 8, wherein the first set of SAXS measurement experiment settings further includes a detector pixel size and a sample-to-detector distance.
10. The computer-implemented method of any one of claims 1 to 9, wherein the first set of SAXS measurement experiment settings further includes a measurement time.
11. determining a second experimental one-dimensional SAXS intensity profile for each of the calculated one-dimensional SAXS intensity profiles based on a second set of SAXS measurement experimental settings including a second X-ray beam intensity profile at the two-dimensional detector; 11. The computer-implemented method of claim 1, wherein the first experimental one-dimensional SAXS intensity profile is used to perform machine learning of the first SAXS measurement data analysis model, and the second experimental one-dimensional SAXS intensity profile is used to perform machine learning of a second SAXS measurement data analysis model associated with a second set of SAXS measurement experimental configurations.
12. acquiring SAXS data corresponding to experimental measurements with a SAXS instrument; and processing the acquired SAXS data using a first SAXS measurement data analysis model trained by the method of claim 1.
13. Acquiring SAXS data (F1) corresponding to experimental measurements by a SAXS instrument; (F2) processing the acquired SAXS data with a structural model classifier trained according to the method of claim 6; selecting the structural model that best fits the experimental SAXS measurement data; and (F3) processing the acquired SAXS data using the regression model trained by the method of claim 7, with the selected structural model as a single structural model, to determine the value of at least one of the variation parameters.
14. 14. The computer-implemented method of claim 13, further comprising: determining (F4) parameters of the sample model by fitting the acquired SAXS data to the sample model, the determining being initiated based on the determined values and the selected structural model.
15. acquiring SAXS data corresponding to experimental measurements with a SAXS instrument having a predetermined experimental setup; selecting a first or second SAXS measurement data analysis model trained by the method of claim 12 based on the predetermined experimental setting; and processing the acquired SAXS data using a selected one of the first or second SAXS measurement data analyses.
16. - acquiring (G1) SAXS data corresponding to experimental measurements with a SAXS instrument having an experimental setup; Implementing the method of claim 1 (G5), the experimental settings are set as a first set of SAXS measurement experimental settings, and performing machine learning of a first SAXS measurement data analysis model; and processing (G6) the acquired SAXS data using the learned first SAXS measurement data analysis model.
17. A data processing unit comprising a processor configured to carry out the method according to any one of claims 1 to 16.
18. A small-angle X-ray scattering device comprising the data processing unit according to claim 17.
19. A computer program product comprising instructions that cause a computer to carry out the method according to any one of claims 1 to 14 when the program is executed by a computer.