Computer-implemented methods of analyzing experimental data and computer-implemented methods of generating training data to train machine learning models for analyzing experimental data

CN122743552APending Publication Date: 2026-09-11MAX PLANCK GESELLSCHAFT ZUR FOERDERUNG DER WISSENSCHAFTEN EV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480081661.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-20
Filing Date
2024-06-05
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

然而,在OOD泛化技术的发展中存在严重不足,特别是在涉及振动光谱、NMR波谱和质谱的分子分析中,以及在临床化学分析中

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122743552A_ABST
    Figure CN122743552A_ABST
Patent Text Reader

Abstract

A computer-implemented method for analyzing experimental data is provided. The method includes: receiving a first dataset comprising experimental data to be analyzed. The method further includes: receiving one or more variation datasets, wherein each variation dataset includes information about variation in experimental observations. Furthermore, the method includes: extrapolating the experimental data by applying variation to the experimental data, wherein the applied variation corresponds to one or more variations in the experimental observations specified by information contained in the one or more variation datasets. The method further includes: generating a synthetic dataset based on the extrapolated experimental data, analyzing the synthetic dataset, and providing output based on the results of the analysis of the synthetic dataset. The received one or more variation datasets are selected according to one or more types of noise associated with the experimental data to be analyzed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Computer-implemented methods for analyzing experimental data are provided, including methods for generating training data to train a machine learning model for analyzing experimental data, methods for training a machine learning model for analyzing experimental data, data processing apparatus, computer-implemented machine learning models for analyzing experimental data, uses of the computer-implemented machine learning models according to this disclosure for analyzing experimental data, computer-implemented methods for classifying experimental data subjected to variation using machine learning models, and methods for identifying individuals from samples of biological fluids based on individuals. Therefore, this disclosure relates to experimental data analysis using machine learning models and / or with the assistance of machine learning models. Background Technology

[0002] Advances in molecular analysis have increasingly enabled the probing of biological systems. Distinguishing physiologically relevant states from quantitative molecular fingerprints has opened new opportunities for phenotypic analysis. Consequently, significant efforts have been made to develop standardized procedures, including simplified biological sampling, post-collection processing, and quantitative measurements of sensitivity. However, empirical datasets remain susceptible to a wide variety of sources of variability, both analytical and intrinsically biological (see RA Bowen and AT Remaley, “Interferences from blood collection tube components on clinical chemistry assays,” Biochemia Medica, pp. 31–44, 2014). Obtaining datasets that reflect the true distribution of data is often resource-intensive, expensive, and in some cases, virtually impossible. This is especially true in the context of clinical research, which encompasses all pathophysiological levels, studies rare diseases, or conducts longitudinal probing of the same system over time. Consequently, exploratory studies are often limited in sample size, making it difficult for a given “training” set to represent the true, unseen “test” domain. Therefore, when a developed machine learning (ML) model is applied to samples that are independently collected and experimentally measured, the model may not achieve the expected power.

[0003] While traditional methods typically rely on standardized experimental workflows and the creation of computer-aided processing techniques to reduce unwanted empirical noise, complete noise removal appears to be unattainable (see CLM Morais, KMG Lima, M. Singh, and FL Martin, “Tutorial: multivariate classification for vibrational spectroscopy in biological samples,” Nature Protocols, vol.15, pp. 2143–2162, June 2020). Failure to account for noise and distribution shifts that may obscure the biological patterns of genuine interest can mislead ML algorithms into exploiting confounding information unlikely to be reproduced. This failure stems from a violation of the assumption upon which (supervised) ML algorithms are based that training and test data are independently and identically distributed (i.i.d.) (see JGMoreno-Torres, T. Raeder, R. Alaiz-Rodriguez, NV Chawla, and F. Herrera, “A unifying view on dataset shift in classification,” Pattern Recognition, vol. 45, pp. 521–530, Jan. 2012). To interpret the information that collected datasets may contain, understanding and interpreting analytical and biological variability can be crucial for ensuring successful data analysis. The concept of out-of-distribution (OOD) generalization has recently gained attention in ML research to address the shortcomings of the iid hypothesis (see J. Liu, Z. Shen, Y. He, X. Zhang, R. Xu, H. Yu, and P. Cui, “Towards out-of-distribution generalization: A survey,” arXiv, 2023). This paradigm shift acknowledges the unpredictability of unseen data, prompting the exploration of methods that better adapt to distribution offsets to generalize beyond the training set. OOD generalization has been extensively explored in computer vision and natural language processing tasks (see X. Zhang, L. Zhou, R. Xu, P. Cui, Z. Shen, and H. Liu, “Towards unsupervised domain generalization,” arXiv, 2022).However, there are serious shortcomings in the development of OOD generalization techniques, especially in molecular analysis involving vibrational spectroscopy, NMR spectroscopy and mass spectrometry, as well as in clinical chemistry analysis.

[0004] The following publications describe data augmentation methods for spectral data in deep chemometrics based on convolutional neural networks, according to which some datasets are augmented by adding random mutations of offsets, multipliers, and slopes:

[0005] EJBjerrum et al., Data Augmentation of Spectral Data for Convolutional Neural Network (CNN) Based Deep Chemometrics, arXiv, 2017. Summary of the Invention

[0006] The objective technical problem of this disclosure is to enrich the prior art. Optionally, this objective technical problem may be: to allow experimental data to be analyzed with improved reliability, even when the experimental data includes only a limited number of data points.

[0007] The objective technical problem is solved by the subject matter of the independent claims. Optional features and embodiments are given in the dependent claims and the description.

[0008] A computer-implemented method for analyzing experimental data is provided. The method includes: receiving a first dataset comprising experimental data to be analyzed, and receiving one or more variant datasets, each variant dataset including information about the variants of experimental observations. The method further includes: extrapolating the experimental data by applying variants to the experimental data, wherein the applied variants correspond to one or more variants of the experimental observations specified by information contained in the one or more variant datasets. Furthermore, the method includes: generating a synthetic dataset based on the extrapolated experimental data, analyzing the synthetic dataset, and providing output based on the results of the analysis of the synthetic dataset. The received one or more variant datasets are selected according to one or more types of noise associated with the experimental data to be analyzed.

[0009] In addition, a computer program is provided that includes instructions which, when executed by a computer, cause the computer to perform the method according to this disclosure.

[0010] In addition, a computer-readable medium, optionally a computer-readable storage medium and / or a data carrier signal, is provided, which includes instructions that, when executed by a computer, cause the computer to perform the method according to this disclosure.

[0011] Additionally, a data processing device is provided, which is configured to perform the method according to this disclosure.

[0012] In another aspect, a computer-implemented method is provided for generating training data to train a machine learning model for analyzing experimental data. The method includes: receiving one or more variant datasets, each variant dataset comprising experimental observations subjected to mutations that at least partially affect the experimental data to be analyzed. The method further includes: analyzing the mutations of the experimental observations contained in the one or more variant datasets, and determining variability characteristics based on the analyzed mutations. Furthermore, the method includes: receiving a first dataset comprising at least one reference sample of experimental data and predetermined analysis results for the at least one reference sample; and generating one or more synthetic datasets comprising multiple synthetic data points by applying the determined variability characteristics to the first dataset.

[0013] Furthermore, a computer-implemented method for training a machine learning model for analyzing experimental data is provided. The method includes: receiving a training dataset comprising training data, wherein the training data includes one or more synthetic datasets having the same or similar mutations as one or more variant datasets, the one or more variant datasets being used to generate one or more synthetic datasets based on a first dataset, the first dataset including at least one reference sample of experimental data and predetermined analysis results for the at least one reference sample. The method further includes: training the machine learning model with the received training dataset.

[0014] Furthermore, a computer-implemented method is provided for analyzing experimental data using a machine learning model, wherein the machine learning model is configured to classify the analyzed mutated experimental data, the mutation of which exceeds the mutation of a reference sample of experimental data used as a first dataset to generate training data for training the machine learning model.

[0015] Additionally, a computer-implemented method for analyzing experimental data using a machine learning model is provided, wherein the machine learning model is trained according to the method disclosed herein.

[0016] In addition, a computer program is provided that includes instructions which, when executed by a computer, cause the computer to perform the method according to this disclosure.

[0017] In addition, a computer-readable medium is provided, optionally a computer-readable storage medium and / or a data carrier signal, the computer-readable medium including instructions that, when executed by a computer, cause the computer to perform the method according to the present disclosure.

[0018] In addition, a data processing device is provided, which is configured to perform the method according to this disclosure.

[0019] Additionally, a computer-implemented machine learning model is provided for analyzing experimental data, wherein the machine learning model is configured to analyze experimental data subjected to mutations exceeding the mutations of a reference sample of experimental data used as a first dataset to generate training data for training the machine learning model.

[0020] The disclosure also provides a computer-implemented machine learning model for analyzing experimental data.

[0021] In addition, a computer-implemented method is provided for classifying mutated experimental data using a machine learning model, wherein the mutation of the experimental data to be classified exceeds the mutation of the reference sample of the experimental data used to generate training data to train the machine learning model.

[0022] Additionally, a method for identifying individuals based on samples of biological fluids is provided. The method includes receiving experimental data representing samples of biological fluids characteristic of an individual. The method further includes determining an analysis result of the received experimental data using a machine learning model according to this disclosure, the machine learning model being trained on experimental data of reference samples of biological fluids derived from the individual, wherein the analysis result corresponds to a match or mismatch between the received experimental data representing the sample and the experimental data of the reference sample used to train the machine learning model. Furthermore, the method includes providing information about individual identification based on the match or mismatch between the received experimental data representing the sample and the experimental data of the reference sample used to train the machine learning model.

[0023] Furthermore, a computer-implemented method is provided for generating data for analyzing experimental data. The method includes: receiving one or more variant datasets, each variant dataset including experimental observations subjected to variation that at least partially affects the experimental data to be analyzed. Furthermore, the method includes: analyzing the variation of the experimental observations contained in the one or more variant datasets, and determining variability characteristics based on the analyzed variation. Additionally, the method includes: receiving a first dataset including at least one reference sample of experimental data and predetermined analysis results for the at least one reference sample; and generating one or more synthetic datasets including multiple synthetic data points by applying the determined variability characteristics to the first dataset.

[0024] "Analyzing experimental data" can involve extracting information from data that has been acquired or is being acquired during an experiment. The experimental process can include scientific experiments. Experimental data may be affected by various measurement errors and / or noise and variation from other sources, which may stem from the measurement and / or the nature of the sample generation process and / or the sample itself being examined during the experiment.

[0025] A variation dataset is a dataset that includes information about the variation of experimental observations. The variation of experimental observations should be understood as the noise experienced by the experimental observations. A variation dataset may include multiple data values ​​representing the variation of specific data points of the experimental observations. Variation data can represent a “point cloud,” which should be understood as a “cloud” of data points spread out by noise, related to specific data points representing at least a portion of the experimental observations. A variation dataset may include the point cloud itself. Alternatively or additionally, a variation dataset may include a mathematical description of the variation experienced by the results of the underlying experimental observations. Experimental observations can be or include the results of experimental measurements and / or experimental data from any other source. Experimental observations can be spectroscopic measurement data, which includes spectral information about samples from which spectral information is obtained.

[0026] "Extrapolated experimental data" can refer to adding data points and / or other data that were not originally part of the experimental data. Extrapolated experimental data can include adding one or more point clouds to one or more data points included in the experimental data, wherein the one or more point clouds are defined by information contained in one or more variant datasets.

[0027] "Generating a synthetic dataset" can be understood as creating a dataset based on a first dataset and one or more variant datasets, where extrapolated data is used to create the synthetic dataset. The term "synthetic" means that the synthetic dataset is generated artificially involving extrapolated experimental data and does not only include measured experimental data.

[0028] "Analyzing synthetic datasets" should be understood as extracting at least the information intended to be obtained by analyzing experimental data from synthetic datasets, just as one would analyze experimental data. Analyzing synthetic data may include the same steps performed routinely based on (raw) experimental data, except that a synthetic dataset is being analyzed. Optionally, the process of analyzing synthetic data may differ from the routine process of analyzing experimental data in terms of the steps performed. In other words, in addition to replacing the experimental data with a synthetic dataset, the analysis process may include additional modifications relative to the routine process of analyzing experimental data.

[0029] One or more variance datasets are selected based on one or more types of noise associated with the experimental data to be analyzed. This should be understood as: one or more received variance datasets are selected or have been selected taking into account one or more types of noise associated with the experimental data to be analyzed. In other words, one or more variance datasets can be selected in a contextualized manner, such that the context or association between the variance datasets and the experimental data to be analyzed is taken into account. One or more received variance datasets can be dependent on one or more types of noise associated with the experimental data in such a way that the variation in experimental observations affected by the information contained in one or more variance datasets originates from the same type of noise, which is expected or known to typically affect the experimental data to be analyzed.

[0030] Machine learning models can be understood as algorithms based on artificial intelligence. They can be understood as algorithms that define an artificial system that, as part of a learning phase, has learned relevance from examples or based on training data, and can generalize these relevances after the learning phase ends (so-called machine learning). This algorithm can have a statistical model based on the training data. This means that the aforementioned machine learning does not memorize examples from the training data, but rather identifies patterns and regularities in the learning or training data. In this way, the algorithm can also evaluate unknown data outside the training dataset (so-called learning transfer or generalization).

[0031] A machine learning model can have at least one artificial neural network. An artificial neural network can be based on several interconnected units or nodes, called artificial neurons. Connections between neurons can transmit information (sometimes called signals) to one or more other neurons. An artificial neuron receives input information, processes it, and can then output information to one or more neurons connected to it based on the processing of the input information. Input and output information can be real numbers, so the output information of a corresponding neuron can be calculated by a function (especially a nonlinear function) according to the sum of its input information. Connections can also be called edges. Neurons and edges can have weights, which are adjusted during the learning process or training procedure using a loss function to be optimized. Weights increase or decrease the output information. Neurons can have thresholds, such that output information is sent only when the output information exceeds the threshold. Neurons can be grouped into layers. Different layers can perform different transformations on their inputs. Information can be sent from the first layer (input layer) to the last layer (output layer), possibly via intermediate layers, or possibly after passing through one or more layers several times. More specifically, neurons can be arranged in multiple layers, where neurons in one layer can (optionally only) connect to neurons in the immediately preceding and immediately following layers. The layer that receives external data is the input layer. The layer that produces the final result is the output layer. There can be zero or more hidden layers in between. A single-layer network can also be used. There are several connection patterns between two layers. Two layers can be fully connected, meaning every neuron in one layer can connect to every neuron in the next layer. However, layers can also be connected through "pooling," where a group of neurons in one layer connects to a single neuron in the next layer, thus reducing the number of neurons in the next layer. Networks with only this type of inter-layer connection form a directed acyclic graph and are called feedforward networks. Alternatively, networks that allow connections between neurons in the same layer or previous layers are called recurrent networks.

[0032] However, synthetic datasets are not limited to machine learning applications. The generated synthetic datasets can be used for basic statistical analysis purposes, such as calculating the mean, standard deviation, and range of values ​​after introducing additional levels of variability into some experimental data. The generated synthetic dataset data can also be used to perform more accurate outlier detection, data clustering, and / or data visualization.

[0033] This disclosure provides the following advantages: it can improve the reliability of experimental data analysis. In particular, when analyzing experimental data with only a very limited number of samples or even just one sample, this disclosure can improve the reliability of the analysis process when performed on a synthetic dataset compared to performing the analysis based on a first dataset.

[0034] Furthermore, this disclosure provides the following advantages: experimental data can be enriched by a real data point distribution that is expected for this type of experimental data but is not shown in the experimental data due to the small number of available samples.

[0035] This disclosure provides the following advantages: it can facilitate experimental data analysis using machine learning models. In particular, this disclosure can improve the accuracy and / or reliability of experimental data analysis (e.g., experimental data classification) even when only a small number of reference samples of experimental data are used to train the machine learning model.

[0036] This disclosure provides the following advantages: although the data analyzed may be subject to a higher level of variation than the potentially limited first dataset used to train the machine learning model, the machine learning model is able to analyze the experimental data with high accuracy and / or high reliability.

[0037] This disclosure also provides the advantage that knowledge about the variation in experimental data that may affect the analysis can be considered and applied during the training of the machine learning model. This allows the machine learning model to cope with experimental data that has a higher degree of variation than the reference sample of experimental data used to train the machine learning model. Therefore, this disclosure provides the advantage that the machine learning model can be robust and can be used for out-of-distribution generalization.

[0038] One or more variance datasets can be selected such that the variance in the experimental observations contained in one or more datasets corresponds to noise that typically affects experimental data of the same type as the experimental data included in the first dataset. In other words, the selected noise can be chosen to match noise that typically affects the experimental data to be analyzed; however, this noise may not exist due to the small number of available samples. This can improve the accuracy of experimental data analysis.

[0039] One or more variant datasets can be selected in a contextualized manner based on the context associated with the experimental data to be analyzed and / or the context associated with the analysis of the experimental data to be performed. Context can arise from the type of experimental data and / or the preparation of the samples used to obtain the experimental data and / or the measurements used to obtain the experimental data. Different processes may introduce different kinds of noise. Context dependence can specifically consider variant datasets that reflect the types of noise expected to affect or already affecting the samples and / or the processes used to obtain the experimental data.

[0040] One or more variant datasets received are selected and / or provided at least in part by an external source. Optionally, the selection can be performed by a user and can be provided by means of user input. Optionally, an external computing device (e.g., a remote server) can be configured to select variant datasets to be received for analysis of experimental data. The selection of variant datasets to be received may involve examining the experimental data to be analyzed. This provides the advantage of allowing efficient selection of variant datasets to be received.

[0041] The method may further include: determining one or more noise types to be considered when analyzing measurement data before receiving one or more variant datasets; and selecting one or more variant datasets to be received based on the determined noise types. In other words, the selection of one or more variant datasets may at least partially constitute part of the method for analyzing experimental data. In other words, the selection of one or more variant datasets may be at least partially performed by a data processing device that performs the method for analyzing experimental data. The method may further include: requesting the selected one or more variant datasets from an external data source. This provides the advantage that adaptive selection of one or more variant datasets can be integrated into the process of analyzing experimental data through a data processing device that performs the method for analyzing experimental data.

[0042] Experimental observations from one or more variant datasets can be obtained independently of the experimental data included in the first received dataset. In other words, experimental observations from one or more variant datasets can be obtained completely independently of the acquisition of the experimental data to be analyzed. Experimental observations from one or more variant datasets can be obtained from a sample independent of the sample from which the experimental data to be analyzed is obtained. The sample used to obtain experimental observations from one or more variant datasets can be a sample that allows the acquisition of the distribution of data points originating from a specific type of noise. Optionally, the sample used to obtain experimental observations from one or more variant datasets can be a sample that allows the acquisition of the distribution of data points originating from a specific type of noise and has little or no interference from other types of noise besides the specific type of noise. The acquired data points can then be used as experimental observations from one or more variant datasets specifying the specific type of noise. Then, when the specific type of noise should be considered in the analysis of the experimental data to be analyzed, these one or more variant datasets can be selected or received.

[0043] Extrapolating experimental data may include or consist of: adding a measurement point cloud to at least one measurement point in the experimental data, wherein the measurement point cloud is based on at least one of one or more variant datasets. Different types of noise and / or different variant datasets can affect different parts of the experimental data, such as optionally affecting different spectral characteristics of the spectral data to be analyzed. Extrapolation may include: adding a statistical distribution of the data points to one or more specific data points contained in the experimental data.

[0044] One or more variant datasets can be selected such that the experimental observations in one or more variant datasets and the experimental data included in the first dataset belong to the same type of experimental data. The selection of variant datasets can be tailored to reflect this type of noise from samples of the same type from which the experimental data to be analyzed is derived. This allows for the addition of accurate statistical distributions of data points when extrapolating experimental data.

[0045] The type of experimental data may involve spectral data. Experimental data may include data acquired through at least one of the following techniques:

[0046] - Vibrational spectroscopy;

[0047] - Optical spectroscopy in the visible and / or ultraviolet spectral range;

[0048] - X-ray diffraction method;

[0049] - NMR spectroscopy;

[0050] - Mass spectrometry;

[0051] - Clinical chemistry testing;

[0052] - Electrophysiological methods; and

[0053] - Flow cytometry.

[0054] Different types of noise can affect different spectral features of the spectral data that form part of the experimental data. Therefore, different variations in the dataset can lead to variations in the different spectral features of the experimental data. However, some types of noise, and corresponding variations in the dataset, can affect all spectral features of the experimental data being analyzed.

[0055] Experimental data can be obtained from biofluids in biological organisms. Biofluids can include at least one of the following types:

[0056] - Blood;

[0057] - Blood plasma;

[0058] - Serum;

[0059] - Saliva;

[0060] - Urine or urine exprimate.

[0061] - Cerebrospinal fluid;

[0062] - Live or chemically fixed animal, human, or plant tissues;

[0063] - Live or chemically fixed bacteria, animal, human, and plant cells;

[0064] - Viruses or viral particles; and

[0065] - Multimolecular structure.

[0066] The variation in experimental observations within a specific variability dataset, which may be contained in one or more variability datasets, can originate from one of the following sources of variation:

[0067] - Tolerance for the reproducibility of quantitative measurements used to obtain experimental observations.

[0068] - Pre-analytical variability that occurs during the preparation of one or more samples used to obtain experimental observations;

[0069] - Used to obtain the variation between different phenotypes from which one or more samples for experimental observations originate;

[0070] - The inherent variation within the biological system under study, including variation between different individuals and variation within individuals over time.

[0071] Variation can arise from differences in sample storage, such as differences in storage duration and / or storage conditions. Variation can arise from differences in sample handling, such as handling by different machines and / or human operators and / or differences in handling duration and / or handling conditions. Variation can arise from differences in obtaining experimental data from samples, such as differences in the equipment used to perform measurements and / or differences in the settings of the equipment used to perform measurements and / or differences in the performance of the equipment used to perform measurements.

[0072] The number of synthetic data points can be greater than the number of at least one reference sample included in the experimental data in the first dataset. This allows for the provision of a broader synthetic dataset for training machine learning models than the first dataset, which is not possible when using the first dataset to train machine learning models.

[0073] The received one or more variation datasets can be selected such that the variation of the experimental observations contained in the one or more variation datasets corresponds to the actual and / or expected variation affecting the experimental data to be analyzed. This provides the advantage that the training data can be adapted to the variation affecting the experimental data to be analyzed.

[0074] Variation affecting the experimental data to be analyzed can originate from one or more of the following sources of variation: (i) the tolerance of the quantitative measurement reproducibility of the measurement technique used to obtain the experimental data; (ii) pre-analytical variability that occurs during the preparation of one or more samples used to obtain the experimental data; (iii) variation between different phenotypes from which one or more samples used to obtain the experimental data originate; and (iv) variation between sample types from which the experimental data were extracted; and (v) biological variation between different biological systems or within the same biological system over time.

[0075] Each variant dataset may include experimental observations affected by variants originating from one or more predetermined variant sources. Alternatively or additionally, each variant dataset may optionally include experimental observations affected only by variants originating from one or more predetermined variant sources.

[0076] Each variant dataset may include experimental observations affected by variants originating from one or more predetermined variant sources, wherein the experimental observations are substantially unaffected by variants originating from variant sources other than the predetermined variant sources.

[0077] Experimental data may include data obtained by at least one of the following techniques: (i) vibrational spectroscopy, (ii) optical spectroscopy in the infrared, visible and / or ultraviolet spectral range, (iii) X-ray diffraction, (iv) NMR spectroscopy, (v) mass spectrometry, (vi) clinical chemistry testing, (vii) electrophysiological methods and (viii) flow cytometry.

[0078] Experimental data can be obtained from biofluids of biological organisms. Biofluids may include at least one of the following types: (i) blood, (ii) plasma, (iii) serum, (iv) saliva, (v) urine or expelled urine after squeezing, (vi) cerebrospinal fluid, (vii) live or chemically fixed animal, human, or plant tissues, (viii) live or chemically fixed bacterial, animal, human, or plant cells, (ix) viruses or virus particles, and (x) multimolecular structures. Multimolecular structures may include at least one of the following: (i) amyloid proteins, (ii) antibodies, (iii) antigens, and (iv) other molecular complexes.

[0079] Reference samples of experimental data included in the first dataset can be obtained from the same biological fluid type as the experimental data to be analyzed by a machine learning model trained using one or more synthetic datasets. Alternatively, reference samples of experimental data included in the first dataset can be obtained from a different biological fluid type than the experimental data to be analyzed by a machine learning model trained using one or more synthetic datasets. The machine learning model can be trained using training data with variations that reflect these possible variations in the experimental data.

[0080] One or more variant datasets can be selected such that the experimental observations contained in one or more variant datasets are affected by the same and / or similar types of variants as the experimental data to be analyzed.

[0081] The variation of experimental observations in a particular variation dataset contained in one or more variation datasets may originate from one of the following sources of variation: (i) the tolerance of quantitative measurement reproducibility of the measurement technique used to obtain the experimental observations, (ii) pre-analytical variability that occurs during the preparation of one or more samples used to obtain the experimental observations, and (iii) the variation between different phenotypes from which one or more samples used to obtain the experimental observations originate.

[0082] Analyzing the variation of experimental observations included in one or more variation datasets may include: analyzing the variation of experimental observations in each of the one or more variation datasets independently of each other. Alternatively or additionally, determining the variability characteristics based on the analyzed variation of experimental observations may include: determining the variability characteristics based on the experimental observations in each of the one or more variation datasets independently of each other.

[0083] Applying the determined variability characteristics to the first dataset may include extrapolating each of at least some data points from one or more data points of a reference sample of experimental data included in the first dataset to multiple synthetic data points, based on the variability characteristics.

[0084] Synthetic datasets can be provided as training data, which are used to train machine learning models that analyze experimental data based on out-of-distribution generalization.

[0085] The variation in the experimental data to be classified can correspond to the variation in one or more synthetic datasets used to train a machine learning model, wherein the one or more synthetic datasets are based on a first dataset and one or more mutated datasets, the first dataset including at least one reference sample of the experimental data and a predetermined analysis result of the at least one reference sample, and the one or more mutated datasets including experimental observations that have undergone variation, the variation at least partially affecting the experimental data to be classified. The machine learning model can be a machine learning model according to this disclosure.

[0086] The disclosure provided for methods used to analyze experimental data should be regarded as being made public for computer programs, computer-readable media, data processing devices, methods for generating training data, methods for training machine learning models for analyzing experimental data, methods for analyzing experimental data using machine learning models, computer-implemented machine learning models, computer-implemented methods for classifying experimental data, methods for identifying individuals from samples of biological fluids based on individuals, and methods for generating data for analyzing experimental data, and vice versa.

[0087] Those skilled in the art will understand that the features described above, as well as those in the following description and drawings, are not only disclosed in the explicitly disclosed embodiments and combinations, but also other technically feasible combinations and individual features are included in this disclosure. Hereinafter, several alternative embodiments and specific examples are described with reference to the accompanying drawings to illustrate this disclosure, but not to limit this disclosure to the described embodiments.

[0088] The following disclosure presents further alternative implementations, optional features, and possible advantages of this disclosure without limiting its scope. Attached Figure Description

[0089] Figures 1A to 9 depict methods according to alternative embodiments of the present disclosure.

[0090] Figure 10 schematically depicts an overview of the problem scenario and CODI methodology according to an alternative embodiment.

[0091] Figures 11A to 11D schematically illustrate the application of CODI to introduce measurement variability into exemplary experimental IR spectra.

[0092] Figures 12A to 12E illustrate CODI's enhanced personalized fingerprint recognition through more accurate long-term molecular mapping analysis.

[0093] Figures 13A to 13F illustrate CODI’s ability to achieve classification flexibility across biological sample variants.

[0094] Figures 14A to 14C show the CODI that restored lost categorical power on independent case-control test sets.

[0095] Figure 15 illustrates the impact of the modeled sources of variability on classification accuracy in longitudinal molecular surveillance.

[0096] Figures 16A and 16B illustrate a comparison of CODI with domain-agnostic enhancement schemes. Detailed Implementation

[0097] The accompanying drawings are described in detail below.

[0098] Figure 1 schematically illustrates a computer-implemented method 100 for analyzing experimental data according to an alternative embodiment.

[0099] Method 100 includes: receiving 102 a first dataset, the first dataset including experimental data to be analyzed.

[0100] Method 100 further includes receiving 104 or more variation datasets, wherein each variation dataset includes information about the variation of experimental observations. The received one or more variation datasets are selected based on one or more types of noise associated with the experimental data to be analyzed. The one or more variation datasets can be selected such that the variation of experimental observations contained in the one or more datasets corresponds to noise that typically affects experimental data of the same type as that included in the first dataset. The received one or more variation datasets can be selected in a contextualized manner based on the context associated with the experimental data to be analyzed and / or based on the context associated with the analysis of the experimental data to be performed.

[0101] Method 100 further includes: extrapolating 106 experimental data by applying variation to experimental data, wherein the applied variation corresponds to one or more variations of experimental observations specified by information contained in one or more variation datasets. Extrapolating experimental data may include or consist of: adding a measurement point cloud to at least one measurement point of the experimental data, wherein the measurement point cloud is based on at least one of one or more variation datasets.

[0102] In addition, method 100 includes: generating a synthetic dataset 108 based on extrapolated experimental data, analyzing the synthetic dataset 110, and providing output 112 based on the results of the analysis of the synthetic dataset.

[0103] The received one or more variant datasets are selected based on one or more types of noise associated with the experimental data to be analyzed.

[0104] One or more variant datasets received may be selected and / or provided at least in part by external sources.

[0105] One or more variant datasets can be selected such that the experimental observations of one or more variant datasets and the experimental data included in the first dataset are of the same type of experimental data.

[0106] Method 100 may further include: determining one or more noise types to be considered when analyzing measurement data before receiving one or more variant datasets, and selecting one or more variant datasets to be received based on the determined one or more noise types.

[0107] Experimental observations from one or more variant datasets can be independent of experimental data included in the first received dataset.

[0108] The type of experimental data may involve spectral data. Experimental data may include data acquired through at least one of the following techniques:

[0109] - Vibrational spectroscopy;

[0110] - Optical spectroscopy in the visible and / or ultraviolet spectral range;

[0111] - X-ray diffraction method;

[0112] - NMR spectroscopy;

[0113] - Mass spectrometry;

[0114] - Clinical chemistry testing;

[0115] - Electrophysiological methods; and

[0116] - Flow cytometry.

[0117] Experimental data can be obtained from biofluids in biological organisms. Biofluids can include at least one of the following types:

[0118] - Blood;

[0119] - Blood plasma;

[0120] - Serum;

[0121] - Saliva;

[0122] - Urine or urine expelled after being squeezed;

[0123] - Cerebrospinal fluid;

[0124] - Live or chemically fixed animal, human, or plant tissues;

[0125] - Live or chemically fixed bacteria, animal, human, and plant cells;

[0126] - Viruses or viral particles; and

[0127] - Multimolecular structure.

[0128] The variation in experimental observations within a specific variability dataset, which includes one or more variability datasets, can originate from one of the following sources of variation:

[0129] - Tolerance for the reproducibility of quantitative measurements used to obtain experimental observations;

[0130] - Pre-analytical variability that occurs during the preparation of one or more samples used to obtain experimental observations;

[0131] - Used to obtain the variation between different phenotypes from which one or more samples for experimental observations originate.

[0132] Figure 1B schematically depicts a data processing apparatus 120 according to an alternative embodiment, wherein the data processing apparatus is configured to perform the method according to Figure 1A.

[0133] Figure 2 schematically illustrates a computer-implemented method 200 for generating training data to train a machine learning model for analyzing experimental data, according to an optional implementation.

[0134] Method 200 includes: receiving 202 one or more variant datasets, wherein each variant dataset includes experimental observations that have undergone variation, which at least partially affects experimental data to be analyzed. The received one or more variant datasets may be selected such that the variation in the experimental observations contained in the one or more variant datasets corresponds to the actual variation and / or expected variation affecting the experimental data to be analyzed.

[0135] Method 200 further includes: analyzing the variation of experimental observations contained in one or more variation datasets, and determining the variability characteristics based on the analyzed variation.

[0136] Method 200 further includes: receiving 206 a first dataset, the first dataset including at least one reference sample of experimental data and predetermined analysis results for the at least one reference sample.

[0137] Method 200 further includes: 208 generating one or more synthetic datasets comprising multiple synthetic data points by applying the determined variability characteristics to the first dataset.

[0138] The number of synthetic data points can be greater than the number of at least one reference sample included in the experimental data in the first dataset.

[0139] Variation affecting the experimental data to be analyzed can originate from one or more of the following sources:

[0140] - Tolerance for the reproducibility of quantitative measurements in measurement techniques used to acquire experimental data;

[0141] - Pre-analytical variability that occurs during the preparation of one or more samples used to obtain experimental data; and

[0142] - The variation between different phenotypes from which one or more samples are derived to obtain experimental data;

[0143] - Variations between sample types from which experimental data are extracted.

[0144] Each variant dataset may include experimental observations affected by variants originating from one or more predetermined variant sources. Alternatively or additionally, each variant dataset may optionally include experimental observations affected only by variants originating from one or more predetermined variant sources.

[0145] Each variant dataset may include experimental observations affected by variants originating from one or more predetermined variant sources, wherein the experimental observations are substantially unaffected by variants originating from variant sources other than the predetermined variant sources.

[0146] Experimental data may include data acquired through at least one of the following techniques:

[0147] - Vibrational spectroscopy;

[0148] - Optical spectra in the visible and / or ultraviolet spectral range;

[0149] - X-ray diffraction method;

[0150] - NMR spectroscopy;

[0151] - Mass spectrometry;

[0152] - Clinical chemistry testing;

[0153] - Electrophysiological methods; and

[0154] - Flow cytometry.

[0155] Experimental data can be obtained from biofluids in biological organisms. Biofluids can include at least one of the following types:

[0156] - Blood;

[0157] - Blood plasma;

[0158] - Serum;

[0159] - Saliva;

[0160] - Urine or urine expelled after being squeezed;

[0161] - Cerebrospinal fluid;

[0162] - Live or chemically fixed animal, human, or plant tissues;

[0163] - Live or chemically fixed bacteria, animal, human, and plant cells;

[0164] - Viruses or viral particles; and

[0165] - Multimolecular structure.

[0166] Reference samples of experimental data included in the first dataset can be obtained from the same biological fluid type as the experimental data to be analyzed by a machine learning model trained using one or more synthetic datasets. Alternatively, reference samples of experimental data included in the first dataset can be obtained from a different biological fluid type than the experimental data to be analyzed by a machine learning model trained using one or more synthetic datasets.

[0167] One or more variant datasets can be selected such that the experimental observations contained in one or more variant datasets are affected by the same and / or similar types of variants as the experimental data to be analyzed.

[0168] The variation in experimental observations within a specific variability dataset, which may be contained in one or more variability datasets, can originate from one of the following sources of variation:

[0169] - Tolerance for the reproducibility of quantitative measurements used to obtain experimental observations;

[0170] - Pre-analytical variability that occurs during the preparation of one or more samples used to obtain experimental observations;

[0171] - Used to obtain the variation between different phenotypes from which one or more samples for experimental observations originate.

[0172] Analyzing the variation of experimental observations included in one or more variation datasets may include: analyzing the variation of experimental observations in each of the one or more variation datasets independently of each other. Alternatively or additionally, determining the variability characteristics based on the analyzed variation of experimental observations may include: determining the variability characteristics based on the experimental observations in each of the one or more variation datasets independently of each other.

[0173] Applying the determined variability characteristics to the first dataset may include extrapolating each of at least some data points from one or more data points of a reference sample of experimental data included in the first dataset to multiple synthetic data points, based on the variability characteristics.

[0174] A synthetic dataset is provided as training data, which is used to train a machine learning model that analyzes experimental data based on out-of-distribution generalization.

[0175] Figure 3 schematically illustrates a computer-implemented method 300 for generating training data to train a machine learning model for analyzing experimental data, according to an optional implementation.

[0176] Method 300 includes: receiving 302 a training dataset including training data, wherein the training data includes one or more synthetic datasets having the same or similar mutations as one or more variant datasets, the one or more variant datasets being used to generate one or more synthetic datasets based on a first dataset, the first dataset including at least one reference sample of experimental data and predetermined analysis results for the at least one reference sample.

[0177] Method 300 also includes: training a machine learning model using the received training dataset.

[0178] Figure 4 schematically illustrates a computer-implemented method 400 for analyzing experimental data using a machine learning model, wherein the machine learning model is configured to classify the analyzed mutated experimental data, the mutation of which exceeds the mutation of a reference sample of experimental data used as a first dataset to generate training data for training the machine learning model.

[0179] Figure 5 schematically depicts a computer-implemented method 500 for analyzing experimental data using a machine learning model, wherein the machine learning model is trained according to the method in Figure 4.

[0180] Figure 6 schematically illustrates a data processing device 600 configured to perform a method according to any one of Figures 2 to 4.

[0181] Figure 7 schematically depicts a computer-implemented method 700 for classifying mutated experimental data using a machine learning model, wherein the mutation of the experimental data to be classified exceeds the mutation of the reference sample of the experimental data used to generate training data to train the machine learning model.

[0182] The variation in the experimental data to be classified may correspond to the variation in one or more synthetic datasets used to train a machine learning model, wherein the one or more synthetic datasets are based on a first dataset and one or more mutated datasets, the first dataset including at least one reference sample of the experimental data and a predetermined analysis result of the at least one reference sample, and the one or more mutated datasets including experimental observations that have undergone variation, which at least partially affects the experimental data to be classified.

[0183] Figure 8 schematically depicts a method 800 for identifying individuals from samples of individual-based biological fluids according to an alternative implementation.

[0184] Method 800 includes: receiving experimental data from a sample of biological fluid characterizing an individual.

[0185] The method further includes determining the analysis results of the received experimental data by using a machine learning model according to an alternative embodiment of the present disclosure, the machine learning model being trained on experimental data of a reference sample of biological fluid derived from an individual, wherein the analysis results correspond to a match or mismatch between the received experimental data of the characterizing sample and the experimental data of the reference sample used to train the machine learning model.

[0186] Method 800 further includes providing information 806 about the identification of an individual based on the matching or mismatch between the received experimental data of the representation sample and the experimental data of the reference sample used to train the machine learning model.

[0187] Figure 9 illustrates a computer-implemented method 900 for generating data for analyzing experimental data, according to an optional implementation.

[0188] Method 900 includes: receiving 902 one or more variant datasets, wherein each variant dataset includes experimental observations that have undergone mutations that at least partially affect the experimental data to be analyzed.

[0189] Method 900 further includes: analyzing the variation of experimental observations contained in one or more variation datasets, and determining the variability characteristics based on the analyzed variation.

[0190] Method 900 further includes: receiving 906 a first dataset, the first dataset including at least one reference sample of experimental data and predetermined analysis results for the at least one reference sample.

[0191] Method 900 further includes generating 908 one or more synthetic datasets comprising multiple synthetic data points by applying the determined variability characteristics to the first dataset.

[0192] Other optional details and alternative embodiments of this disclosure are described below without limiting the scope of this disclosure.

[0193] We present and empirically test a hybrid experimental and computational modeling framework. We explore OOD generalization within the context of molecular analysis and propose treating the variance generated in the analytical procedure as a component of real-world observations (Figure 10). We introduce Contextualized External Distribution Integration (CODI), a modeling framework that, in a seemingly paradoxical way, incorporates measurement variability and the inherent complexity of biological systems, transforming them into valuable, exploitable properties. CODI first involves experimental data to assess their distributional characteristics. Following characterization, we generate synthetic data through computer simulations, intentionally incorporating these distributional characteristics into the independent datasets under study. These synthetic data can model one or more systems of interest while extending the distribution of the original seed data to include information about the sources of variance for critical OODs and missing values.

[0194] Figure 10 schematically depicts an overview of the problem scenario and CODI methodology according to an alternative embodiment. Given a biological, medical, or molecular system of interest, the task is to classify different groups of samples (e.g., healthy observations versus diseased observations). However, several biological and (pre)analytical variations arising from the empirical workflow can affect the captured measurements in different ways. CODI utilizes independent sources of variability in its representations and incorporates them into a training set of labeled experimental observations. This process generates a larger simulated sample set with a more representative data distribution. Training an ML classifier on the simulated samples enables it to learn decision boundaries for classifying samples in a more fully informed manner, thereby increasing the likelihood of generalization to unseen test samples. This approach makes the sample representations more robust to variations in the empirical workflow.

[0195] To establish this concept and evaluate it in a real-world setting, we applied our approach to experimental infrared (IR) spectral data to aid in in vitro blood-based diagnostics. An advantage may lie in cross-molecular fingerprinting, where quantitative analysis of measurements captures the breadth of variation in the molecular spectra of complex samples as an indicator of overall health and disease. We tested our approach within the framework of three independent longitudinal clinical studies (with follow-up periods spanning up to 8 years) [25–27] and a case-control study detecting four common cancers

[28] . Our results demonstrate that integrating CODI into the classification pipeline enables the creation of larger, more representative datasets, allowing ML algorithms to more effectively capture reproducible signals in biological datasets. Finally, we show how the proposed framework can significantly improve classification outputs on unseen, independently measured test samples, ensuring robust predictions despite variations in data distribution.

[0196] Characterizing the variability observed empirically

[0197] We previously introduced computer simulation models for generating 1D spectra of complex biological samples, focusing on infrared (IR) absorption spectra

[29] . Our preliminary work explored the impact of different levels of biological variability among individuals on classification efficiency under simulated case-control conditions. Building on this, we extend the model beyond the theoretical framework. We generate data simulating longitudinal and / or case-control settings, considering a wide variety of possible sources of variability that we have characterized experimentally. To evaluate its practical application, we explore the modeling framework’s ability to computationally generate larger training sets that are more robust to biological aspects and to analyzing variability. CODI is a simple statistical procedure. Essentially, the method can rely on capturing the differences (u) between a set of experimental observations. i - v j These are then referred to as calibration measurements (e.g., control samples of repeated measurements under different laboratory conditions). By selecting calibration measurements that exhibit a level of variability from a given pool of measurements, we can scale the differences using random variables that follow a probability distribution and combine them over the entire set of calibration measurements utilized. This aggregation can then be added to the individual experimental measurements xxxk to create new simulated measurements. This simulation method can be repeatedly applied to generate queues of measurements of arbitrary size. The generated queues, as a whole, will hold measurements in u i and v j The variability observed between them is reflected in x k Above. If the calibration measurement set reflects the training set of given measurements. If a new source of variability is not observed in the data, then a new level of variability will be introduced into the training set for that measurement. This strategy allows for the creation of realistic synthetic data without adjusting any free parameters controlling the generation of the data (except for the number of observations generated). A detailed description of this modeling approach is provided in the Methods section and Supplementary Information below.

[0198] In our CODI implementation, we introduced several different sets of calibration measurements to model the various sources of variability that may be observed in IR spectral measurements in blood-based media (Figure 11A). Among these calibration measurements are empirical variability characteristics stemming from intrinsic biological factors, variability in sample collection and processing, and instrument-specific measurement noise and drift (Methods, Supplementary Information).

[0199] When biological variability is involved, the calibration measurement value u i v jExperimental measurements selected for the same individual over time capture the level of intra-individual biological variability (Fig. 11A, top left), as previously described

[27] . Alternatively, selecting calibration measurements for different individuals yields the level of inter-individual biological variability, as previously described

[29] . Further, other exemplary variability arising from different clinical sample collection sites (e.g., clinics), clinical study protocols, and sample handling procedures can be effectively represented by selecting calibration measurement characteristics from samples originating from different clinical studies (Fig. 11A, top right). The same concept can be extended to simulate real-world variability arising from experimental procedures such as sample storage temperature and duration, dispensing procedures, and measurement equipment drift. For example, quality control (QC) samples can be subjected to a wide variety of handling and storage conditions. Performing QC measurements under different operating conditions of the measuring equipment, including, for example, recalibration, routine maintenance, or changes in the surrounding environment, will enable the QC measurement dataset to simulate potential variability in laboratory procedures and instrument drift (Fig. 11A, bottom left). Furthermore, independent measurements of technical repetitions (e.g., pure water) performed over extended time periods help to facilitate a clearer distinction between instrument noise and laboratory variability (Figure 11A, bottom right).

[0200] The primary goal of characterizing the diverse sources of variability is to realistically simulate the data distributions that may be encountered in an empirical workflow. It is important to recognize that this is achieved by utilizing measurements that are independent of the original training set and unrelated to the specific problem posed (i.e., class-invariant). Similar calibration sets can be adapted to characterize variability relevant to the conditions under investigation when this concept is extended to other molecular systems or measurement modalities.

[0201] Introducing experimental variability into computer simulations as an exemplary embodiment, we applied the CODI modeling framework to five experimental spectra of plasma spectra to generate a larger set of simulated measurements that reflects an increased level of variance (Fig. 11B). The five raw spectra can be considered as the training set, where each measurement represents a labeled category (Fig. 11B, left). Using the five measurements as seed inputs, the CODI framework allows us to generate a larger and more representative training set of measurements (Fig. 11B, right). Principal component analysis (PCA) applied to the raw and simulated measurements shows that the source of variability for each measurement has a different impact on the spatial distribution of the seed data in the first two components (Fig. 11C). In other words, each set of calibrated measurements (which models different variability characteristics) influences linearly independent data features. Once all four sources of measurement variability are incorporated into the raw seed data (Fig. 11C, right), the simulated measurements occupy a larger point cloud while still maintaining their distinct cluster centroids. Similarly, examining the standard deviation between the empirical sources of variability modeled for each modeled source and the simulated measurements revealed that the standard deviation of the simulated measurements was higher than that of each individual empirical source of variability (Figure 11D). While it may seem counterintuitive that simulated datasets with increased variability compared to existing experimental observations can provide added value, this variance contains valuable, usable information. This principle relies on the assumption that simulated measurements represent unexplained (out-of-distribution) fluctuations within a smaller set of measurements that are likely to occur when a larger set of experimental observations is available.

[0202] Figures 11A through 11D schematically illustrate the application of CODI to introduce measurement variability into illustrative experimental IR spectra. Figure 11A: Four different possible sources of variability are characterized from the calibration set of experimental observations. Figure 11B: By applying CODI, four sources of variability are introduced into five illustrative blood-based experimental spectra (left panel) to generate a larger set of simulated spectra (right panel). Figure 11C: Principal component analysis (PCA) is performed on the five exemplary experimental spectra (left panel), the larger set of simulated spectra with only one of the four characterized sources of variability introduced (middle panel), and the set of simulated spectra with all four sources of variability introduced into the five illustrative experimental spectra (right panel). Figure 11D: The standard deviation of each characterized source of variability within the fingerprint spectral range is compared to the standard deviation of the simulated measurements (including all four sources of variability), resulting in an increase in the overall standard deviation.

[0203] We have now demonstrated how CODI can be used to introduce additional sources of variability into existing measurement datasets, enriching their information content. The following sections will examine the value of this method in other applications.

[0204] Application in longitudinal research environments

[0205] In longitudinal clinical studies, which involve the collection of biological samples destroyed by attrition and loss to follow-up over time, a significant effort can be required to collect a sufficiently large dataset. Typically, individuals participate in an initial baseline sample collection, followed by an extended waiting period for subsequent collections from the same individuals. When only a small sample is initially available for each individual, extrapolating meaningful insights to follow-up collections and measurements becomes challenging due to the dynamic nature of the empirical procedure, as described above. To examine whether our proposed method provides added value when the available sample for analysis is severely limited, we employ the CODI framework in the context of longitudinal analysis (Figure 12).

[0206] We first utilized three independent clinical studies that tracked individuals over extended periods (Figure 12A). The Lasers4Life-LG study cohort

[27] consisted of 31 individuals who repeatedly donated blood samples at irregular follow-up intervals. The study began with a 7-week baseline monitoring period during which 288 samples were collected through repeated donations. Three additional donations followed the initial baseline donation period, spanning up to 4.5 years, during which one sampling point was considered for each individual. In the BioPersMed study

[25] , a subcohort of 44 individuals participated repeatedly over an 8-year follow-up period with 2-year intervals between donations. In the KORA study

[26] , a subcohort of 2015 individuals participated in two donations with a 6.5-year follow-up interval. Plasma was processed from all samples and measured by IR absorption spectroscopy (Methods section).

[0207] Figures 12A through 12E illustrate CODI's enhancement of personalized fingerprinting through more accurate long-term molecular profiling analysis. Figure 12A: Setup of three independent longitudinal clinical studies in which the same individuals repeatedly participated in venous blood sampling over time. Experimental IR spectroscopic measurements of plasma were performed and used in this analysis. Figure 12B: Individual identification power in the three study cohorts using only a single baseline IR measurement for each individual. The black bars depict classification accuracy using experimental (“Exp.”) baseline measurements for training. The blue bars depict accuracy using simulated (“Sim”) measurements generated by CODI (using each individual's baseline IR measurement as seed input). For both training methods, the classifier was tested on the same experimental follow-up measurements for each individual. Figure 12C: Dependence of identification accuracy on the number of individuals in the cohort (i.e., the number of classes). Individuals were randomly selected at different cohort sizes, and CODI was applied using a single baseline IR measurement for each individual as seed input. The classifier was tested on experimental follow-up measurements for each individual included in the training set. Figure 12D: Dependence between recognition accuracy and follow-up time axis in applications involving Lasers4Life-LG and BioPersMed cohorts. Figure 12E: Modeling with an increasing number of experimental baseline measurements for each individual as training seeds. The right-hand bars depict recognition accuracy on the simulated training set, while the black bars depict recognition accuracy on the experimental training set. Model testing was performed on the same experimental follow-up after the first six baseline measurements for each individual. The figure below shows the difference in accuracy observed between the experimental and simulated training sets.

[0208] Enabling long-term molecular mapping analysis

[0209] The concept of identifying individuals based on different biological fluids has been demonstrated using IR spectroscopy, NMR spectroscopy, and mass spectrometry as fingerprinting patterns [27, 30-32]. We previously demonstrated that plasma and serum-based IR fingerprints can identify individuals over a period of up to 6 months

[27] . This application inherently relies on the stability of measurements during the study period, and despite unavoidable experimental drift, it is often necessary to compare measurements obtained at different times

[33] . In previous studies [27, 30-32], several measurement data points from the same individual were required to adequately train a multi-class classifier to identify individuals (typically 8-42 samples). We set out to test the potential value of our CODI framework in enhancing longitudinal studies, using the individual identification task as a readout metric. To test the limitations of the framework, we considered a scenario where training data was severely limited, relying only on a single experimental observation for each class, i.e., one measurement for each individual (Fig. 12B). We used the first baseline measurement for each individual to train a multi-class classifier to identify individuals from their subsequent follow-up measurements. We then compared this predictive efficiency with a classifier trained on a set of simulated measurements created using the CODI framework, which uses the experimental baseline as initial seed data. Within the CODI framework, we modeled the four sources of variability previously described (Figure 11A), characterized by data sources independent of the experimental seed data (Supplementary Information). This allowed us to generate 1000 simulated measurements for each individual to train the classifier. This study showed that the classifier trained on simulated measurements had a significantly improved predictive efficiency compared to the classifier trained on experimental measurements (Figure 12B). Individual identification accuracy improved from 0.53 to 0.82, from 0.26 to 0.76, and from 0.03 to 0.52 in the three cohorts, respectively. This demonstrates that (informationally) incorporating data variance does indeed enable the classifier to generalize better to unseen test samples. To examine which sources of data variability were most helpful in improving classification efficiency, we re-performed the above analysis but systematically eliminated one of the four sources of variability incorporated into the CODI modeling framework (Figure 15). We found that biological variability within individuals over time and measurement variability in the same quality control samples during the measurement process are the most critical factors for classification success. Variability stemming from technical replication and clinical sampling has the least impact on classification compared to the two sources mentioned above. It is important to emphasize that the population sizes differ between cohorts. Therefore, the accuracy decline observed in the KORA cohort (involving 2015 individuals) is not directly comparable to that in the Lasers4Life-LG cohort (involving only 31 individuals). This is because the larger the dataset, the more likely their fingerprints are to overlap, making the task of identifying individuals more challenging.Nevertheless, it is quite surprising and encouraging to observe that when combined with the proposed modeling method, and requiring only a single baseline venous blood sample per individual, nearly half of the 2015 individuals could be identified from the IR molecular fingerprint. The observation that identification accuracy decreases with increasing population size prompted us to further investigate this dependency (Figure 12C). We first trained the classifier on simulated measurements using the first experimental baseline measurements from only two individuals, and tested it on follow-up as previously described. This procedure was repeated several times using two additional randomly selected individuals. We then performed the same procedure on individuals of 4, 8, 16, etc. The analysis shows that identification accuracy follows a near-perfect logarithmic trend with population size. This interesting finding, derived from information theory, can be further explored to quantify the information content of different molecular fingerprints. Notably, these results are reproducible on three independent cohorts, indicating that similar identification accuracy can be achieved when the dataset involves similar population sizes. In the above application, follow-up measurements from all individuals were pooled together, and the identification accuracy was averaged, independent of the follow-up time axis. This raises the question: Does individual identification accuracy depend on the time interval between follow-up measurements and baseline? In other words, is it more difficult to identify individuals 8 years after baseline evaluation than at a 2-year follow-up? To investigate this, we grouped the data according to the time difference between follow-up measurements and baseline and examined whether any temporal trend was observed in identification accuracy (Fig. 3d). This analysis was only possible in the Lasers4Life-LG and BioPersMed cohorts, as the available KORA cohorts only involved one follow-up. Here, we reveal that identification accuracy does not depend on the length of time between follow-up and baseline. Even within an 8-year follow-up period, accuracy remained relatively stable. While it is important to recognize that the number of samples tested in later follow-up years is limited (Fig. 12A), this is the first experimental result regarding fingerprint stability over such a long period. In summary, the above research was made possible by using the CODI framework, which enables applications previously infeasible with limited experimental observations.

[0210] Comparison of Domain-Independent Augmentation Schemes

[0211] The CODI modeling framework inherently relies on prior information about the potential sources of measurement variability. In contrast, domain-agnostic augmentation methods employ general transformations on the input seed data to model new observations (e.g., introducing additive or multiplicative noise). By eliminating the need for prior information, this strategy is easier to implement than CODI. To examine whether CODI is superior to other augmentation methods, we re-performed the above analysis by applying several methods to augment spectral measurements (Supplementary Figures 16A and 16B). We found that the CODI method for generating the training set significantly outperformed all other data augmentation methods. This highlights the value of incorporating contextualized prior information into the data augmentation process to enable classification to generalize beyond the original training set.

[0212] Personalized multi-baseline modeling

[0213] In our exploration of the value of the CODI framework, we relied on a single baseline measurement for each individual. We then ask: to what extent can the classification be robust when more training instances are available for each category? Specifically, considering that a single baseline measurement may be an outlier, we examine the dependence between the number of training instances for each individual and recognition accuracy (Figure 12E). To properly study this analysis, individuals must be repeatedly sampled within a given baseline monitoring period. Of the three clinical studies, only the Lasers4Life-LG study facilitated this setting

[27] . Here, we used data from the first 6 baseline measurements for each individual from the Lasers4Life-LG cohort as experimental seeds for training. We then simulated a training set consisting of 1000 measurements for each baseline and investigated how recognition accuracy depended on how many experimental baselines were modeled for each individual.

[0214] The classifier was tested on the remaining follow-ups after the first 6 baselines for each individual. This study shows that the identification accuracy following the simulation-based approach can indeed be improved when modeling for more than one baseline measurement for each individual (Fig. 12E, right bar). An accuracy close to 0.85 was achieved when only one baseline measurement was used. Surprisingly, including just one additional baseline resulted in a near-perfect improvement in predictive efficiency, achieving an accuracy of 0.96. As a comparable benchmark, we again examined the dependency between identification accuracy and the number of baselines modeled for each individual, but this time training the classifier directly on experimental measurements (Fig. 12Ee, left bar). The classifier was first trained on one experimental baseline for each individual, and then on two, up to six, baselines for each individual. Testing was performed on the remaining follow-ups after the first 6 training baselines for each individual, as described above. Here, the simulation-based approach again demonstrates a significant advantage over the experimental modeling approach—but only when few observations are available for each class (≤3 baselines per individual). Once ≥4 experimental baselines are available for training for each individual, the experimental modeling method achieves near-perfect predictive efficiency, and therefore CODI no longer shows an advantage. This highlights the impact of our proposed modeling paradigm in environments with only limited experimental datasets. Once sufficiently large experimental datasets are available, simulation-based training methods may not be more advantageous than training directly on experimental data. In summary, these findings suggest that the CODI framework can establish a more reliable “baseline” for each individual—one that is more resistant to analysis and biological variation and can more robustly generalize to ML. We further demonstrate that IR molecular fingerprints are highly stable and individual-specific. Previously, this had only been demonstrated over a 6-month timeframe

[27] . In this study, we extend these findings to a medically relevant 8-year timeframe. These results lay the foundation for future applications of blood-based IR fingerprints as a personalized monitoring modality for human health over time, which may require only one or two baseline samplings.

[0215] Cross-sample generalization

[0216] Molecular mapping applications involve the use of a wide variety of sample specimens—such as serum or plasma as cell-free products of whole-body blood (Figure 13A). The selection of appropriate specimens is typically done during the study design phase, considering factors such as ease of collection and biological relevance [34, 35]. However, preemptive selection can be limiting, as insights gained from one sample may not generalize when transferred to another. For example, suppose a dataset of plasma spectra is available. Subsequently, unlabeled spectra derived from serum samples need to be classified and compared. This raises an interesting question: how well would a classifier trained on plasma spectra perform when tested on serum spectra? The straightforward answer is that classification is likely to fail due to the underlying molecular differences between samples

[36] . Effective classification requires training instances from diverse samples, each sufficiently representative to capture the distribution of a particular class—a highly resource-intensive process. As a proof of principle, we demonstrate here the potential versatility of the CODI framework for such applications while minimizing the need for large-scale biological dataset collection.

[0217] Plasma and serum IR spectra share many features due to their relatively similar molecular spectra (Fig. 13B). The main variation stems from the plasma preparation process, which involves the use of ethylenediaminetetraacetic acid (EDTA), as demonstrated in previous work

[27] . To achieve effective classification flexibility between samples, their differences must be adequately characterized. This can be achieved by calculating the differences between experimental plasma and serum measurements from the same collected blood samples (Fig. 13C). Following the CODI framework, we incorporated this characterized difference into an experimental dataset of independent plasma spectra to generate simulated spectra that resemble mixtures between samples (Fig. 13D). Next, we again employed the task of identifying individuals as readout indicators to evaluate the ability to generalize across samples. We adopted the Lasers4Life-LG cohort (Fig. 13E), in which serum and plasma were obtained simultaneously from the same individual at the time of all blood sampling donations. The datasets were split into training and test sets, with the training set consisting of 4–12 donations per individual and the test set consisting of the remaining follow-up donations. As a baseline, we first examined the effectiveness of training two classifiers, one trained on experimental plasma spectra and the other on simulated plasma spectra, both tested on plasma spectra (Fig. 13F, left). This study confirms the aforementioned finding—that, with sufficiently large experimental training sets, both experimental and simulation-based classifiers perform similarly. We then applied the same classification procedure as above, but this time tested on serum measurements (Fig. 13F, right). For this analysis, we employed the CODI framework to generate a training set of plasma / serum mixtures. Notably, we found a significant advantage of the simulation-based approach over the experimental approach. For the classifier trained on experimental plasma, a significant drop in prediction efficiency was observed, reaching an accuracy of 0.34. The classifier trained on the simulated mixture data almost completely recovered its initial classification efficiency—achieving an accuracy of 0.89. This unexpected finding demonstrates that the CODI framework can create datasets that are robust even to variations in biological sample characteristics.

[0218] In summary, this proof-of-principle analysis further demonstrates the potential of the CODI framework to overcome conventional limitations in biological and biomedical research, thereby enabling new applications. First, it eliminates the need to recollect large numbers of samples when sample collection procedures go awry. One can characterize the differences between samples using a limited set of measurements collected in a class-independent and task-independent manner. The CODI framework can then be used to extend ML applications to different sample variations. However, here we have only demonstrated this potential on serum and plasma. If the molecular composition and physiological responses of the samples differ greatly (e.g., blood-based mediators versus urine or saliva-based mediators), the approach may not be as effective as shown in this paper. Promising avenues for future exploration may involve adapting classifiers trained on EDTA plasma to citrate samples

[37] .

[0219] Figures 13A through 13F illustrate the CODI framework for achieving classification flexibility across biological sample variants. Figure 13A: Plasma and serum are collected as cell-free products of whole venous blood at the same sampling site. Figure 13B: Experimental spectra of measurements from several plasma and serum samples from the same individual. Figure 13C: Calculation of differences between plasma and serum spectra processed from the same whole blood sample to indicate characteristic variations between samples. Figure 13D: The CODI framework enables the generation of simulated spectra of plasma / serum mixtures by utilizing characteristic variations between samples as a calibration set. Figure 13E: Setup of the Lasers4Life-LG cohort, where the same individual repeatedly participates in venous blood sampling over time. Donations are divided into training and test sets, and plasma and serum are processed from all donations. Figure 13F: Individual identification accuracy based on the Lasers4Life-LG cohort for training and testing. The left figure depicts the accuracy of classifiers trained on experimental and simulated plasma fingerprints—both classifiers were tested on experimental plasma fingerprints. The right figure depicts the accuracy of classifiers trained on experimental plasma fingerprints and simulated fingerprints of plasma / serum mixtures—both classifiers were tested on experimental serum fingerprints.

[0220] Generalization of independently acquired datasets

[0221] An important aspect of determining how well a medical diagnostic assay can perform is testing it on unseen samples. In biodiagnostic applications, cross-validation procedures are commonly used to obtain estimates of the true (external) classification performance. However, if there are biases in the collected datasets, such as confounding information caused by the measurement “batch effect,” the estimated performance may not be reproducible when the classifier is actually externally validated

[38] . To test this in relevant medical applications, we consider our previous work

[28] , in which four cancer types were classified against asymptomatic, cancer-free controls. In contrast to our previous work using cross-validation

[28] , the samples here were initially split into training and test sets for independent measurements (Fig. 14A, left panel). As a baseline, we first relied on a cross-validation procedure to investigate the performance of classifying each cancer type (Fig. 14B, left bar). Cross-validation was performed only on the training set of the experimental samples and the receiver operating characteristic (ROC) curves for validating the splits for each cancer type were examined. Next, we trained the classifier on the training set of the experimental samples and tested it on the test set of the experimental samples (Fig. 14B, right bar). This study shows that the classification efficiency for the four cancer types decreases when tested on samples measured later. This validates the previous view that cross-validation estimation cannot be fully reproduced using this training-test split classification setting.

[0222] Next, we pose the question: can the CODI framework help in this scenario? In principle, by introducing class-invariant empirical variability into the training set of measurements, we can actually make the classifier's learning task more difficult. Potentially, this would allow the classifier to appropriately weight features that are more robust to measurement artifacts, making them dependent on information that might recur in unseen data. To test this, we introduced an additional level of variance (Supplementary Information) into the training set of measurements using the CODI framework. In four cancer types, we revealed that an improvement in predictive efficiency was indeed observed when the classifier was tested on the held-out test set (Figure 14B, middle bar). The most impressive improvement was in the lung cancer application, where the area under the ROC curve (AUC) was almost completely recovered, comparable to the previous cross-validation estimate. For the remaining cancer types, the CODI framework still provided an advantage, albeit to a lesser extent than in lung cancer. This may be partly attributed to measurement artifacts occurring in the training set that happened to be relevant to the outcome of interest, causing the AUC estimate to be overly optimistic during cross-validation. This may also be partly due to the typically small sample size used to test the classifier, and the fact that the cases and controls included in the randomly selected test samples are more difficult to distinguish than those in the training set (e.g., due to inherent physiological variations that interfere with cancer signals). However, the CODI framework consistently delivers improved classification outputs on independently measured test sample data compared to training directly on experimental observations.

[0223] The effect of experimental training cohort size To investigate under what conditions the CODI framework enables more robust classifier training, we repeated the case-control study described above, but changed the number of experimental observations used for training (Fig. 14C). In addition to the previous cancer application, we also addressed a case-control application from the KORA cohort

[39] —focusing on detecting common health physiological states (Fig. 14A, right). We randomly selected samples at different cohort sizes to train several classifiers on each subset of the selected samples (Supplementary Information). First, we trained directly on the experimental observations and always tested on the reserved test set (Fig. 14C, black curve). Unsurprisingly, the smaller the training set, the worse the classifier performed on the experimental test set. Then, we used the CODI framework to generate a simulated dataset that used the experimental observations of each sample count as seed input (Fig. 14C, gray curve). Here, we showed that CODI almost always outperformed the experimental modeling method across different sample counts available as a basis for training. Notably, for the detection of dyslipidemia and type 2 diabetes, both diseases exhibited strong molecular biases in the IR fingerprint, with CODI providing the greatest improvement when the number of available training samples was small. For the detection of prediabetes and hypertension, no significant advantage was observed in incorporating CODI into the classification workflow. This may stem from the fact that the experimental training data were already very similar to the test data distribution, or because the variability introduced by CODI failed to effectively capture the distributional bias present in the test data. While no advantage was observed for these two diseases, including CODI did not adversely affect classification. This observation suggests that integrating CODI into the classification workflow may be an effective standard practice—as it either improves predictive performance or at least does not impair it.

[0224] In summary, our findings demonstrate the promising potential of the proposed CODI framework. The value of this method has been proven in several biomedical applications, where it has enabled improved ML classification outputs for numerous practical applications.

[0225] Figures 14A through 14C illustrate the CODI for recovering lost classification power on independent case-control test sets. Figure 14A: Eight binary classifications covering a wide variety of health conditions were set up in two independent clinical studies. Plasma IR spectroscopy was performed on all samples. Training and test sets were measured under different measurement device conditions and in different measurement activities. Figure 14B: Cancer detection was studied in three different settings for evaluating classification efficiency. For simulation-based training, the CODI framework was used to incorporate measurement variability into the training set measurements. ROC curves (top) for each cancer detection applied to the validation partition of the data, and the estimated AUC (bottom) are plotted. Figure 14C: Classification efficiency when training the classifier on experimental observations (black curve) and when applying the CODI framework to train the classifier on simulated data (grey curve), using different training sample counts as the training basis. Classifier testing was performed only on the reserved experimental observations.

[0226] discuss

[0227] Multimolecular mapping and computational modeling offer promising avenues for advancing our understanding of biological systems. In this study, we introduce CODI, a modeling framework designed to enrich collected datasets to facilitate robust analysis that enhances the probing capabilities of systems. We rigorously experimented with and tested the framework in several experimental settings within the context of molecular mapping analysis to demonstrate its effectiveness. We examined how different empirical and biological factors contribute to the variability in data measured via IR spectroscopy and revealed the framework’s strengths in overcoming the limitations of unrepresentative observational datasets. Indeed, the datasets generated by CODI enable ML algorithms to better capture the latent information present in the studied datasets, guiding them to distinguish which features are most relevant and reproducible. This strategy is particularly valuable when collecting large, representative datasets is limited. This is exemplified in the context of studying pathophysiological phenomena through molecular mapping analysis (e.g., omics). In such cases, biological experiments or medical research require significant input, including probing large numbers of subjects, substantial time commitments for phenotypic evolution, and additional constraints related to complex sample collection and processing [1–4, 40, 41]. Another layer of complexity arises from biological variation at the organismal level, which is inherent and unavoidable—due to the dynamics of biological systems and human physiology (e.g., circulation, renewal, rhythmic oscillations, aging) [4, 42–46]. The quantitative measurement procedures often involved further exacerbate these challenges. Factors such as wear and tear on measurement equipment components, routine maintenance, and sensitivity to environmental conditions can all contribute to the known “batch effect,” which is often specific to quantitative analysis methods [5, 6, 13, 47, 48]. Ultimately, these challenges hinder the generalization of insights to unseen, later-collected and measured samples. The concept of CODI, which aims to facilitate generalization, draws on data augmentation techniques that rely on modifying existing experimental observations to create synthetic data. This technique has wide applications in image classification, where data augmentation includes geometric transformations (e.g., rotation, skewing, cropping), color adjustments, and the introduction of random noise into the training set of images to capture latent variation in real-world applications

[23] . Data augmentation has also been applied to biosignal measurements from electroencephalography (EEG), electromyography (EMG), Raman spectroscopy, and near-infrared (IR) spectroscopy [49-56]. While existing methods typically involve the introduction of random noise, signal warping, and decomposition of the available dataset, our approach extends the concept of data augmentation. We utilize additional independent measurements that were not initially part of the original dataset of interest to model the real-world empirical variation of the processes involved in the molecular analysis.

[0228] In our study, we discussed the topic of barely supervised learning

[57] —where the labeled training sample set for each category is limited to a very small number of observations. Given the experimental limitations of population sampling over timeframes of several years, we examined whether CODI could minimize the number of sequential samples of the same individual over time. Using IR fingerprinting as an example, we were surprised to find that only a single baseline measurement was needed to adequately follow up and identify individuals in the population at later time points. Although identifying the same individual in a heterogeneous population is only a rough approximation of identifying physiologically relevant biases, it provides a fundamental complement to the concept of longitudinal probing. Our general framework can be rapidly adopted to potentially save unnecessary sampling, thus informing future prospective studies.

[0229] Further applications of CODI in infrared spectral (IR) fingerprinting demonstrate its versatility and potential impact in assisting data analysis. We observed significant comparability of experimental data collected over the past decade, highlighting the method's ability to improve classification power on independently measured test sets. The framework's adaptability extends to proof-of-principle applications involving training a classifier on one sample medium (plasma) and applying it to another (serum). This application demonstrates the potential of the CODI framework to streamline cross-sample dataset analysis in a variety of biological and biomedical applications. This strategy can be particularly valuable when collecting datasets from retrospective studies or online repositories to help ensure comparability of samples with another envisioned application (i.e., domain-adaptive applications) [58, 59]. Another promising use of CODI could be to harmonize data from different measurement devices (e.g., spectrometers manufactured by different manufacturers), potentially improving the transferability of ML models between them.

[0230] It is important to emphasize that our proposed method should not be positioned as an alternative to improving research design and better standardizing classical analytical procedures. These aspects remain crucial when establishing molecular mapping analysis platforms to meet the iID assumption between training and test datasets. Beyond efforts to ensure the comparability of training and test datasets, our method is motivated by a concern that this assumption may be violated due to unavoidable sources of error and variability. An inherent limitation of the proposed method is its reliance on prior knowledge about possible empirical sources of variability. This presents a challenge, as gathering such information may require extensive experimental evaluation of biological and analytical variability. Furthermore, for any successful application, the characterized sources of variability must adequately represent the true possible range of empirical variability. Therefore, continuous optimization through more controlled experiments to include several independent sources of variance has the potential to further enhance the framework's practicality. However, once the range of possible variances has been successfully characterized for a given system and measurement procedure, the same characterization can be reused in different applications. While our research primarily focuses on blood-based IR spectroscopy to aid in in vitro diagnostics, the practical applications of the CODI framework are not limited to this. The principles and mathematical foundations of this framework are general enough to be translated into examinations of various biological systems, medical problems, measurement modalities, and ML tasks. Applications involving NMR spectroscopy, mass spectrometry, and Raman spectroscopy are direct extensions that can be explored using this modeling framework. Furthermore, the framework shows promise in applications related to cell typing and the integration of single-cell multimodal omics data—given the inherent challenges associated with obtaining accurate measurements of cell type populations at scale, where out-of-distribution measurement events are prevalent [60, 61].

[0231] In summary, the proposed framework lays the foundation for future explorations to enhance the robustness of molecular analysis.

[0232] Materials and methods

[0233] The supplementary information below provides a more detailed description of the CODI modeling framework, our application, dataset, experimental procedures, and ML analysis. The methods disclosed according to the alternative implementation (which can be described as a computer simulation model behind CODI) can be designed for scalability, allowing them to be applied to a variety of applications and different measurement modalities. The modeling framework involves utilizing seed observations. These represent the intrinsic properties of the phenomenon of interest. For example, seed observations can represent the average observation of one type of sample (e.g., healthy controls), while another corresponds to a different category (e.g., disease samples). By adding random functions f1, f2...f... m Introducing variability into seed observations s i The measured values ​​are modeled as the following generalized statistical variable Y:

[0234] .

[0235] Therefore, repeatedly applying the above model will generate a queue of simulated measurements of arbitrary size, with the queue numbered in s... i Centered on, and incorporating f1, f2...f m The introduced mutations. From f1, f2...f m The introduced variance can be characterized by either ab initio or bottom-up modeling, each representing a source of expected data variance. However, the former is often specialized and problem-specific. Alternative descriptive approaches that rely on collecting a dataset of calibration measurements containing the expected level of variance can be readily applied to a variety of problems. For example, quality control samples may be subjected to varying durations of cryopreservation, number of freeze / thaw cycles, and aliquots by different operators. The quality control samples can then be measured repeatedly under different measurement equipment conditions. The variance observed in this calibration dataset will reflect potential sources of variance in the processing and measurement of samples from the original dataset modeling a phenomenon (e.g., biological fluids in cases and controls). This variance can then be introduced by defining a function f1. Other potential sources of data variance (e.g., biological variability) can be modeled using additional calibration datasets that reflect the sources of variability.

[0236] Four separate study cohorts were used in our CODI application: Lasers4Life-LG

[27] , BioPersMed

[25] , KORA

[26] , and Lasers4Life-Cancer

[28] . Lasers4Life-LG involved collecting serum and plasma from 31 healthy individuals, initially sampled up to 13 times over 7 weeks, with additional follow-up at 6 months, as detailed in a previous publication

[27] . Since that initial publication, the same individuals have been invited to participate in two additional sample donations at 3.5 and 4.5 years after their initial participation. BioPersMed is a population-based cohort study underway at the Medical University Graz, Austria

[25] . Participants were retested at 2-year intervals. In this study, we utilized plasma samples and medical data from 44 healthy individuals (out of 1022 participants). KORA is a population-based cohort study in southern Germany

[26] . The cohort consisted of a sample of participants stratified by age and sex, randomly selected from resident registries within the study area. In this study, we utilized plasma samples and medical data from the second and third visits (designated KORA-F4 and KORA-FF4, respectively). The available KORA-F4 data consisted of 3,044 samples, while the KORA-FF4 data consisted of 2,140 samples. A subset of 2,015 individuals participated in both samplings, while 1,154 individuals participated in only one sampling. Lasers4Life-Cancer is a case-control cohort involving several cancer types, collecting both serum and plasma

[28] . The samples used in this study largely overlap with those in our previous studies involving detection in four cancer types (lung, prostate, bladder, and breast cancer)

[28] —although the measurement procedures differ (Supplementary Information). Plasma and serum samples from different individuals have been collected since the initial publication and included in this study. Case samples were collected prior to cancer-related treatment (i.e., no treatment). Asymptomatic controls were matched with cancer cases by age, sex, and body mass index. Experimental measurements of liquid samples were performed on a Fourier transform infrared (FTIR) spectrometer, as in previous studies [27, 28, 39]. Samples were injected into flow cells, and IR spectra were recorded in transmission mode. The measurement equipment underwent routine maintenance, with some components replaced as needed during measurements of all clinical samples (Supplementary Information). For applications involving training with a single observation per category, multi-class classification for individual identification was performed using the nearest neighbor algorithm. For multi-class applications involving training with more than one observation per category, the linear discriminant analysis (LDA) algorithm was applied. For binary classification in case-control applications, logistic regression with ridge penalty was used.All categorical power metrics are reported on the reserved test sample.

[0237] Additional Information

[0238] Computer simulation model

[0239] In the following sections, we provide the computer simulation model behind CODI and a general framework for the methods of alternative implementations according to this disclosure to facilitate its scalability to other applications and molecular fingerprint modalities. Technical details regarding specific applications of our model are provided below.

[0240] Broadly speaking

[0241] CODI may rely on incorporating information about the potential sources (biological, technological, etc.) of variation that may occur in the experimental setting into the existing set of experimental observations. In, where each x i It is a numerical vector. In its generalized form, the source of variability can be represented by functions f1, f2, ..., f m Each function represents a model of a different aspect of variation. These functions are assumed to be random vectors in the same space as the experimental observations, and to take a probability distribution centered at 0.

[0242] Based on the experimental observation set x i We can extract another set of seed observations derived from the experiment. This modeles the intrinsic properties of subgroups (e.g., different categories) in the original dataset. Using these, the resulting simulated measurements can be modeled as a statistical variable Y, using the following generalized form:

[0243] .

[0244] Therefore, repeatedly applying the above model will generate a queue of simulated measurements of arbitrary size, with the queue numbered in s... i Centered on, and incorporating f1, f2...f m Introduced variants.

[0245] Set seed measurement value

[0246] Define the seed measurement s derived from the experiment. i The input dataset x depends on the available experimental measurements. i And the type of measurement event we want to simulate.

[0247] In a straightforward implementation, we can set s i = x iThis formula allows us to introduce a certain level of variability for each measurement-specific observation. Repeatedly applying the model to each xxxi will create several sets of measurements, each centered on each experimental observation. In another definition, we can set... .

[0248] This allows us to define the experimental set x i The average measurement introduces a certain level of variability, thereby creating a queue of measurements centered on its expected value.

[0249] If x i If all measurements in a dataset reflect a specific group or category of samples (e.g., healthy individuals), then the simulated cohort as a whole will reflect the distribution of measurements for that category of samples. Clearly, when dataset x... i When switching to measurements that represent different categories (such as disease cases), the simulated cohort will reflect a distribution centered on a different outcome.

[0250] Introduction of variability modeling functions

[0251] From functions f1, f2...f m The introduced variance can be characterized by either de novo computation or a bottom-up model, each representing a source of expected data variance. However, the former is highly specialized and problem-specific. Alternative descriptive methods, relying on collecting a dataset of calibrated measurements containing the expected level of variance, can be readily applied to a wide range of problems.

[0252] Assuming an independent calibration dataset It is available and reflects the source of variability for a given measurement. Here, f1 will take the following form:

[0253] .

[0254] By using a Gaussian random variable centered at 0 Scaling and combining l measurements relative to the average value Individual bias, the random function f1 will have the same characteristics as dataset b. i The same variance. This has been formally and empirically demonstrated in previous studies, where f1 produces the level of variance of the measured values ​​observed across different individuals

[29] . If the sources of variance under discussion do not follow a Gaussian distribution, then the random variable It can follow a more suitable probability distribution.

[0255] Other functions f2, f3...f can be similarly evaluated by utilizing different calibration datasets from other sources that reflect the variability of measured values. mModeling was performed. In this study, we used a calibration dataset in an example application of the model that reflects the level of biological variability within or between individuals, as well as several sources of variability for analysis.

[0256] The model variants used in our model application and how the calibration dataset is defined will be described in detail in the following sections of this paper.

[0257] Simulated case-control measurements

[0258] For the analysis of simulated case and control measurements involved in our applications (e.g., cancer detection), the input dataset takes the following form, where superscripts indicate cases or controls:

[0259] .

[0260] Based on Equation 1, two model definitions were applied, and the settings were configured once. One-time setup This yields the following model form:

[0261] .

[0262] Then, the two model variants are applied repeatedly for each disease detection application to generate datasets containing case and control measurements, respectively.

[0263] Simulated longitudinal capture measurements

[0264] For applications involving longitudinal measurements simulating the same individual, the input dataset takes the following form:

[0265] .

[0266] The superscript indicates different individuals, and the subscript indicates the number of visits they received.

[0267] Then, for each individual, we can use their baseline measurements (visit 0) to model a set of measurements centered on that baseline, in the following form:

[0268] .

[0269] In addition to using only baseline measurements, the same concept can be extended to incorporate additional measurements obtained for each individual during subsequent visits. Therefore, the process can generate any number of measurements for each individual—whether based on a single baseline measurement or more.

[0270] Simulating biological variability among individuals

[0271] To model the variation in measurements among different individuals, the calibration dataset b was used.i The dataset should include measurements from different individuals obtained through experiments. Ideally, all measurements captured for different individuals should be performed within a similar experimental workflow, following the same analytical protocol. This helps ensure that the variance of the calibration dataset is primarily driven by biological differences between individuals, excluding the influence of other sources of data variance. Here, f1 can be set to model biological variability between individuals and will be defined in the same way as described in Equation 2.

[0272] Simulating intra-individual biological variability

[0273] To model the variation in measurements of the same individual over time, the calibration dataset b was used. i It should include measurements obtained from several experiments for each individual. Similar to the longitudinal dataset definition provided in Equation 4, we can formulate another calibration dataset containing one entry to model intra-individual bias in the following form:

[0274] .

[0275] In other words, the calibration vector is calculated per individual and can incorporate the entire set of individual-specific biases into a single dataset.

[0276] Using the dataset defined in Equation 7, function f1 can be set to model intra-individual biological variability and will be defined in a manner similar to Equation 2, as follows:

[0277] .

[0278] In this way, function f1 models the intra-individual variability observed in multiple individuals.

[0279] Simulation analysis of variability

[0280] To model the variability that may arise in a given analytical procedure, additional experimental measurements that have undergone experimental variation and are introduced during the analytical procedure can be utilized. This includes variability that may arise from sample preparation, environmental factors, and instrument errors. Typically, describing the reproducibility of collected data can be achieved by performing multiple measurements on replicates. The specific definition of how replicates are handled and measured depends on the analytical workflow and measurement techniques applied.

[0281] Measurements of duplicates can be performed on samples with a high level of technical reproducibility (e.g., water samples) or on quality control samples that exhibit chemical compositional stability similar to that of the intended application (e.g., blood-based samples pooled from thousands of individuals). For example, for blood-based quality control samples, variations can further include duration of cryopreservation, number of freeze / thaw cycles, aliquoting of samples by different operators, different laboratory room temperatures, and storage of samples in tubes from different manufacturers.

[0282] When performing measurements, technical details specific to the measuring equipment can be intentionally altered, such as filter replacement, software version, and cuvette lifespan. Repeated samples can also be performed across multiple measuring devices and environmental conditions (including temperature and humidity) to further enrich the calibration dataset.

[0283] The variability arising from repeated replicate measurements will then be modeled in the same manner as described in Equation 2. For example, f2 could cover all sources of variability in measurements reflected in blood-based quality control samples, while f3 could cover sources of variability in measurements reflected in water measurements.

[0284] These factors are modeled in a way that best simulates the situations that may be encountered in the application, making the given training dataset more resilient to molecular variations or measurement device operating conditions that may arise in the analysis workflow.

[0285] The definition should include which variability functions?

[0286] The successful application of the CODI modeling framework depends on carefully selecting appropriate sources of variability f1, f2...f m This choice depends on the application of interest and the sources of variability in the measurements that are expected to be observed. For example, if the intended application is to model case-control measurements, including inter-individual biological variability is more critical than including intra-individual biological variability. Conversely, if the goal is to model how individual-specific measurements vary over time, including intra-individual biological variability is necessary.

[0287] The sources of analytical error can be independent of the intended application of the analytical procedure (e.g., measurement equipment drift or sample storage variability). Therefore, such variability functions can be included regardless of the application settings.

[0288] During the model validation phase, when experimental test samples are available, it must be ensured that the calibration dataset used does not include any test samples. This prevents any information or statistical properties of the test samples from leaking into the training samples.

[0289] In our applications, variation between quality control samples, variation between different studies, and variation between technically repeated measures are always included in the model. Intra-individual biological variability is included only in longitudinal applications, while inter-individual biological variability is included only in case-control applications. Variation between plasma and serum is included only when a model trained on plasma measurements is transferred to its application on serum measurements.

[0290] The following sections of this article provide a description of which calibration measurements were used for each application.

[0291] Clinical research

[0292] Lasers4Life-LG Research

[0293] The Lasers4Life-LG study cohort consisted of 31 nominally healthy individuals who were initially sampled up to 13 times over 7 weeks, with additional follow-up at 6 months, as detailed in previous publications

[27] . Since that initial publication, the same individuals were invited to participate in two additional samplings at 3.5 and 4.5 years after their initial participation. Eight individuals participated in each of the latter two samplings. Plasma and serum were collected from all participants throughout the study. The study was approved by the Ethics Committee of Ludwig Maximilian University of Munich (LMU), and all participants provided written informed consent (Protocol #17-532).

[0294] BioPersMed Research

[0295] The BioPersMed study is a population-based cohort study being conducted at the Medical University of Graz, Austria.

[25] Participants were retested at 2-year intervals. In this study, we used plasma samples and medical data from 44 nominally healthy individuals out of 1022 participants. All 44 individuals were included in the baseline sampling and followed up for up to 8 years. A further detailed description of the study design has been previously published.

[25] The study was approved by the Ethics Committee of the Medical University of Graz, Austria (EC Nr. 24-224 ex 11 / 12; Project Application No. 4008 22).

[0296] KORA Research

[0297] The KORA study was a population-based cohort study in southern Germany

[26] . The study consisted of age- and sex-stratified participants randomly selected from resident registries within the study area. In this study, we utilized plasma samples and medical data from the second and third participant visits (designated KORA-F4 and KORA-FF4, respectively). The available KORA-F4 data consisted of 3,044 samples, while the KORA-FF4 data consisted of 2,140 samples. A subset of 2,015 individuals participated in both samplings, while 1,154 individuals participated in only one sampling. The mean follow-up interval was 6.5 years. Data collection methods and standardized sample collection have been described in detail elsewhere [26, 62-64]. The KORA-F4 and KORA-FF4 study methods were approved by the Bavarian Chamber of Physicians, Munich (EC number 06068).

[0298] Lasers4Life - Cancer Research

[0299] The samples used for cancer analysis came from the case-control Lasers4Life-Cancer Study. The samples used in this study largely overlapped with those used in our previous detection studies involving four common cancer types (lung, prostate, bladder, and breast cancer)

[28] . Newly collected plasma and serum samples from different individuals were included to increase the sample size and thus improve the statistical robustness of the results. Case samples were collected prior to any cancer-related treatment (i.e., no treatment). Asymptomatic individuals served as controls. For each cancer type, controls were matched with cancer cases by age, sex, and body mass index (BMI). Written informed consent was provided by all participants in accordance with study protocols #17-141 and #17-182, both of which were approved by the LMU Ethics Committee in Munich. The clinical trial was registered in the German Clinical Trials Register (ID DRKS00013217).

[0300] Experimental Procedure

[0301] Infrared spectroscopy

[0302] Infrared spectroscopy measurements of liquid samples were performed using a Fourier transform infrared (FTIR) spectrometer (MIRA analyzer, CLADE GmbH, Esslingen, Germany). After injecting the sample into a flow cell (window material: calcium fluoride) with an optical path length of approximately 8 μm, the IR spectrum of the sample was recorded in transmission mode, followed by the reference spectrum of the transmission medium (i.e., calcium fluoride-saturated water). The IR spectrum was measured at 950 cm⁻¹.-1 With 3050cm -1 Within the spectral range between 4cm -1 The spectrum was obtained at a resolution of 1000. The raw data provided by the MIRA analyzer is the sample absorption spectrum after subtracting the reference spectrum. The resulting infrared spectrum was preprocessed as previously described

[28] .

[0303] The instrument is maintained annually by the manufacturer. The following components, which may affect the IR spectrum, should be replaced during annual or routine user maintenance: light source (after three years); desiccant cartridge (once a year or as needed); flow cell when the optical path exceeds its limit (typically after several thousand samples); 2µm flow cell pre-filter (typically after 200-300 samples).

[0304] Sample processing

[0305] All clinical samples involved were processed into serum or plasma and stored at -80°C. Transport to the measurement laboratory was performed on dry ice and further stored at -80°C. The original samples (0.3 to 1.0 mL) were thawed at 4°C and centrifuged at 2000 g for 10 minutes. The supernatant was aliquoted into measurement tubes (50–100 μL per tube) and refrozen until measurement.

[0306] Purchased mixed human serum was used as the quality control (QC) sample. Order 3 liters (100 mL flasks, BioWest Niagara, France) and store in the original flasks at -80°C. Before use, thaw at 4°C, filter through a 0.45 µm filter, and aliquot into 50–100 µL aliquots, storing at -80°C until use. Each sample measurement batch begins and ends with a QC serum measurement.

[0307] In addition, a QC sample was measured after every five clinical samples. A measurement batch consisted of 25 to 40 samples. Slight variability was measured between flasks despite using the same batch of QC serum, which may be due to differences in storage time and processing techniques.

[0308] To assess the technical variability of the FTIR spectrometer, pure water samples were measured. Although not measured daily, hundreds of water measurements were performed over many years. The resulting bias reflects the technical noise of the measurement process due to the preprocessing of the measurement data by the FTIR instrument (i.e., subtraction of the reference spectrum).

[0309] Measurement of clinical samples

[0310] Samples from the Lasers4Life-LG study were measured in a completely randomized order at all visits, but serum and plasma were measured separately (single-flow chamber). Samples from the BioPersMed study were measured in a completely randomized order (single-flow chamber).

[0311] KORA-F4 and KORA-FF4

[0312] The study samples were measured in two separate measurement events, with an average interval of 2.7 years between each sampling date. Three flow cells were used throughout the two measurement events (KORA-F4: single flow cell; KORA-FF4: two flow cells). For the Lasers4Life-Cancer Study, the entire study was divided into training and test sets before any measurements were performed. Within these sets, the measurement order was completely randomized. The training set samples were measured using the three flow cells. A four-week interval preceded the measurement of the test set samples. During this interval, completely independent measurements were routinely performed.

[0313] Measurements of the test set were performed on a separate (fourth) flow cell to simulate potential measurement drift. All samples were measured in small batches of 25-40 samples plus QC serum, as described above. The equipment was cleaned after each measurement batch according to the manufacturer's recommendations. The instrument was subjected to a performance validation test daily, either before the first batch measurement or after technical issues were resolved.

[0314] Simulation and classification analysis

[0315] Individual identification in longitudinal monitoring

[0316] A simulated dataset of longitudinal measurements was generated using the model form given in Equation 6. For the analyses shown in Figures 12B through 12D, the baseline measurements for each individual were used as the initial seed input. In the analysis depicted in Figure 12E, the model form in Equation 6 was applied iteratively, using up to the first 6 baselines for each individual as seed input.

[0317] Four variability functions were used to introduce variability into each seed input: (1) intra-individual biological variability; (2) clinical sampling variability; (3) quality control variability; and (4) technical reproducibility variability. Intra-individual variability was always modeled using data from only two of the three longitudinal cohorts used (as shown in Figure 12A). Intra-individual variability was characterized from the other two cohorts when individual identification classification was studied in one of the three cohorts. Since intra-individual biological variability relies on follow-up measurements (Equation 8), this step ensured that no information was leaked from the test set (follow-up) into the training set. For each experimental baseline, 1000 simulated measurements were generated to ensure a sufficiently large sample size to reach a plateau in classification performance.

[0318] Prior to classification, data standardization steps (mean 0, standard deviation 1) and principal component analysis (PCA) were applied. For data standardization, the mean and standard deviation were calculated only from the training set to standardize the training and test sets. PCA was applied to the training set—preserving components that explained 99.99% of the total variance. Then, before training, the loading vectors from PCA were applied to the training and test sets. After the preprocessing before classification, the linear discriminant analysis (LDA) algorithm was used for multi-class classification when more than one observation per class was available for training. When only one instance per class was available in the training set, the k-nearest neighbor (KNN) algorithm (k=1) was used for classification. Classifier testing was always performed on follow-up holdout samples that were not used to create the training dataset. Confidence intervals (CIs) were constructed by bootstrap resampling of the test data, following the approach applied in previous work

[65] . The [2.5, 97.5] percentile boundaries were chosen to construct the 95% CI intervals for classification accuracy.

[0319] Generalization between plasma and serum samples

[0320] For applications involving the creation of simulated spectra of plasma and serum mixtures (Figure 13A), a similar simulation and classification workflow, as described in the preceding sections, was applied. However, in addition to the four sources of variability previously described, an additional function was introduced to model the deviation between serum and plasma measurements. The calibration dataset used to characterize the differences between plasma and serum IR spectral measurements (shown in Figure 13C) was calculated from the Lasers4Life-Cancer study, where plasma and serum were collected at the same sampling occasion. An individual identification application was then applied to the Lasers4Life-LG cohort (Figure 13F)—a separate cohort involving individuals different from those in the Lasers4Life-Cancer cohort.

[0321] Classifier testing on independent test sets

[0322] L2-regularized logistic regression was used for binary (case-control) classification analysis. The data were initially split into training and test sets (Figure 14A). Three different settings were used to evaluate classification efficiency. First, 10-fold stratified cross-validation with 10 repetitions was performed on the training set of samples. Classification efficiency on the validation split of cross-validation was evaluated by calculating the area under the receiver operating characteristic (AUC). The AUC was averaged over the validation split and reported along with its standard deviation. In the second setting, the classifier was trained directly on the experimental data. The AUC was then calculated on the unseen, initially retained test set. In the third setting, the classifier was trained on a simulated sample set based on the training set as seed data. The AUC was again calculated on the unseen, initially retained test set. When creating the simulated dataset, four variability functions were used to introduce variability into each seed input: (1) inter-individual biological variability; (2) clinical sampling variability; (3) quality control variability; and (4) technical repetition variability. The data sources used to calculate the level of variability among individuals involve the training set of the samples, as well as samples from other independent studies—that is, samples from sources outside the test sets of all clinical studies utilized. The model form in Equation 4 is applied, where the mean measurements of cases and controls are used as seed measurements for each binary classification task. For each task, 100,000 samples are generated for each class.

[0323] Analysis software

[0324] The analysis was performed using a custom script written in Python (v.3.8.8). The open-source packages NumPy (v.1.21.2), scikit-learn (v.0.24.1), and matplotlib (v.3.5.1) were used.

[0325] 1. Choose the variability function

[0326] The sources of variability we used in our CODI application were characterized by a set of calibrated measurements that reflect the level of data variance. As previously mentioned, these calibrated measurements contained empirical variability characteristics stemming from intrinsic biological factors, sample collection and processing, and instrument-specific measurement noise and drift. Specifically, in our longitudinal analysis, we characterized four main sources of variability: (1) intra-individual biological variability; (2) clinical sampling variability; (3) quality control variability; and (4) technical repeatability variability. Here, we examined how each of these sources of variability contributed to the success of classification. Figure 15 depicts an analysis where the classification task was to identify individuals given a baseline measurement for each individual / category—and test the classification on follow-up measurements. Similar to the analyses depicted in Figures 12A and 12B, this analysis was also performed on three separate longitudinal cohorts. The CODI modeling framework was used to generate simulated training data based on all four sources of variability (Figure 15, left bar). We then systematically removed one source of variability considered in the simulation model (retaining the other three) to examine whether classification accuracy was affected (Figure 15, remaining bars). We found that classification accuracy was significantly affected in the absence of intra-individual biological variability or quality control variability. In other words, including these two sources of variability is crucial for successful classification across several questions. Variability introduced by clinical sampling and technical recalibration measurements had a small impact on classification accuracy. However, no significant loss of classification accuracy was observed when all four sources of variability were included in the simulation model. Therefore, we included all four sources in our analysis.

[0327] Figure 15 illustrates the impact of the modeled sources of variability on classification accuracy in longitudinal molecular surveillance. Each bar represents the classification accuracy achieved using CODI simulation data (with different combinations of included sources of variability).

[0328] 2. Comparative Analysis of CODI-Independent Enhancement Techniques

[0329] The CODI modeling framework inherently relies on prior information about the possible sources of variability in samples / measurements. In contrast, domain-independent augmentation methods employ general transformations (e.g., additive or multiplicative noise) on the input data to model new measurements, thus eliminating the need for prior information. Here, we present a comparative analysis between CODI and other domain-independent methods that generate simulated training sets. Specifically, for spectral datasets, domain-independent augmentation methods typically include introducing additive white noise, randomly scaling intensity by multiplicative factors, linear slope variability, and vertical offset shifts [55,56] (Figure 16A).

[0330] Figure 16B depicts an analysis in which several domain-agnostic augmentation methods were employed to generate simulated training data for a classification task identifying individuals based on a single baseline measurement for each class. As a benchmark, we include the performance of classifiers trained directly on experimental data and simulated data generated via CODI. Each augmentation strategy (including CODI) was applied to the same (seed) experimental training data, generating 1000 measurements per class, and tested on the same follow-up experimental measurements. Unlike CODI, domain-agnostic augmentation methods require tuning of free parameters that control the extent to which they influence the seed input. In the three longitudinal cohorts conducted in this study, we found a significant advantage over all other methods on data generated via CODI.

[0331] While domain-independent augmentation techniques are easier to implement and do not require prior knowledge, they generally offer minimal or no improvement in classification accuracy compared to training directly on experimental data. In contrast to domain-independent augmentation methods, CODI selectively integrates unrepresented variants, allowing the generation of more representative simulated data and enhancing the classifier's ability to generalize beyond the original experimental training set.

[0332] Figures 16A and 16B depict a comparison of CODI and domain-independent augmentation schemes. Figure 16A: Given an experimental spectrum of input seed (top left), several methods are used to generate spectra with additional variance. These methods include CODI (bottom left) and other domain-independent augmentation schemes. White noise is introduced by repeatedly generating random Gaussian vectors and adding them to the input seed (top middle). Multiplicative noise is introduced by repeatedly scaling the input seed with a random scaling factor (bottom middle). An intensity shift is introduced by vertically shifting the intensity to the input seed with a random factor (top right). A slope shift is introduced by repeatedly manipulating the linear slope of the input seed (bottom right). Figure 16B: An individual identification classification task is studied in three study cohorts. Baseline measurements of each individual in the cohort are used as training seed measurements. The classifier is trained on experimental data (left bar), data generated by CODI (second bar from the left), and data generated by domain-independent augmentation schemes (the remaining bars). For each domain-independent augmentation scheme, several parameters controlling data generation are examined, as listed in the legend on the right.

[0333] This disclosure relates to the following documents:

Claims

1. A computer-implemented method for analyzing experimental data, the method comprising: - Receive a first dataset, which includes experimental data to be analyzed; - Receive one or more variant datasets, where each variant dataset includes information about the variants of the experimental observations; - Extrapolate experimental data by applying variation to experimental data, wherein the applied variation corresponds to one or more variations of experimental observations specified by information contained in one or more variation datasets; - Generate a synthetic dataset based on extrapolated experimental data; - Analyze the synthetic dataset; and - Outputs are provided based on the analysis results of the synthetic dataset; Its features are, The received one or more variant datasets are selected based on one or more types of noise associated with the experimental data to be analyzed.

2. The method of claim 1, wherein the one or more variant datasets are selected such that the variants of experimental observations contained in the one or more datasets correspond to noise that typically affects experimental data of the same type as the experimental data included in the first dataset.

3. The method of claim 1 or 2, wherein one or more received variant datasets are selected in a contextualized manner based on the context related to the experimental data to be analyzed and / or based on the context related to the analysis of the experimental data to be performed.

4. The method according to any one of the preceding claims, wherein the received one or more variant datasets are selected and / or provided at least in part by an external source.

5. The method according to any one of the preceding claims, further comprising the following step before receiving the one or more variant datasets: - Identify one or more types of noise to consider when analyzing measurement data; and - Select one or more variant datasets to receive based on one or more identified noise types.

6. The method according to any one of the preceding claims, wherein the experimental observations of the one or more variant datasets are independent of the experimental data included in the received first dataset.

7. The method according to any one of the preceding claims, wherein the extrapolated experimental data comprises or consists of the following: Add a measurement point cloud to at least one measurement point of the experimental data, wherein the measurement point cloud is based on at least one of one or more variant datasets.

8. The method according to any one of the preceding claims, wherein the one or more variant datasets are selected such that the experimental observations of the one or more variant datasets and the experimental data included in the first dataset are of the same type of experimental data.

9. The method of claim 8, wherein the type of experimental data relates to spectral data.

10. The method according to any one of the preceding claims, wherein the experimental data includes data acquired through at least one of the following techniques: - Vibrational spectroscopy; - Optical spectroscopy in the visible and / or ultraviolet spectral range; - X-ray diffraction method; - NMR spectroscopy; - Mass spectrometry; - Clinical chemistry testing; - Electrophysiological methods; and - Flow cytometry.

11. The method according to any one of the preceding claims, wherein the experimental data is obtained from a biological fluid of a biological organism.

12. The method of claim 11, wherein the biofluid comprises at least one of the following types: - Blood; - Blood plasma; - Serum; - Saliva; - Urine or urine expelled after being squeezed; - Cerebrospinal fluid; - Live or chemically fixed animal, human, or plant tissues; - Live or chemically fixed bacteria, animal, human, and plant cells; - Viruses or viral particles; and - Multimolecular structure.

13. The method according to any one of the preceding claims, wherein the variation of experimental observations in a particular variation dataset within one or more variation datasets originates from one of the following variation sources: - Tolerance for the reproducibility of quantitative measurements used to obtain experimental observations; - Pre-analytical variability that occurs during the preparation of one or more samples used to obtain experimental observations; - Used to obtain the variation between different phenotypes from which one or more samples for experimental observations originate; and - The inherent variation within the biological system under study, including variation between different individuals and variation within individuals over time.

14. A computer program comprising instructions that, when executed by a computer, cause the computer to perform the method according to any one of the preceding claims.

15. A computer-readable medium, optionally a computer-readable storage medium and / or data carrier signal, comprising instructions that, when executed by a computer, cause the computer to perform the method according to any one of claims 1 to 13.

16. A data processing apparatus configured to perform the method according to any one of claims 1 to 13.

17. A computer-implemented method for generating training data to train a machine learning model for analyzing experimental data, the method being characterized in that it comprises: - Receive one or more variant datasets, wherein each variant dataset includes experimental observations that have undergone mutations that at least partially affect the experimental data to be analyzed; - Analyze the variation of experimental observations contained in one or more variation datasets, and determine the variability characteristics based on the analyzed variation; - Receive a first dataset, the first dataset including at least one reference sample of experimental data and predetermined analysis results for at least one reference sample; and - By applying the determined variability characteristics to the first dataset, one or more synthetic datasets comprising multiple synthetic data points are generated.

18. The method of claim 17, wherein the number of synthetic data points is greater than the number of at least one reference sample of experimental data included in the first dataset.

19. The method of claim 17 or 18, wherein the received one or more variation datasets are selected such that the variation of experimental observations contained in the one or more variation datasets corresponds to the actual variation and / or expected variation affecting the experimental data to be analyzed.

20. The method according to any one of claims 17 to 19, wherein the variation affecting the experimental data to be analyzed originates from one or more of the following sources of variation: - Tolerance for quantitative measurement reproducibility of measurement techniques used to acquire experimental data; - Pre-analytical variability that occurs during the preparation of one or more samples used to obtain experimental observations; - The variation between different phenotypes from which one or more samples are derived to obtain experimental data; - Extracting variations between sample types from experimental data; and - The inherent variation within the biological system under study, including variation between different individuals and variation within individuals over time.

21. The method of claim 20, wherein each variation dataset includes experimental observations affected by variations originating from one or more predetermined variation sources, and / or Each variant dataset may optionally include experimental observations that are affected only by variants originating from one or more predetermined sources of variants.

22. The method of claim 20 or 21, wherein each variant dataset includes experimental observations affected by variants originating from one or more predetermined variant sources, wherein the experimental observations are substantially unaffected by variants originating from variant sources other than the predetermined variant sources.

23. The method according to any one of claims 17 to 22, wherein the experimental data includes data acquired by at least one of the following techniques: - Vibrational spectroscopy; - Optical spectroscopy in the visible and / or ultraviolet spectral range; - X-ray diffraction method; - NMR spectroscopy; - Mass spectrometry; - Clinical chemistry testing; - Electrophysiological methods; and - Flow cytometry.

24. The method according to any one of claims 17 to 23, wherein the experimental data is obtained from a biological fluid of a biological organism.

25. The method of claim 24, wherein the biofluid comprises at least one of the following types: - Blood; - Blood plasma; - Serum; - Saliva; - Urine or urine expelled after being squeezed; - Cerebrospinal fluid; - Live or chemically fixed animal, human, or plant tissues; - Live or chemically fixed bacteria, animal, human, and plant cells; - Viruses or viral particles; and - Multimolecular structure.

26. The method of claim 25, wherein a reference sample of the experimental data in the first dataset is obtained from the same type of biological fluid as the experimental data to be analyzed by a machine learning model trained using one or more synthetic datasets.

27. The method of claim 25, wherein a reference sample of the experimental data in the first dataset is obtained from a different type of biological fluid than the experimental data to be analyzed by a machine learning model trained using one or more synthetic datasets.

28. The method according to any one of claims 17 to 27, wherein one or more variant datasets are selected such that experimental observations contained in one or more variant datasets are affected by the same and / or similar types of variants as the experimental data to be analyzed.

29. The method according to any one of claims 17 to 28, wherein the variation of experimental observations in a particular variation dataset within one or more variation datasets originates from one of the following sources of variation: - Tolerance for the reproducibility of quantitative measurements used to obtain experimental observations; - Pre-analytical variability that occurs during the preparation of one or more samples used to obtain experimental observations; - Used to obtain the variation between different phenotypes from which one or more samples for experimental observations originate; and - The inherent variation within the biological system under study, including variation between different individuals and variation within individuals over time.

30. The method according to any one of claims 17 to 29, wherein analyzing the variation of experimental observations included in one or more variation datasets comprises: Analyze the variation of experimental observations in each of one or more variation datasets independently; and / or Determining variability characteristics based on the variation of the analyzed experimental observations includes: determining variability characteristics independently of each of the experimental observations in one or more variability datasets.

31. The method according to any one of claims 17 to 30, wherein applying the determined variability characteristics to the first dataset comprises: Based on the variability characteristics, each of at least some data points from one or more data points of the reference sample of experimental data included in the first data is extrapolated to multiple synthetic data points.

32. The method according to any one of claims 17 to 31, wherein a synthetic dataset is provided as training data, the training data being used to train a machine learning model for analyzing experimental data based on out-of-distribution generalization.

33. A computer-implemented method for training a machine learning model for analyzing experimental data, the method comprising: - Receive a training dataset including training data, wherein the training data includes one or more synthetic datasets having the same or similar mutations as one or more variant datasets, the one or more variant datasets being used to generate one or more synthetic datasets based on a first dataset, the first dataset including at least one reference sample of experimental data and predetermined analysis results for at least one reference sample; and - Train a machine learning model using the received training dataset.

34. A computer-implemented method for analyzing experimental data using a machine learning model, wherein the machine learning model is configured to classify the analyzed mutated experimental data, the mutation exceeding the mutation of a reference sample of experimental data used as a first dataset for generating training data for training the machine learning model.

35. A computer-implemented method for analyzing experimental data using a machine learning model, wherein the machine learning model is trained by the method of claim 33.

36. A computer program comprising instructions that, when executed by a computer, cause the computer to perform the method according to any one of the preceding claims.

37. A computer-readable medium, optionally a computer-readable storage medium and / or data carrier signal, comprising instructions that, when executed by a computer, cause the computer to perform the method according to any one of claims 17 to 35.

38. A data processing apparatus configured to perform the method according to any one of claims 17 to 35.

39. A computer-implemented machine learning model for analyzing experimental data, wherein the machine learning model is configured to analyze experimental data subjected to mutations exceeding the mutations of a reference sample of experimental data used as a first dataset for generating training data for training the machine learning model.

40. A computer-implemented machine learning model for analyzing experimental data, wherein the machine learning model is trained by the method of claim 33.

41. The use of the computer-implemented machine learning model according to claim 39 for analyzing experimental data.

42. A computer-implemented method for classifying mutated experimental data using a machine learning model, wherein the mutation of the experimental data to be classified exceeds the mutation of a reference sample of the experimental data used to generate training data to train the machine learning model.

43. The method of claim 42, wherein the variation of the experimental data to be classified corresponds to the variation of one or more synthetic datasets used to train a machine learning model, wherein the one or more synthetic datasets are based on a first dataset and one or more mutated datasets, the first dataset including at least one reference sample of the experimental data and a predetermined analysis result of the at least one reference sample, the one or more mutated datasets including experimental observations subjected to variation, the variation at least partially affecting the experimental data to be classified.

44. The method of claim 42 or 43, wherein the machine learning model is the machine learning model of claim 39 or 40.

45. A method for identifying an individual based on a sample of an individual's biological fluid, the method comprising: - Receive experimental data from samples of biological fluids that characterize an individual; - The analysis results of the received experimental data are determined by using the machine learning model according to claim 22 or 23, the machine learning model being trained on experimental data of a reference sample of biological fluid derived from an individual, wherein the analysis results correspond to a match or mismatch between the received experimental data representing the sample and the experimental data of the reference sample used to train the machine learning model. and - Information about an individual's identification is provided based on the matching or mismatch between the received experimental data of the representation sample and the experimental data of the reference sample used to train the machine learning model.

46. ​​A computer-implemented method for generating data for analyzing experimental data, the method being characterized in that it comprises: - Receive one or more variant datasets, wherein each variant dataset includes experimental observations that have undergone mutations that at least partially affect the experimental data to be analyzed; - Analyze the variation of experimental observations contained in one or more variation datasets, and determine the variability characteristics based on the analyzed variation; - Receive a first dataset, the first dataset including at least one reference sample of experimental data and predetermined analysis results for at least one reference sample; and - By applying the determined variability characteristics to the first dataset, one or more synthetic datasets comprising multiple synthetic data points are generated.