Method and system for classifying observed molecular interactions

A machine learning-based method simplifies and speeds up the classification of analyte-ligand interactions, addressing user-dependent challenges in existing methods by providing rapid and reliable identification of candidate analytes for drug discovery.

JP2026035619APending Publication Date: 2026-03-04CYTIVA SWEDEN AB
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-11-06
Publication Date
2026-03-04

AI Technical Summary

Technical Problem

Classifying analyte-ligand interactions for drug discovery is tedious and user-dependent, especially when dealing with large numbers of samples, as existing methods struggle to consistently identify candidate analytes for further processing.

Method used

A method utilizing a trained machine learning algorithm to classify analyte-ligand interactions by extracting interaction parameters from response datasets, allowing for rapid and less user-dependent evaluation of molecular interactions.

Benefits of technology

The method provides a simpler and faster way to classify analyte-ligand interactions, reducing user dependency and enabling efficient identification of candidate analytes for further processing, particularly in drug discovery experiments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026035619000001_ABST
    Figure 2026035619000001_ABST
Patent Text Reader

Abstract

To provide a method for interacting a first analyte sample set with a ligand, obtaining a response data set, extracting at least one interaction parameter from the response data, and for each analyte sample solution, providing the interaction parameter to a learned machine learning algorithm to classify observation results from an analytical sensor system.SOLUTION: The trained machine learning algorithm classifies each analyte sample solution into at least one quality classification group indicating the interaction between the analyte sample solution and the ligand based on the interaction parameter. The machine learning algorithm is trained using the set of interaction parameters extracted from the response data obtained from the interaction between the second set of analyte sample solutions and the at least one ligand and the at least one quality classification group indicative of the interaction of the analyte sample solution with the ligand.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates generally to methods and systems for classifying observed molecular interactions. [Background technology]

[0002] Analytical sensor systems configured to monitor molecular interactions, such as biomolecular interactions, in real time are often based on label-free biosensors, such as optical biosensors. A representative example of such a biosensor system is the BIACORE® instrument, which uses surface plasmon resonance (SPR) to detect interactions between sample molecules and molecular structures immobilized on a sensing surface. As the sample passes over the sensor surface, the evolution of the binding directly reflects the rate at which the interaction occurs. Sample injection is followed by a buffer flow, and the detector response reflects the rate at which the complex dissociates from the surface.

[0003] A typical output from a BIACORE® system or similar biosensor system is a response graph or detection curve that describes the evolution of a molecular interaction over time, including portions of the association and dissociation phases. This response graph or detection curve is typically displayed on a computer screen and is often referred to as a binding curve or "sensorgram."

[0004] The BIACORE® system and similar sensor systems can be used to determine multiple interaction parameters for molecules used as ligands and analytes, including kinetic rate constants for the binding (association) and dissociation of intermolecular interactions and the affinity of the interactions.

[0005] For example, in the early stages of fragment drug discovery, it is rare to be able to identify candidate analytes from a single binding measurement alone. A binding level screen (BLS) provides an overview of the library of candidate analytes and allows the identification of analytes for further analysis. In a typical BLS, a single concentration of each analyte is run over the target and reference. Through analysis of the response data, potential candidate analytes can be selected for binding strength and binding behavior. In the next step, an affinity screen, potential candidate analytes are tested over multiple concentration series to ensure that binding is concentration-dependent, allowing the affinity of the analyte / ligand interaction to be determined or at least estimated.

[0006] However, classifying analyte samples as hits in binding level screens (BLS) and using them for further processing can be tedious and difficult even for experienced users, especially when hundreds of samples are obtained in a single evaluation. Because different users may classify BLS samples differently (i.e., hits or not), the evaluation of a library of candidate analytes can be diverse and user-dependent. In affinity screen analyses, users may need to exclude some concentrations from the analysis and must decide which model to use to determine the KD (equilibrium dissociation constant between the analyte and the ligand). As with BLS, this can be difficult even for experienced users, and the results obtained from affinity screens can be user-dependent. By performing either or both BLS and affinity screen analyses, analyte candidates can be identified for further processing, for example, in drug discovery experiments. Summary of the Invention [Problem to be solved by the invention]

[0007] One object of the present disclosure is to provide a simpler and faster method, or at least an alternative method, for classifying analyte-ligand interactions, for example to identify candidate analytes for further processing in drug discovery experiments. The method is suitable for isolating a large number of molecular interactions. Another object is to provide an analytical system for performing such molecular interaction classification. [Means for solving the problem]

[0008] The invention is defined by the accompanying independent claims. Non-limiting embodiments will become apparent from the dependent claims, the accompanying drawings and the following description.

[0009] According to a first aspect, there is provided a method for classifying observations from an analytical sensor system configured to observe molecular interactions between an analyte and a ligand, the method comprising obtaining response datasets representative of the time evolution of the molecular interactions, sequentially interacting a first set of analyte sample solutions with the ligand to obtain a corresponding response dataset for each observed molecular interaction, extracting at least one interaction parameter from at least a portion of each obtained response dataset, and providing the extracted at least one interaction parameter for each analyte sample solution to a trained machine learning algorithm, which classifies each analyte sample solution into at least one quality classification group indicative of the interaction of that analyte sample solution with the ligand based on the extracted at least one interaction parameter.

[0010] The machine learning algorithm is trained using a set of interaction parameters extracted from at least a portion of the acquired response dataset obtained from the interaction between the second set of analyte sample solutions and the at least one ligand, and at least one quality classification group indicative of the interaction between the analyte sample solution and the ligand for each analyte sample solution in the second set of analyte sample solutions.

[0011] The observed intermolecular interaction can occur at the sensing surface of a mass-detecting biosensor system, e.g., an optical biosensor, where one molecule (ligand) can attach / immobilize to the sensor surface and another molecule (analyte) passing over the sensor surface, allowing the time evolution of the binding between the analyte and the ligand to be observed. Analytical sensor systems based on other detection principles, such as electrochemical systems, are also possible.

[0012] The first analyte sample solution set may comprise at least one analyte sample solution. The analytes may be, for example, different candidate analytes in a drug discovery assay. Typically, when performing a binding level screen (BLS) or affinity screen assay, hundreds of analyte sample solutions may be analyzed to provide an overview of a library of candidate analytes and identify candidate analytes that are likely to interact with a particular ligand and be relevant for further processing.

[0013] The user of the analytical sensor system can determine which interaction parameters to extract from the response data and how many interaction parameters to extract and provide to the trained machine learning algorithm. Alternatively, the method can predetermine which interaction parameters to extract and how many interaction parameters to extract and provide to the trained machine learning algorithm. Generally, a larger number of interaction parameters results in more reliable classification. The determination of which interaction parameters to extract and the extraction of those interaction parameters can be performed by an expert system or ANN trained with multiple training response data. Depending on the purpose of the analysis and the ligand being analyzed, the number and type of interaction parameters extracted and provided to the trained machine learning algorithm can vary.

[0014] The analyte sample solution is classified by the trained machine learning algorithm into at least one quality classification group indicative of the interaction of the analyte sample solution with the ligand based on at least one extracted interaction parameter.

[0015] In one example, the number of different quality classification groups can be two, with one group containing analytes that interact well enough with the ligand of interest and a second group containing analytes that interact poorly with the ligand of interest, and the group containing analytes that interact well enough with the ligand can be selected for further analysis.

[0016] In another example, the number of quality classification groups could be 100, with group 1 containing analytes that interact poorly with the ligand and group 100 containing analytes that interact very well with the ligand. The user may regulate the stringency of the method, for example, by setting a cutoff at >50, so that only analytes classified into classification groups 51 through 100 are included in further analysis. Again, what constitutes a sufficiently good quality interaction and where the cutoff is set depends on the objectives of the analysis / experiment.

[0017] Classifying an analyte sample solution or series of solutions into at least one quality classification group indicative of the interaction of the analyte sample solution with the ligand may include a first classification into a main group (e.g., relevant for further analysis or not), and a second group providing information as to why the analyte sample solution was classified into that main group, such as uncertain binding, too low or too high refractive index, non-1:1 binding to the ligand, no or very low binding response (below a predetermined cut-off line), high or unexpected gradient in association, slow dissociation, poor mixing, R>Rmax, analyte being a superstoichiometric binder, concentration above KD (equilibrium dissociation constant between analyte and ligand), concentration too low to determine KD, unreliable parameters in steady-state analysis, etc.

[0018] A second set of analyte sample solutions is used when training the machine learning algorithm. The second set of analyte sample solutions can differ from the first set of analyte sample solutions in both the number of analyte sample solutions and the type of analyte. The ligands used when training the machine learning algorithm can be the same as those used to interact with the first set of analyte sample solutions. Alternatively, the ligands used can be different from those used to interact with the first set of analyte sample solutions.

[0019] The second analyte sample solution set can include a plurality of different analyte sample solutions. Preferably, the second analyte sample solution set can include 10 or more, 100 or more, or 500 or more different analyte sample solutions. The greater the number of analyte sample solutions used for training, the better the classification of the analyte sample solutions in the first analyte sample solution set.

[0020] The interaction parameters used to train the machine learning algorithm and how many interaction parameters are used per analyte sample solution may depend on the intended use of the machine learning algorithm. Generally, the more interaction parameters used for training, the better the classification of the analyte sample solutions in the first set of analyte sample solutions. The determination of the interaction parameters and their extraction may be performed by an expert user of the analytical system, or may be performed by an expert system or ANN that has been trained on multiple training response data sets.

[0021] The number and type of interaction parameters used to train the machine learning algorithm may be the same as those provided to the trained machine learning algorithm for the analyte sample solutions in the first set of analyte sample solutions. Alternatively, the number and / or type of interaction parameters used may be different. The machine learning algorithm may be trained with a larger set of interaction parameters than will be used to actually classify the analyte sample solutions in the first set of analyte sample solutions.

[0022] The machine learning algorithm is also provided with at least one quality classification group for each analyte sample solution in the second set of analyte sample solutions, which is indicative of the interaction between that analyte sample solution and the ligand. The determination of the quality classification group for each analyte sample solution in the second set of analyte sample solutions can be performed by an expert user of the analytical system, or can be performed by an expert system or ANN that has been trained on multiple training response datasets.

[0023] In one example, the number of different quality classification groups can be two, with one group containing analytes that interact well enough with the target ligand and a second group containing analytes that interact poorly with the target ligand, with a classification of good enough indicating that the analyte may be suitable for further analysis.

[0024] In another example, the number of quality classification groups can be 100, with group 1 containing analytes that interact poorly with the ligand and group 100 containing analytes that interact very well with the ligand.

[0025] Classifying an analyte sample solution or series of solutions into at least one quality classification group indicative of the interaction of the analyte sample solution with the ligand may include a first classification into a main group (e.g., relevant for further analysis or not), and a second group providing information as to why the analyte sample solution was classified into that main group, such as uncertain binding, refractive index that is too low or too high, non-1:1 binding to the ligand, no or very low binding response (below a pre-determined cut-off line), high gradient in association, slow dissociation, poor mixing, R>Rmax, analyte being a superstoichiometric binder, concentration above KD, concentration too low to determine KD, unreliable parameters in steady-state analysis, etc.

[0026] The machine learning algorithm thus trained then classifies analyte sample solutions from the first set of analyte sample solutions into quality classification groups indicative of the interaction between the analyte sample solution and the ligand. In BLS analyses, where hundreds of analyte sample solutions are analyzed to provide an overview of a library of analyte candidates, the method provides rapid classification of relevant or incompatible analyte candidates for further processing, and in affinity screens, the method provides additional estimates of concentration dependence and affinity (KD). Furthermore, the evaluation of candidate analytes can become less user-dependent, and an experienced user may not be able to assist in the evaluation of observed molecular interactions.

[0027] The at least one interaction parameter extracted from the acquired response data sets of the analyte sample solutions in the first set of analyte sample solutions may comprise 80-100% of the acquired response data.

[0028] The at least one interaction parameter extracted from the acquired response data sets of the analyte sample solutions of the second set of analyte sample solutions and used for training the machine learning algorithm may comprise 80-100% of the acquired response data.

[0029] The at least one extracted parameter extracted from the acquired response data sets for the analyte sample solutions in the first set of analyte sample solutions may comprise any one or more of the following interaction parameters extracted at a given time point in the response data: a) initial binding response, b) late binding response; c) stable initial response, d) stable late response; e) the initial binding response divided by the molecular weight of the analyte; f) the late binding response divided by the molecular weight of the analyte; g) the difference between the late binding response and the early binding response divided by the early binding response; h) Measurement fluctuation of any of the interaction parameters a) to g). i) the measured gradient of any of the interaction parameters a) to g), and / or When the analyte sample solution is composed of multiple analyte sample solutions containing the same analyte but with different analyte concentrations, the interaction parameters include the following: j) Equilibrium dissociation constant calculated using maximum binding capacities from multiple analyte sample solutions. The equilibrium dissociation constant (KD) of parameter j) is obtained from an affinity screen analysis obtained from a series of analyte concentrations (i.e., the same analyte). The maximum binding capacity (Rmax) used can be a constant or a fixed Rmax, where Rmax is the response when all binding sites on the surface are occupied by the analyte. A fitted Rmax is one obtained when Rmax is used as a free fit parameter. A constant Rmax is one where Rmax is not fitted but is substituted into the equation as a constant. This can be used when the fitted Rmax is predicted to be uncertain and a constant Rmax is obtained from a more definitive determination, for example, by adjusting the fitted Rmax with an analyte of higher affinity by Mw (molecular weight).

[0030] The at least one extracted parameter extracted from the second acquired response datasets for the analyte sample solutions of the second set of analyte sample solutions and used for training the machine learning algorithm may comprise any one or more of the following interaction parameters extracted at a given time point in the response data: a) initial binding response, b) late binding response; c) stable initial response, d) stable late response; e) the initial binding response divided by the molecular weight of the analyte; f) the late binding response divided by the molecular weight of the analyte; g) the difference between the late binding response and the early binding response divided by the early binding response; h) Measurement fluctuation of any of the interaction parameters a) to g). i) the measured gradient of any of the interaction parameters a) to g), and / or When the analyte sample solution is composed of multiple analyte sample solutions containing the same analyte but with different analyte concentrations, the interaction parameters include the following: j) Equilibrium dissociation constant calculated using maximum binding capacities from multiple analyte sample solutions.

[0031] The method may further comprise extracting at least one reference-subtracted interaction parameter for each analyte sample solution of the first set of analyte sample solutions and providing the extracted interaction parameter to the trained machine learning algorithm.

[0032] The machine learning algorithm may be further trained using at least one reference-subtracted interaction parameter for each analyte sample solution of the second set of analyte sample solutions.

[0033] The method may further comprise extracting at least one reference interaction parameter for each analyte sample solution of the first set of analyte sample solutions and providing the reference interaction parameter to the trained machine learning algorithm, wherein the at least one reference interaction parameter may comprise any one or more of the following: k) the reference initial combined response; l) late binding response of the reference; m) Reference stable initial response, n) Reference stable late response, o) the difference between the reference late binding response and the reference early binding response divided by the reference early binding response; and / or If the analyte sample solution is composed of multiple analyte sample solutions of the same analyte but with different analyte concentrations, the reference interaction parameters comprise: p) Equilibrium dissociation constant calculated using maximum binding capacities from multiple analyte sample solutions.

[0034] The machine learning algorithm may be further trained for each analyte sample solution of the second set of analyte sample solutions using at least one reference interaction parameter, which may comprise any one or more of the following: k) the reference initial combined response; l) late binding response of the reference; m) Reference stable initial response, n) Reference stable late response, o) the difference between the reference late binding response and the reference early binding response divided by the reference early binding response; and / or If the analyte sample solution is composed of multiple analyte sample solutions of the same analyte but with different analyte concentrations, the reference interaction parameters comprise: p) Equilibrium dissociation constant calculated using maximum binding capacities from multiple analyte sample solutions.

[0035] The method may further comprise extracting at least one negative control interaction parameter obtained from the interaction of the negative control solution with the ligand and providing the at least one negative control interaction parameter to the trained machine learning algorithm.

[0036] The machine learning algorithm may be further trained using at least one negative control interaction parameter obtained from the interaction of the negative control solution with the ligand.

[0037] The method may further comprise extracting at least one positive control interaction parameter obtained from the interaction of the positive control solution with the ligand and providing the at least one positive control interaction parameter to the trained machine learning algorithm.

[0038] The machine learning algorithm may be further trained using at least one positive control interaction parameter obtained from the interaction of the ligand with a positive control solution.

[0039] The machine learning algorithm may be selected from the group consisting of a decision tree, k-nearest neighbors (kNN), random forest, K-means, gradient boosting algorithm, artificial neural network (ANN), deep learning algorithm, or a combination thereof.

[0040] According to a second aspect, there is provided an analytical sensor system for detecting molecular binding interactions between an analyte and a ligand and classifying the observations, the system comprising: a sensor device with a detector for monitoring a molecular interaction between a ligand and an analyte sample solution over time, a data generation unit configured to generate response data representative of the time evolution of the observed molecular interaction, an extraction unit configured to extract at least one interaction parameter from at least a portion of the generated response data, and a data processing unit configured to receive the at least one interaction parameter and to classify the analyte sample solution into at least one quality classification group indicative of an interaction between the analyte sample solution and the ligand using a trained machine learning algorithm based on the at least one interaction parameter.

[0041] The analytical sensor system may further comprise a training center for training a machine learning algorithm by providing the machine learning algorithm with interaction parameters extracted from response data from a plurality of observed analyte-ligand interactions, together with at least one quality classification group indicative of the analyte and ligand interactions. [Brief explanation of the drawings]

[0042] [Figure 1] 1 shows a schematic diagram of an analytical system for detecting molecular binding interactions and classifying observations. [Figure 2] 1 shows the detection curve versus time from the interaction between the analyte and the ligand. [Figure 3] Response levels of binding interactions between different analytes and ligands at the same concentration are shown. [Figure 4] Figures 4a and 4b show plots of the concentration series and the concentration at the reporting point BL. [Figure 5] An ideal detection curve is shown, characterized by a late binding response, an early binding response, a stable early period, and a stable late period. [Figure 6]1 shows an example of a detection curve showing the binding of an analyte to a ligand with poor binding quality. [Figure 7] 1 shows an example of a detection curve showing the binding of an analyte to a ligand with poor binding quality. [Figure 8] Figures 8a and 8b show a series of analyte solutions and their dose-response curves at concentrations both below and above the KD. The series are grouped for fitting the data using a free fit of the Rmax parameter. [Figure 9] Figures 9a and 9b show a series of analyte solutions and their dose-response curves, all concentrations below the KD, where the affinity is determined using a free fit of the Rmax parameter. The series is divided into groups for fitting the data at a constant Rmax. [Figure 10] 1 is a flowchart illustrating steps of a method for classifying observations from an analytical sensor system configured to observe molecular interactions between an analyte and a ligand. [Figure 11a] 1 shows graphical elements that support learning and classification. [Figure 11b] 1 shows graphical elements that support learning and classification. [Figure 11c] 1 shows graphical elements that support learning and classification. [Figure 11d] 1 shows graphical elements that support learning and classification. [Figure 11e] 1 shows graphical elements that support learning and classification. DETAILED DESCRIPTION OF THE INVENTION

[0043] This disclosure relates to analytical sensor methods, particularly those based on biosensors, for investigating molecular interactions. Results are presented in the form of response data, often presented as a response curve, or sensorgram, in real time as the interaction progresses.

[0044] Biosensors can be based on a wide variety of detection methods. Typically, such methods include, but are not limited to, mass-sensing methods (e.g., piezoelectric, optical, thermo-optical, surface acoustic wave (SAW) device methods), and electrochemical methods (e.g., potentiometric, conductometric, amperometric, and capacitance methods). Representative optical detection methods include mass-surface concentration-sensing methods (e.g., reflective optical methods such as internal and external reflection methods), angle-sensing methods, wavelength-sensing methods, and resolved phase-sensing methods, such as ellipsometry and evanescent wave spectroscopy (EWS), including surface plasmon resonance (SPR) spectroscopy, Brewster angle refractometry, critical angle refractometry, frustrated total reflection (FTR), evanescent wave ellipsometry, scattered total internal reflection (STIR), light guide sensors, and evanescent wave imaging, e.g., critical angle resolved imaging, Brewster angle resolved imaging, and SPR angle resolved imaging. Furthermore, photometric methods, such as those based on evanescent fluorescence (TIRF, total internal reflection fluorescence), phosphorescence, and waveguide interferometry, can also be used.

[0045] One commonly used detection principle is surface plasmon resonance (SPR) spectroscopy. An example of an SPR-based biosensor is sold under the name BIACORE® (hereinafter referred to as a "BIACORE instrument"). Such biosensors use SPR-based mass detection to provide substantially real-time, label-free analysis of binding interactions between surface-bound ligands and target analytes.

[0046] The BIACORE instrument includes a sensor chip containing a light-emitting diode (LED), a glass plate covered with a thin gold film, an integrated fluidic cartridge that provides a liquid flow over the center chip, and a photodetector array. Incident light from the LED is totally internally reflected at the glass-gold interface and detected by the photodetector array. At a specific angle of incidence (the "SPR angle"), surface plasmon waves are generated in the gold layer and detected as a loss in intensity, or "dip," in the reflected light. The SPR phenomenon for the BIACORE instrument relies on the resonant coupling of monochromatic p-polarized light incident on the thin metal film through a prism and glass plate to the vibrations of conduction electrons, called plasmons, in the thin metal film on the opposite side of the glass plate. These vibrations generate an evanescent field extending from the surface into the liquid flow at distances on the order of one wavelength (approximately 1 μm). When resonance occurs, the collective excitation of electrons in the metal film causes light energy to be lost to the metal film, and the reflected light intensity decreases at a precisely defined angle of incidence (SPR angle), which depends on the refractive index within the reach of the evanescent field near the metal surface.

[0047] As mentioned above, the SPR angle depends on the refractive index of the medium near the gold layer. In BIACORE instruments, dextran is typically bound to the gold surface, and a ligand for analyte binding is bound to the surface of the dextran layer. The analyte of interest is injected in solution onto the sensor surface through a fluidic cartridge. Because the refractive index near the gold film depends on (i) the refractive index of the solution (a constant) and (ii) the amount of material bound to the surface, the binding interaction between the bound ligand and the analyte can be observed as a function of the change in the SPR angle.

[0048] Figure 1 shows a schematic diagram of an analytical sensor system 20 for detecting molecular binding interactions and classifying observations. In this example, the system comprises a sensor device (chip 1) having a sensing surface (gold film 2) supporting capture molecules 3 (e.g., antibodies) exposed to a sample stream having an analyte 4 (e.g., antigen) through a flow path 5. Monochromatic p-polarized light 6 from a light source 7 (LED) is coupled by a prism 8 to a glass-metal interface 9 where the light is totally internally reflected. The intensity of the reflected light beam 10 is detected by a detector 11 (e.g., a photodetector array).

[0049] A typical output from a BIACORE instrument is a "sensorgram," which is a plot of response data (measured in "resonance units" (RU)) as a function of time. An increase of 1000 RU corresponds to approximately 1 ng / mm 2 This corresponds to an increase in mass on the sensor surface. When a sample containing an analyte contacts the sensor surface, the ligand bound to the sensor surface interacts with the analyte in a step called "association." This step is shown in the sensorgram as an increase in RU, since the sample first contacts the sensor surface. Conversely, "dissociation" usually occurs when the sample flow is replaced by, for example, a buffer flow. This step is shown in the sensorgram as a decrease in RU over time, since the analyte dissociates from the surface-bound ligand.

[0050] A representative sensorgram of a BIACORE instrument is shown in Figure 2, depicting a sensing surface with an immobilized ligand (e.g., an antibody) interacting with an analyte in a sample. The y-axis represents the response (here in resonance units (RU)), and the x-axis represents time (here in seconds). Initially, buffer is passed over the sensing surface, providing the "baseline response" of the sensorgram. The increase in signal observed during sample injection is due to analyte binding (i.e., association), leading to a plateau where the response signal levels off. At the end of sample injection, the sample is replaced with a continuous flow of buffer, and the decrease in signal reflects the dissociation or release of the analyte from the surface. The slope of the association / dissociation curve provides useful information about the interaction kinetics, and the height of the response signal represents the surface concentration (i.e., the response due to the interaction is related to the mass concentration change on the surface). The analytical system 20 shown in Figure 1 includes a data generation unit 12 for generating response data representing the time evolution of the interaction.

[0051] Detection curves and sensorgrams generated by biosensor systems based on other detection principles have a similar appearance.

[0052] For example, in the early stages of fragment drug discovery, it is rare to be able to identify candidate analytes from a single binding measurement alone. A binding level screen (BLS) provides an overview of the library of candidate analytes and allows the identification of analytes for further analysis. In a typical BLS, a single concentration of each analyte is run over the target and reference. Through analysis of the response data, lead candidate analytes can be selected for binding strength and binding behavior. In the next step, an affinity screen, the lead candidate analytes are tested at multiple concentrations to ensure that binding is concentration-dependent, allowing the affinity of the analyte / ligand interaction to be determined, or at least estimated.

[0053] However, classifying analyte samples as hits in a binding level screen (BLS) and using them for further processing can be tedious and difficult even for experienced users, especially when hundreds of samples are obtained in a single evaluation. Figure 3 shows the response levels of binding interactions of multiple different analytes at the same concentration to a ligand from a typical BLS analysis. Because different users may classify BLS samples differently (i.e., hits or not), the evaluation of a library of candidate analytes can be diverse and user-dependent. In affinity screen analyses, users may need to exclude some concentrations from the analysis and must decide which model to use to determine KD.

[0054] In affinity screen analysis, a set of acquired response data (sensorgrams) obtained from a series of analyte concentrations (i.e., using the same analyte) is analyzed. Machine learning algorithms (ML) can be trained to remove outliers from the series of sensorgrams so that further analysis of the sensorgrams is not based on data containing such outlier sensorgrams. If too many sensorgrams in a series are determined to be of poor quality, it may be desirable to reject the entire series. ML determines sensorgram inclusion and exclusion and can also be trained by an expert user to include information about why particular sensorgrams were rejected.

[0055] In affinity screens, classification of sensorgrams can be performed using multiple sensorgrams, and for each set of sensorgrams, a set of features is calculated and a mathematical model is fitted to the set of sensorgrams.

[0056] The set of features thus computed may include three or more of the following: - Maximum binding capacity Rmax divided by the standard error of Rmax; - late binding response B divided by Rmax; - The mean squared error MSE between the detection curve and the fitting mathematical model is the late binding response B 2 divided by; - Classification of each sensorgram into quality classification groups.

[0057] The maximum binding capacity Rmax used can be a constant or a fitted Rmax. The equilibrium dissociation constant KD is fitted.

[0058] A set of sensorgrams can include at least one sensorgram, and if it includes two or more sensorgrams (e.g., 2 to 10 sensorgrams), each sensorgram in the set represents a molecular interaction (between the same ligand and analyte) at a different molecular concentration, i.e., a different analyte concentration.

[0059] The mathematical model that is fitted to the set of sensorgrams is a model that describes the molecular interactions at equilibrium, and can be, for example, a 1:1 steady-state affinity model or a heterologous ligand steady-state affinity model.

[0060] The parameter values ​​of KD, Rmax, and offset are divided by their respective standard errors. The standard error of a parameter is a measure of how close the parameter is to the fitted mathematical model. Rmax can be the Rmax calculated from the fitting. The late binding response B is the maximum binding response in the detection curve relative to the baseline value of the detection curve. The late binding response B feature can be normalized by dividing its value by Rmax, and the mean squared error MSE feature is calculated by squaring its value. 2 These features are normalized to be applicable to any kind of input data. When a steady-state model with a constant Rmax is applied, there is no standard error that can be calculated for the parameters.

[0061] A machine learning algorithm can be trained to classify each sensorgram into a quality classification group by fitting a mathematical model to the set of sensorgrams and calculating a set of features from the set of sensorgrams and the fitted mathematical model, with parameters determined by an expert user.

[0062] Sensorgrams are classified into quality classification groups that indicate the quality of the sensorgram. Poor-quality sensorgrams can adversely affect the evaluation and should be excluded from further analysis of the observed interaction. Poor-quality curves can have, for example, unstable baselines, air spikes, or sub-baseline responses, and can be caused by, for example, a contaminated flow system or running buffer, the ligand capture method used, or too low a detergent concentration in the running buffer. What constitutes a sufficiently good quality detection curve and whether to include it in the affinity analysis of the observed interaction depends on the purpose of the analysis / experiment.

[0063] For example, the number of different quality classification groups is two, one group has sensorgrams of sufficiently good quality, and the second group has sensorgrams of poor quality, and the group with sensorgrams of sufficiently good quality can be selected for affinity analysis of observed molecular interactions.

[0064] In another example, the number of quality classification groups is 100, with group 1 having sensorgrams of poor quality and group 100 having sensorgrams of very good quality. The user may regulate the stringency of the method, for example, by setting a cutoff of >50, so that only sensorgrams classified into classification groups 51 through 100 are included in the kinetic analysis of the observed interactions. Again, what constitutes a sensorgram of sufficiently good quality and where the cutoff is set will depend on the objectives of the analysis / experiment.

[0065] The set of computed features can be three or more. Using more features in the method can improve the classification of sensorgrams if the features complement each other, highlighting what an experienced user looks for to distinguish good enough quality from not so good quality. Too many features that do not contribute to the classification can lead to less good learning and classification. Also, additional features not mentioned above can be computed and included in the sensorgram classification step.

[0066] The mathematical model can be selected from a 1:1 steady-state affinity model or a heterologous ligand steady-state affinity model.

[0067] Based on the classification, it can be decided which sensorgrams to use for affinity analysis of the observed molecular interactions.

[0068] When using the fixed / constant Rmax model, the fixation level is set to the predicted Rmax, and the evaluation software automatically refits the series after the ML prediction. The model predicts a fixed Rmax and can be trained by an expert user to provide annotations for why it is a fixed Rmax. The fixed Rmax currently uses only two features from the sensorgram, but additional features such as the slope of the association step and the initial binding response can be added to detect superstoichiometric cases.

[0069] An ML model predicts whether a series of sensorgrams pass or fail, and the reasons for failure. To train a machine learning algorithm, an expert provides annotations for the failing series with predefined subclasses, while leaving the passing series unlabeled. An ML model can consist of two models, one that handles pass / fail and the other that handles the reasons for failure.

[0070] Figures 4a and 4b show the concentration series and the response versus concentration at the reporting point BL. As with BLS, this can be challenging even for experienced users, and the results obtained from affinity screens can be user-dependent. By performing either or both BLS and affinity screen analyses, candidate analytes can be identified for further processing, for example, in drug discovery experiments.

[0071] Described below are methods and systems for classifying molecular interactions between analytes 4 and ligands 3, such as to identify candidate analytes for further processing in drug discovery experiments. The methods and apparatus are suitable for classifying large numbers of such molecular interactions.

[0072] The method is illustrated in FIG. 10, where an exemplary analytical sensor system 20 for detecting molecular binding interactions between an analyte 4 and a ligand 3 and classifying the observations is shown in FIG. 1. A first set of analyte sample solutions is allowed to interact (100) with a ligand 3 of interest. The first set of analyte sample solutions may comprise at least one analyte sample solution. The analytes may be, for example, a plurality of different candidate analytes in a drug discovery assay. Typically, when performing a binding level screen (BLS) or affinity screen, hundreds of analyte sample solutions may be analyzed to provide an overview of a library of candidate analytes and allow identification of candidate analytes that are likely to interact with a particular ligand and be relevant for further processing.

[0073] A sensor device 1 equipped with a detector 11 is arranged to monitor (101) over time the intermolecular interaction between a ligand 4 and a sample solution of analyte 3. The monitored intermolecular interaction can occur at the sensing surface 2 of a mass-detection biosensor system, e.g., an optical biosensor, where one molecule (ligand 3) is attached / immobilized to the sensor surface and another molecule (analyte 4) passing over the sensor surface, allowing the time evolution of the analyte-ligand binding to be monitored. Analytical sensor systems based on other detection principles, such as electrochemical systems, are also possible.

[0074] Multiple sample solutions may have approximately the same concentration of analyte. In practice, the exact concentration of an analyte is rarely known. In practice, using a nominal fixed concentration value allows for concentration variation, but this makes interpretation of results more uncertain.

[0075] The data generation unit 12 is configured to generate corresponding response data representative of the time evolution of the observed intermolecular interactions. By using the extraction unit 13, at least one interaction parameter can be extracted (102) from at least a portion of each set of acquired response data. Each response data set ideally represents the intermolecular interaction between the ligand 3 and the analyte 4. However, the response data set may also comprise "noise" due to, for example, undesired and uncertain binding to the sensing surface 2, differences in the refractive index between the analyte sample solution and the running buffer used, etc.

[0076] During injection of the analyte sample, an increase in signal due to the binding (i.e., association) of analyte 4 is observed, followed by a plateau in the response signal, which becomes a steady state. When the analyte dissociates from the ligand, the dissociation state is reached.

[0077] The data processing unit 14 is configured, when provided with (103) at least one interaction parameter, to classify (104) the analyte sample solution into at least one quality classification group indicative of the interaction of the analyte sample solution with the ligand based on the at least one interaction parameter using a trained machine learning algorithm.

[0078] The response data can be provided as a data set, or, in the case of using a surface plasmon resonance biosensor such as BIACORE®, can be provided graphically as a detection / response curve (sensorgram).

[0079] At least one interaction parameter may comprise 80-100% of the acquired response data for each detection curve / sensorgram used.

[0080] In some embodiments, it is preferable to use the entire response data set, rather than only a portion thereof, as the extracted interaction parameters. The interaction parameters can be based, for example, on 90-100% of the total response data. The portion of the response data set that contains the most useful information is the response data observed before analyte binds to the ligand, during the interaction, and during the period thereafter (during dissociation).

[0081] The at least one extracted parameter may comprise any one or more of the following interaction parameters extracted at a given time point (e.g., a given time period) of the response data: a) initial binding response, b) late binding response; c) stable initial response, d) stable late response; e) the initial binding response divided by the molecular weight of the analyte; f) the late binding response divided by the molecular weight of the analyte; g) the difference between the late binding response and the early binding response divided by the early binding response; h) Measurement fluctuation of any of the interaction parameters a) to g). i) the measured gradient of any of the interaction parameters a) to g), and / or When the analyte sample solution is composed of multiple analyte sample solutions containing the same analyte but with different analyte concentrations, the interaction parameters include the following: j) Equilibrium dissociation constant calculated using maximum binding capacities from multiple analyte sample solutions.

[0082] Such a predetermined period may include, for example, continuous response data observed for 1 second, 30 seconds, or 1 minute or more, and is preferably between 1 second and 1 minute.

[0083] The initial binding response of the analyte is the binding level of the molecular interaction detected at the beginning of the injection of the analyte sample solution, which is the time when the analyte sample solution begins to interact with the ligand.

[0084] The late binding response of the analyte is the binding level of the intermolecular interaction detected at the end of the injection of the analyte sample solution.

[0085] The stable initial response is the level of intermolecular interaction detected at the beginning of the dissociation phase following injection of the analyte sample solution.

[0086] The stable late response is the level of intermolecular interaction detected at the end of the dissociation phase following injection of the analyte sample solution, respectively.

[0087] An ideal detection curve between an analyte sample solution and a ligand is shown in Figure 5. The detection curve shows an initial binding response, a late binding response, a stable initial phase, and a stable late phase.

[0088] One feature calculated can be the initial binding response divided by the molecular weight of the analyte. By dividing by molecular weight, the response level is normalized, since small analytes do not have as large a response as large analytes. It is desirable that the relative response be as independent of the size of the analyte 4 as possible.

[0089] The difference between the late analyte binding response and the early analyte binding response divided by the early analyte binding response describes the slope of the entire association phase.

[0090] Here, the measured variation of any of the interaction parameters a) to g) refers to the standard deviation taken over the period during which a particular parameter, for example, a parameter called the initial binding response, is measured.

[0091] Here, the measured gradient of any of the interaction parameters a) to g) refers to the gradient taken over the period during which a particular parameter, for example, a parameter called the initial binding response, is measured.

[0092] The number of interaction parameters provided to the trained machine learning algorithm for each analyte sample solution in the first set of analyte sample solutions can be one or more of the parameters described above. For example, the initial binding response, the difference between the late binding response and the initial binding response divided by the initial binding response or the late binding response, or the late binding response can be used. In an affinity screen, a plot of the initial binding, late binding, stable initial, or stable late response versus concentration provides data that can be fitted with a steady-state model to obtain a KD value.

[0093] The user of the analytical sensor system can determine which and how many interaction parameters are extracted from the response data and provided to the trained machine learning algorithm. Instead, the method predetermines which and how many interaction parameters are extracted and provided to the trained machine learning algorithm. Generally, the more interaction parameters, the more reliable the classification. The determination of which interaction parameters and the extraction of the parameters can be performed by an expert system or ANN trained on multiple training response datasets. Depending on the purpose of the analysis and the ligand being analyzed, the number and type of interaction parameters extracted and provided to the trained machine learning algorithm can vary. Using a larger number of interaction parameters in the method, where the interaction parameters complement each other, can improve the classification performed by the trained machine learning algorithm and highlight what an experienced user focuses on to distinguish related analytes from unrelated analytes. Too many parameters that do not contribute to the classification can result in a less than satisfactory classification. Also, additional parameters not mentioned above may be calculated and included in the step of classifying the analyte sample solution.

[0094] Alternatively, one or more of these interaction parameters may be provided to the trained machine learning algorithm along with the interaction parameters that comprise 80-100% of the acquired response data for the analyte sample solutions in the first set of analyte sample solutions described above.

[0095] The analyte sample solution is classified by the trained machine learning algorithm into at least one quality classification group indicative of the interaction of the analyte sample solution with the ligand based on at least one extracted interaction parameter.

[0096] Analyte sample solutions can be classified with respect to their binding strength or binding behavior.

[0097] In one example, the number of different quality classification groups can be two, one group containing analytes that interact well enough with the ligand of interest and a second group containing analytes that interact poorly with the ligand of interest, and the group containing analytes that interact well enough with the ligand can be selected for further analysis.

[0098] In another example, the number of quality classification groups could be 100, with group 1 containing analytes with poor interactions with the ligand and group 100 containing analytes with very good interactions with the ligand. The user may regulate the stringency of the method, for example by setting a cutoff at >50, so that only analytes classified into classification groups 51 through 100 are included in further analysis. Again, what constitutes a sufficiently good quality interaction and where the cutoff is set depends on the objectives of the analysis / experiment.

[0099] Classifying an analyte sample solution into at least one quality classification group indicative of the interaction between the analyte sample solution and the ligand can include a first classification into a main group (e.g., whether it is relevant for further analysis) and a second group providing information about why the analyte sample solution was classified into that main group, such as poor binding, a refractive index that is too low or too high, non-1:1 binding to the ligand, no or very low binding response (below a predetermined cut line), high gradient in association, slow dissociation, poor mixing, R>Rmax, or the analyte being a superstoichiometric binder. Figure 5 shows ideal detection between a sample solution of analyte 4 and ligand 3. The detection curve shows an initial binding response, a late binding response, a stable initial phase, and a stable late phase. Figures 6 and 7 show curves indicating poor quality analyte-ligand binding. The analyte sample solution in Figure 6 shows a response that extends over a long period after injection and is classified in the main group as being irrelevant for further analysis, and in the second group as indicating secondary, undesired binding of the analyte to the ligand and / or aggregation of the analyte. The analyte sample solution in Figure 7 shows increased binding of the analyte to the ligand during injection, and is classified in the main group as being irrelevant for further analysis, and in the second group as indicating secondary, undesired binding of the analyte to the ligand and / or aggregation of the analyte. For the series of analyte solutions shown in Figures 8a and 8b, fitting of the plotted data reveals that concentrations exist both below and above the KD value, with at least one concentration approaching the saturation response. This indicates that the KD can be obtained by fitting the Rmax parameter to a steady-state model. This series is classified into a group for fitting the data using a free fit of the Rmax parameter. For the dose-response curves shown in Figures 9a and 9b, all concentrations are below the KD value, and initial fitting work yields unreliable KD and Rmax values.The series are divided into groups for fitting the data with a constant Rmax parameter.

[0100] Other interaction parameters than those described above may be extracted and used in the method, either alone or in combination with the interaction parameters described above. For example, the method may further comprise extracting at least one reference-subtracted interaction parameter for each analyte sample solution in the first set of analyte sample solutions and providing it to the trained machine learning algorithm.

[0101] Reference-subtracted interaction parameters can be obtained by interacting a first set of analyte sample solutions with a ligand-free reference (e.g., a reference surface with no immobilized ligand) and subtracting the extracted reference interaction parameters from the corresponding extracted interaction parameters.

[0102] Reference interaction parameters may be used in the method, with at least one reference interaction parameter comprising any one or more of the following: k) the reference initial combined response; l) late binding response of the reference; m) Reference stable initial response, n) Reference stable late response, o) the difference between the reference late binding response and the reference early binding response divided by the reference early binding response; and / or If the analyte sample solution is composed of multiple analyte sample solutions of the same analyte but with different analyte concentrations, the reference interaction parameters comprise: p) Equilibrium dissociation constant calculated using maximum binding capacities from multiple analyte sample solutions.

[0103] The reference interaction parameters may be obtained by interacting a set of first analyte sample solutions with a ligand-free reference (e.g., a reference surface with no immobilized ligand) and extracting the reference interaction parameters from the resulting set of reference response data. Adding such interaction parameters to the method may improve the probability of detecting imprecise binding or poor mixing.

[0104] Negative control interaction parameters obtained from the interaction of a ligand with a negative control solution may be used in the method.

[0105] At least one negative control solution may contain an analyte with known non-affinity for the ligand or may contain no analyte, e.g., a buffer solution. The negative control interaction parameter may be, for example, a boundary level based on the negative control. This is the average response level from the negative control injection multiplied by a predetermined number of standard deviations (typically three), defining a binding level limit below which the presence of an intermolecular interaction cannot be confirmed. The interaction parameter with the negative control is intended to detect poor buffer matching between the sample and the system running buffer, sample and / or system contamination, changes in the ligand, and experimental signals unrelated to intermolecular interactions. The negative control also determines a boundary level, which is used to distinguish potential hits of interest from nonbinders (below the boundary).

[0106] Positive control interaction parameters obtained from the interaction of a ligand with a positive control solution may be used in the method.

[0107] At least one positive control solution contains an analyte sample with known affinity and binding behavior for the ligand. The positive control may have the same molecular concentration as the analyte sample solution, but this is rare in practice. The positive control interaction parameter may be, for example, a boundary level based on the positive control, which is the average binding level for each molecular interaction detected between the ligand and the positive control, defining a known binding level to which an unknown sample may relate, or establishing the stoichiometry of an unknown sample if the stoichiometry of the positive control has been established. The positive control interaction parameter is intended to set a stoichiometric response level by observing the binding capacity and reproducibility of the response of a fixed target.

[0108] In addition to using the interaction parameters described above, when using a surface plasmon resonance analytical system such as the BIACORE®, it may be necessary to provide a solvent-corrected interaction parameter to the trained machine learning algorithm during the analysis of an analyte sample. The need for solvent correction can arise when the amount of ligand is high (compared to a ligand-free reference) and the bulk refractive index contribution of the solvent is high compared to the predicted analyte response. Because the bulk solution is excluded from the volume occupied by the ligand on the active surface, the bulk contributions to the active and reference surfaces will be slightly different, resulting in a small error in the reference-subtracted response. As long as the refractive index of the sample remains constant, the error in the reference-subtracted response will also remain constant and, for practical purposes, negligible. However, if the refractive index of the sample changes, the magnitude of the error will also change. Adding 1% DMSO (dimethyl sulfoxide) to a buffer gives a bulk response of approximately 1200 RU, so small variations in DMSO content can lead to significant variations in the bulk response relative to the predicted response from a low-molecular-weight sample (which can be as low as 1 RU). Such variations are difficult to avoid during sample preparation of drug candidates, fragments, etc. for screening.

[0109] Solvent correction can be determined by injecting a series of blank samples with a range of solvent concentrations onto the active and reference surfaces. A plot of the relative reference-subtracted response on the active surface against the absolute response on the reference surface calibrates the error in reference subtraction for bulk contributions. This calibration is then used to correct the measured sample response. It is recommended to perform solvent correction cycles at the beginning and end of a run and at regular intervals during the run. In this way, a given sample cycle falls between two solvent correction cycles. The correction factor for a given sample cycle is determined by interpolation between the curves of the previous and next correction cycles, compensating for drift in the solvent correction factor over the course of the run.

[0110] The analytical sensor system 20 may further comprise a training center 15 for training a machine learning algorithm, which is trained (200) using a set of interaction parameters extracted from at least a portion of the acquired response datasets obtained from interactions between at least one ligand and a second set of analyte sample solutions and at least one quality classification group indicative of the interaction between the analyte sample solution and the ligand for each analyte sample solution in the second set of analyte sample solutions.

[0111] The trained machine learning algorithm mimics the way an expert user classifies analyte sample solutions into quality classification groups. The training can be a standard machine learning algorithm training method, where the entire dataset is divided into a training set and a validation set, with 80% of the data used for training and 20% of the data used for validation. The machine learning algorithm should be trained on the training data, not the validation set.

[0112] Graphical elements such as those shown in Figures 11a-11c can be used to further train the machine learning algorithm to support parameter assignments including comparison against theoretical binding curves, simultaneous display of binding versus a reference, dose-response plot lines corresponding to KD values ​​based on using different steady-state models, etc. The generated model allows for easy modification of the classification through the graphical elements as shown in Figure 11e.

[0113] In Figures 11a and 11b, ghost curves are shown in the sensorgrams. The ghost curves are typical fragment profiles simulated with a 1:1 binding model including mass transfer. Rmax has been scaled to allow direct comparison of the actual sensorgrams and the simulated curves.

[0114] The method may also include filtering and tabbing of the data to visualize sets of classifications as shown in FIG. 11d.

[0115] The computer program may use graphical elements (Figures 11a-c) to support comparison to theoretical binding curves, simultaneous display of binding and binding relative to a reference, and / or parameter assignment such as lines in dose-response plots corresponding to KD values ​​based on the use of different steady-state models.

[0116] The machine learning algorithm may be selected from the group consisting of decision trees, k-nearest neighbors (kNN), random forests, k-means, gradient boosting algorithms, artificial neural networks (ANNs), deep learning algorithms, or combinations thereof.

[0117] A machine learning algorithm may be trained using data from one analytical sensor system and used to classify data from another analytical sensor system, or alternatively, the data used for training and the data classified by the machine learning algorithm may be obtained by the same system, or at least by the same type of system.

[0118] A second set of analyte sample solutions is used when training the machine learning algorithm. The second set of analyte sample solutions can differ from the first set of analyte sample solutions in both the number of analyte sample solutions and the type of analyte. The ligands used when training the machine learning algorithm can be the same as those used to interact with the first set of analyte sample solutions. Alternatively, the ligands used can be ligands that are not used to interact with the first set of analyte sample solutions.

[0119] The second analyte sample solution set can include a plurality of different analyte sample solutions. Preferably, the second analyte sample solution set can include 10 or more, 100 or more, or 500 or more different analyte sample solutions. The greater the number of analyte sample solutions used for training, the better the classification of the analyte sample solutions in the first analyte sample solution set.

[0120] The interaction parameters used to train the machine learning algorithm and how many interaction parameters are used per analyte sample solution may depend on the intended use of the machine learning algorithm. Generally, the more interaction parameters used for training, the better the classification of the analyte sample solutions in the first set of analyte sample solutions. The determination of the interaction parameters and their extraction may be performed by an expert user of the analytical system, or may be performed by an expert system or ANN that has been trained on multiple training response data sets.

[0121] The types of interaction parameters used to train the machine learning algorithm may be the same as those provided to the trained machine learning algorithm for the analyte sample solutions in the first set of analyte sample solutions. Alternatively, the number and / or types of interaction parameters used may be different. The machine learning algorithm may be trained with a larger set of interaction parameters than will be used to actually classify the analyte sample solutions in the first set of analyte sample solutions.

[0122] The types of parameters extracted from the first analyte sample set can be extracted from a second analyte sample set for training the machine learning algorithm. The machine learning algorithm can also be provided with a reference subtraction interaction parameter, a reference interaction parameter, a negative interaction parameter, a positive interaction parameter, and a solvent compensation, as described above. The machine learning algorithm can also be provided with at least one quality classification group for each analyte sample solution in the second analyte sample set, which indicates the interaction between the analyte sample solution and the ligand. The quality classification group for each analyte sample solution in the second analyte sample set can be determined by an expert user of the analytical system, or by an expert system or ANN that has been trained on multiple training response datasets.

[0123] In one example, the number of different quality classification groups is two, with the first group containing analytes that interact well enough with the target ligand (see Figures 4a and 4b) and the second group containing analytes that interact poorly with the target ligand (see Figures 5 and 6). A classification of good enough indicates that the analyte may be a candidate for further analysis.

[0124] In another example, the number of quality classification groups can be 100, with group 1 containing analytes that interact poorly with the ligand and group 100 containing analytes that interact very well with the ligand.

[0125] Classifying an analyte sample solution or series of solutions into at least one quality classification group indicative of the interaction of the analyte sample solution with the ligand may include a first classification into a main group (e.g., relevant for further analysis or not), and a second group providing information as to why the analyte sample solution was classified into that main group, such as uncertain binding, refractive index that is too low or too high, non-1:1 binding to the ligand, no or very low binding response (below a pre-determined cut-line), high gradient in association, slow dissociation, poor mixing, R>Rmax, analyte being a superstoichiometric binder, concentration above KD, concentration too low to determine KD, unreliable parameters in steady-state analysis, etc.

[0126] The same machine learning algorithm may classify into multiple groups: hit or not, why not a hit, etc. An alternative is to combine two trained machine learning algorithms, one to classify the analyte and the other to explain why the analyte was classified as not relevant for further analysis. Statistical measures of each interaction parameter may be used when training the machine learning algorithms.

[0127] The machine learning algorithm thus trained then classifies analyte sample solutions from the first set of analyte sample solutions into quality classification groups indicative of the interaction between the analyte sample solution and the ligand. In BLS or affinity screen assays, where hundreds of analyte sample solutions are analyzed to provide an overview of a library of analyte candidates, this method provides rapid classification of relevant or incompatible analyte candidates for further processing. Furthermore, the evaluation of candidate analytes may become less user-dependent, and an experienced user may not be able to assist in the evaluation of observed molecular interactions.

[0128] As a non-limiting example, a machine learning algorithm was used (using the SciSharp.Tensorflow.Redist v1.15.1 and Tensorflow.NET v0.13.0 NuGet packages) that included a single hidden layer neural network with 10 hidden neurons and a learning rate of 0.001. In this example, interactions between multiple different analyte sample solutions and the ligand, the hepatitis enzyme NS5B1b (RNA-dependent RNA polymerase, subtype 1b), were observed. A dataset containing data on 457 analyte sample solutions, 64 positive controls, and 136 negative controls was divided into a training set (second analyte sample solution set) and a validation set (first analyte sample solution set), with 80% of the data used for training and 20% used for validation.

[0129] The interaction parameters extracted and used in this example for both training and validation were: initial binding response divided by late binding response of positive control; difference between late binding response and early binding response divided by initial binding response; stable early response; stable late response; difference between late binding response and early binding response; reference late binding response; reference stable late response; difference between reference late binding response and reference early binding response; initial binding divided by mean binding of all negative controls at five standard deviations of all negative controls; and measured variance (standard deviation) of initial binding response.

[0130] Furthermore, for each analyte sample solution used for training, a quality classification group indicating the interaction between the unlit sample solution and the ligand was provided to the machine learning algorithm, where the sample solution was classified into two quality classification groups: the first group contained analytes that interacted sufficiently well with the ligand, and the second group contained analytes that interacted poorly with the ligand.

[0131] When the machine learning algorithm trained in this way was used on a validation dataset, it achieved a specificity of 96% and a sensitivity of 93%.

[0132] By continually training and developing these models, an experienced user can easily modify the classification results and retrain the models to better suit the user's preferences. Furthermore, by using the trained machine learning algorithms described above to classify detected analyte samples, the classification process may be sped up, especially when large amounts of analyte samples are present in the analysis, and the evaluation results may be less user-dependent and less likely to aid an experienced user in evaluating the observed molecular interactions.

[0133] Although the above description includes multiple features, they do not limit the scope of the disclosed concepts, but merely illustrate some exemplary embodiments of the disclosed concepts. The scope of the disclosed concepts fully encompasses other embodiments that may become apparent to those skilled in the art, and the scope of the disclosed concepts is not limited. Reference to an element in the singular does not mean one or only one, but one or more, unless otherwise specified. All structural and functional equivalents known to those skilled in the art of the elements of the above-described embodiments are expressly incorporated and encompassed herein.

Claims

1. 1. A method for analyzing observations from an analytical sensor system (20) configured to observe a molecular interaction between an analyte (4) and a ligand (3) such that a response dataset representative of the time evolution of said interaction is obtained, comprising: interacting (100) a set of first analyte sample solutions with a ligand (3); and obtaining (101) a corresponding response data set for each observed molecular interaction; extracting (102) at least one interaction parameter from at least a portion of each acquired response dataset; training (200) the machine learning algorithm, in turn, comprising: providing (103) for each analyte sample solution the at least one extracted interaction parameter to a trained machine learning algorithm; and classifying (104) each analyte sample solution into at least one quality classification group indicative of the interaction between the ligand (3) and the analyte sample solution based on the at least one extracted interaction parameter; a set of interaction parameters extracted from at least a portion of the acquired response data set obtained from the interaction between the second set of analyte sample solutions and at least one ligand (3); and at least one quality classification group indicative of an interaction between the analyte sample solution and the ligand (3) for each analyte sample solution of the second set of analyte sample solutions.

2. 2. The method of claim 1, wherein the at least one interaction parameter extracted from the response data sets acquired for the analyte sample solutions in the first set of analyte sample solutions comprises 80-100% of the acquired response data.

3. 3. The method of claim 1, wherein at least one interaction parameter extracted from the response datasets acquired for the analyte sample solutions of the second set of analyte sample solutions and used to train the machine learning algorithm comprises 80-100% of the acquired response data.

4. The at least one interaction parameter extracted from the response data sets obtained for the analyte sample solutions in the first set of analyte sample solutions comprises one or more of the following interaction parameters extracted at a given time point in the response data: a) initial binding response; b) late binding response; c) stable initial response; d) stable late response; e) the initial binding response divided by the molecular weight of the analyte; f) the late binding response divided by the molecular weight of the analyte; g) the difference between the late binding response and the early binding response divided by the early binding response; h) the measured variation of any of the interaction parameters a) to g); i) a measured gradient of any of the interaction parameters a) to g), and / or j) an equilibrium dissociation constant calculated using the maximum binding capacities from a plurality of analyte sample solutions each containing the same analyte but at different analyte concentrations, when the analyte sample solution is composed of the plurality of analyte sample solutions.

5. At least one interaction parameter extracted from a second response dataset obtained for the analyte sample solutions in the second set of analyte sample solutions and used to train the machine learning algorithm comprises one or more of the following interaction parameters extracted at a predetermined time point in the response data: a) initial binding response; b) late binding response; c) stable initial response; d) stable late response; e) the initial binding response divided by the molecular weight of the analyte; f) the late binding response divided by the molecular weight of the analyte; g) the difference between the late binding response and the early binding response divided by the early binding response; h) the measured variation of any of the interaction parameters a) to g); i) a measured gradient of any of the interaction parameters a) to g), and / or j) an equilibrium dissociation constant calculated using the maximum binding capacities from a plurality of analyte sample solutions each containing the same analyte but at different analyte concentrations, when the analyte sample solution is composed of the plurality of analyte sample solutions.

6. 6. The method of claim 1, further comprising extracting at least one reference-subtracted interaction parameter for each analyte sample solution of the first set of analyte sample solutions and providing the extracted interaction parameter to the trained machine learning algorithm.

7. 7. The method of claim 1, wherein the machine learning algorithm is further trained using at least one reference-subtracted interaction parameter for each analyte sample solution of the second set of analyte sample solutions.

8. and extracting at least one reference interaction parameter for each analyte sample solution of the first set of analyte sample solutions and providing the extracted at least one reference interaction parameter to the trained machine learning algorithm, the at least one reference interaction parameter comprising: k) the reference initial binding response; l) late binding response of the reference; m) reference stable initial response; n) reference stable late response; o) the difference between the reference late binding response and the reference early binding response divided by the reference early binding response; and / or p) When the analyte sample solution is composed of a plurality of analyte sample solutions each containing the same analyte but at different analyte concentrations, the equilibrium dissociation constant calculated using the maximum binding capacity from the plurality of analyte sample solutions; 8. The method of claim 1, comprising one or more of the following:

9. the machine learning algorithm is further trained using at least one reference interaction parameter for each analyte sample solution of the second set of analyte sample solutions, the at least one reference interaction parameter being: k) the reference initial binding response; l) late binding response of the reference; m) reference stable initial response; n) reference stable late response; o) the difference between the reference late binding response and the reference early binding response divided by the reference early binding response; and / or p) When the analyte sample solution is composed of a plurality of analyte sample solutions each containing the same analyte but at different analyte concentrations, the equilibrium dissociation constant calculated using the maximum binding capacity from the plurality of analyte sample solutions; 9. The method of claim 1, comprising one or more of:

10. The method of any one of claims 1 to 9, further comprising extracting at least one negative control interaction parameter obtained from the interaction between a negative control solution and a ligand and providing the extracted parameter to the trained machine learning algorithm.

11. 11. The method of claim 1, wherein the machine learning algorithm is further trained using at least one negative control interaction parameter obtained from the interaction of a negative control sample solution with a ligand.

12. The method of any one of claims 1 to 11, further comprising extracting at least one positive control interaction parameter obtained from the interaction between a positive control solution and a ligand and providing the extracted parameter to the trained machine learning algorithm.

13. 13. The method of claim 1, wherein the machine learning algorithm is further trained using at least one positive control interaction parameter obtained from the interaction of a positive control sample solution with a ligand.

14. 14. The method of any one of claims 1 to 13, wherein the machine learning algorithm is selected from the group consisting of a decision tree, a k-nearest neighbor (kNN), a random forest, a k-means, a gradient boosting algorithm, an artificial neural network (ANN), a deep learning algorithm, or a combination thereof.

15. An analytical sensor system (20) for detecting molecular binding interactions between an analyte (4) and a ligand (3) and classifying the observations, comprising: a sensor device (1) including a detector (11) for monitoring the intermolecular interaction between a ligand (3) and a sample solution of an analyte (4) over time; a data generation unit (12) configured to generate response data representative of the time evolution of the observed intermolecular interactions; an extraction unit (13) configured to extract at least one interaction parameter from at least a portion of the generated response data; and a data processing unit (14) configured to receive the at least one interaction parameter and classify the analyte sample solution into at least one quality classification group indicative of an interaction between the sample solution and the ligand based on the at least one interaction parameter using a trained machine learning algorithm.

16. 16. The analytical sensor system (20) of claim 15, further comprising a training center (15) for training the machine learning algorithm by providing interaction parameters extracted from response data from a plurality of observed analyte-ligand interactions to the machine learning algorithm along with at least one quality classification group indicative of an analyte-ligand interaction.

17. A computer program comprising a program code for performing the method according to any one of claims 1 to 14 when the computer program runs on a computer.