Systems and methods for curation of diverse biological and / or chemical data
The data curation method addresses inconsistencies in biological and chemical data by harmonizing and aligning disparate sources, resulting in high-fidelity training data sets that enhance drug discovery and development efficiency.
Patent Information
- Application Number
- PCT/US2025/047660
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-10-11
- Filing Date
- 2025-09-24
- Publication Date
- 2026-04-16
AI Technical Summary
Modern drug discovery and development programs are hindered by inconsistencies and errors in biological and chemical data from diverse sources, which inhibit the effective use of these data sets in training models and downstream processes.
A method for data curation that involves identifying multiple data sets, generating a configuration for data transformation, creating a curated data set, and implementing error reduction to harmonize and align data, thereby enhancing the quality and usability of the data for training models.
The method produces high-fidelity training data sets that reduce errors and inconsistencies, enabling more accurate molecular property prediction models with fewer computational resources, leading to improved drug discovery and development processes.
Smart Images

Figure US2025047660_16042026_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHODS FOR CURATION OF DIVERSE BIOLOGICAL AND / OR CHEMICAL DATARELATED APPLICATION
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 706,374 (filed on October 11 , 2024). The entirety of the foregoing provisional application is incorporated by reference herein.BACKGROUND
[0002] Modern drug discovery and / or development programs rely on data obtained from a large and diverse range of sources such as commercial data sets, academic publications, and patent literature. Whilst such sources provide a rich vein of data, their use to develop prediction models for drug discovery and / or development programs is inhibited by inconsistencies in experimental protocols, reporting procedures, and / or errors arising from transcription, conversion, and the like.
[0003] There is therefore a need for improved methods for curating data from diverse biological and / or chemical data sources to provide improved training data, trained models, and / or data auditing.SUMMARY OF DISCLOSURE
[0004] According to an aspect of the present disclosure there is provided method for curation of diverse biological and / or chemical data, the method comprising identifying, by one or more processors, a plurality of data sets derived from one or more data sources, each of the plurality of data sets comprising one or more raw data points related to properties of biological and / or chemical compounds of interest; generating, by the one or more processors, a first configuration for a data transformation process comprising a sequence of data transformation steps, the first configuration comprising one or more parameter values for configuring one or more steps of the sequence of data transformation steps; generating, by the one or more processors, a first curated data set by executing the data transformation process on the plurality of data sets in accordance with the first configuration, wherein the first curated data set comprises a plurality of harmonized data points related to properties of the biological and / or chemical compounds of interest;implementing, by the one or more processors, an error reduction evaluation of the plurality of harmonized data points of the first curated data set; analyzing, by the one or more processors, output of the error reduction evaluation to detect that a data issue exists in relation to one or more data points within the plurality of data sets and / or the first curated data set; and upon detection that the data issue exists, implementing, by the one or more processors, at least one adjustment to the one or more data points and / or the first configuration to reduce or eliminate the data issue to enhance a downstream process that inputs the enhanced first curated data set as a data input.
[0005] According to a further aspect of the present disclosure there is provided a non-transitory computer readable medium comprising instructions which, when executed by one or more processors, cause the one or more processors to identify a plurality of data sets derived from one or more data sources, each of the plurality of data sets comprising one or more raw data points related to properties of biological and / or chemical compounds of interest; generate a first configuration for a data transformation process comprising a sequence of data transformation steps, the first configuration comprising one or more parameter values for configuring one or more steps of the sequence of data transformation steps; generate a first curated data set by executing the data transformation process on the plurality of data sets in accordance with the first configuration, wherein the first curated data set comprises a plurality of harmonized data points related to properties of the biological and / or chemical compounds of interest; implement an error reduction evaluation of the plurality of harmonized data points of the first curated data set; analyzing, by the one or more processors, output of the error reduction evaluation to detect that a data issue exists in relation to one or more data points within the plurality of data sets and / or the first curated data set; and upon detection that the data issue exists, implement at least one adjustment to the one or more data points and / or the first configuration to reduce or eliminate the data issue to enhance a downstream process that inputs the enhanced first curated data set as a data input.
[0006] According to an additional aspect of the present disclosure there is provided a device comprising a memory storing instructions which, when executed by one or more processors, cause the device to identify a plurality of data sets derived from one or more data sources, each of the plurality of data sets comprising one or more raw data points related to properties of biological and / or chemical compounds of interest; generate a first configuration for a data transformation process comprising a sequence of data transformation steps, the first configuration comprising one or more parameter values for configuring one or more steps of thesequence of data transformation steps; generate a first curated data set by executing the data transformation process on the plurality of data sets in accordance with the first configuration, wherein the first curated data set comprises a plurality of harmonized data points related to properties of the biological and / or chemical compounds of interest; implement an error reduction evaluation of the plurality of harmonized data points of the first curated data set; analyzing, by the one or more processors, output of the error reduction evaluation to detect that a data issue exists in relation to one or more data points within the plurality of data sets and / or the first curated data set; and upon detection that the data issue exists, implement at least one adjustment to the one or more data points and / or the first configuration to reduce or eliminate the data issue to enhance a downstream process that inputs the enhanced first curated data set as a data input.
[0007] Further aspects and embodiments of the present disclosure are set out in the appended claims.
[0008] In accordance with the above, and with the disclosure herein, the present disclosure includes effecting a transformation or reduction of a particular article to a different state or thing, e.g., the transformation of a plurality of diverse chemical biological and / or chemical data sets, to a different state or thing, e.g., the generation training data based on a curated data set that has been filtered to remove inconsistencies, errors, and / or noise.
[0009] Still further, the present disclosure includes improvements in computer functionality or in improvements to other technologies at least because the disclosure herein discloses systems and methods for reducing error in underlying computing devices, e.g., by producing enhanced (e.g., curated) data sets filtered from disparate and inconsistent data sets obtained from a diverse range of sources. Still further, once a curated data set is generated, high-fidelity and high-quality training data sets may then be trained or generated therefrom. Such training data sets may then be used to train or update models, such as new or updated machine learning models, in order to produce more accurate and less error prone output, and, as a result, high-quality drug products, such as therapeutics. Still further, the prediction models, when deployed on the underlying system, allows the systems and methods of the present disclosure to execute with fewer iterations, and use fewer computing resources, than prior art related systems and methods. That is, the present disclosure describes improvements in the functioning of the computer itself or “any other technology or technical field” because the increased predictive improvement provided by the prediction model allows the underlying computer system to utilize less processing and memory resources compared to prior art systems and methods. This at least because the predictionmodel, trained on a data set generated from curated data, can generate or determine a molecular property with higher accuracy, without the need for various tests and / or empirical computer simulation across a wide range of tests using multiple compute cycles and data. Therefore, use of the prediction model results in fewer compute cycles, or otherwise iterations, that has less of an impact on the underlying computing device compared to previous prior art systems and methods. Moreover, the disclosure herein allows for identification and use of high-fidelity data sets, which reduces the need for additional computational cycles further.
[0010] Still further, the present disclosure includes specific features other than what is well- understood, routine, conventional activity in the field, and / or otherwise adds unconventional steps that confine the disclosure to a particular useful application, e.g., systems and methods for the curation of diverse biological and / or chemical data, which can be used, for example, to train more accurate molecular property prediction models, provide actionable insights into data issues and / or automatically remediate such data issues, and / or improve downstream tasks such as drug discovery and development.
[0011] Advantages will become more apparent to those of ordinary skill in the art from the following description of the preferred embodiments which have been shown and described byway of illustration. As will be realized, the present embodiments may be capable of other and different embodiments, and their details are capable of modification in various respects. Accordingly, the drawings and description are to be regarded as illustrative in nature and not as restrictive.BRIEF DESCRIPTION OF DRAWINGS
[0012] Embodiments of the present disclosure will now be described, by way of example only, and with reference to the accompanying drawings, in which:
[0013] Figure 1 A shows a method for data curation — i.e., for the curation of diverse biological and / or chemical data — according to an aspect of the present disclosure;
[0014] Figure 1 B shows further steps forming part of the method shown in Figure 1 A according to an embodiment of the present disclosure;
[0015] Figure 1 C shows further steps forming the method shown in Figure 1A according to an embodiment of the present disclosure;
[0016] Figure 2 illustrates an example data curation process to generate a curated data set of properties of compounds which are known to interact with and / or modulate Poly [ADP-ribose] polymerase 1 (PARP1 ) according to an embodiment of the present disclosure;
[0017] Figures 3A-3C show plots of distributional representations of a curated data set according to an embodiment of the present disclosure;
[0018] Figure 3D shows a correlation analysis of pairs of data points within a curated data set according to an embodiment of the present disclosure;
[0019] Figure 3E shows a plot of exponent issues within a curated data set according to an embodiment of the present disclosure;
[0020] Figure 4 shows a portion of a curated data set generated from the example data curation process shown in Figure 2 according to an embodiment of the present disclosure; and
[0021] Figure s shows an example computing system for carrying out the methods of the present disclosure.TECHNICAL FIELD
[0022] The present disclosure relates to curating diverse data sets. More particularly, but not exclusively, the present disclosure relates to generating curated data sets from diverse biological and / or chemical data; more particularly, but not exclusively, the present disclosure relates to generating molecular property training data sets from such curated data sets.DETAILED DESCRIPTION
[0023] Biological and / or chemical data sets for training target specific molecular property models are typically derived from multiple disparate data sources and so may include noisy, incorrect, inconsistent, and / or non-harmonized, data. Whilst combining such data sets may help improve the richness and diversity of data about compounds and their interactions with various targets, the process of combining them into a useful and usable form is made difficult due to the conflicts and inconsistencies that exist between data sets from different sources. This may then inhibit the use of these data sets within drug discovery / development and molecular design tasks since downstream data generated from disparate and potentially conflicting data sources, without adequate curation, may replicate or exacerbate these conflicts and inconsistencies.
[0024] The systems and methods of the present disclosure seek to alleviate these issues and to extract maximum value from the data by automatically harmonizing and aligning raw qualitative and quantitative data from a plurality of disparate data sets into a consistent and coherent form. The data curation method of the present disclosure intelligently combines, filters, and / or transforms data from these disparate sources to create curated / harmonized data sets. The data curation method of the present disclosure is agnostic to the reliability of the data sources and can curate diverse data which can subsequently be used to generate training data for model training (e.g., molecular property model training). Moreover, the systems and methods of the present disclosure provide insights into the curated data and the process(es) for curating the data to help improve the quality, quantity, and usability of the data and process.
[0025] Figure 1 A shows a method 100 for data curation — i.e., for the curation of diverse biological and / or chemical data — according to an aspect of the present disclosure.
[0026] The method 100 comprises the steps of identifying 102 a plurality of data sets, generating 104 a first configuration, generating 106 a first curated data set, implementing 108 an error reduction evaluation, analyzing 110 the output of the error reduction evaluation to detect that a data issue exists, and implementing 112 at least one adjustment to reduce or eliminate the data issue. The method 100 may further include the steps shown in Figure 1 B (path “1 ” in Figure 1A) and / or Figure 1 C (path “2” in Figure 1A).
[0027] At the step of identifying 102, a plurality of data sets derived from one or more data sources are identified. Each of the plurality of data sets comprise one or more raw data points (alternatively referred to as rows, data objects / instances, or tuples) which include qualitative and / or quantitative values related to properties of compounds of interest. Compounds of interest include compounds that are known or predicted to interact with and / or modulate at least one target or at least two targets, compounds that are present within a chemical space that is of interest, and / or compounds having or predicted to have one more desired features, such as high solubility, high stability, low toxicity, etc. Each of the one or more raw data points can comprise a value or a vector of values identified or extracted from a data source. For example, if a data source contains data from three different assays for a compound, then the raw data point for the compound is a vector of three values corresponding to the assay data for the compound obtained from the data source. The raw data points can further include data regarding the experimental setup used to obtain the properties of the compounds of interest, bibliographic data related to the data source, measurement and parameter information, target information, and the like. As such,a raw data point can include quantitative and qualitative data related to a compound and at least one target.
[0028] The plurality of data sources can include experimental biological and / or chemical data sources. For example, the plurality of data sources may include a database which stores results from biological and / or chemical experiments. Each record within a database of experimental biological and / or chemical data can include data identifying a molecule (i.e., a compound, protein, etc.) and quantitative experimental results for the molecule. Metadata regarding the experimental setup (e.g., assay, assay type, etc.) may also be recorded with the database. The data identifying the molecule can be in the form of a molecular representation such as a SMILES string, a molecular graph, or the like. Example quantitative experimental results for the molecule include basic information concerning the molecule, such as molecular weight, overall charge, charge distribution, structural conformation, solubility, crystal structure, etc. Further data concerningthe molecule includes howthe molecule interacts with other molecules (e.g., targets, receptors, etc.) and how the molecule functions within a physiological system. This data can be obtained from in vitro or in vivo experiments, including in vitro test systems, animal models and clinical trials. The data can include structure-activity relationship (SAR) data, bioactivity data, toxicity data, stability data, binding affinity data, etc.
[0029] The plurality of data sources can include publicly and / or commercially available databases such as GOSTAR, ChEMBL, PubChem, and the like. The plurality of data sources can also include literature-based sources such as patent applications, other scientific literature data sources such as scientific publications (e.g., journal articles, conference publications, textbooks, etc.), clinical trial data including clinical trial results and metadata, etc. Experimental data from such sources can be extracted and added to a database or repository of data. A configuration can be used to define which of the plurality of data sources should be included as part of the extraction process.
[0030] A data set can be extracted from a data source by identifying raw data points within the data source which are related to compounds of interest, e.g., raw data points related to compounds which are known or predicted to interact with and / or modulate at least one target. In one embodiment, where the goal is to identify candidate compounds against a particular target, a data set is extracted from a data source by searching for data within the data source related to compounds which are known or predicted to interact with and / or modulate directly or indirectlythe target. For example, extracting raw data points which have experimental data against the target.
[0031] A raw data point (alternatively referred to as a data row, a data object, a data instance, or a tuple) is a set of qualitative and / or quantitative values related to the interaction of a compound and a target which are extracted and / or derived from a data source. Quantitative data can include numerical values of one or more properties of the compound (e.g., a minimum and maximum value of the property reported for the interaction of the compound and the target within the data source). The properties can include one or more of: a physiological activity; a pharmacological activity (including a pharmacokinetic property and / or a pharmacodynamic property); a toxicity or selectivity; a reactivity; a binding affinity, etc. Qualitative data can include: a molecular representation of the compound to which the raw data point relates; a target name; an assay name and type; an effect; an activity; a mechanism of action (MoA); information regarding the parameter and units used to report the quantitative data; an indication of the source of the raw data point; a tissue identifier or name related to the experimental subject; a species of the experimental subject; and bibliographic data (e.g., article title, authors, journal title, volume no., issue no., pages, publication date, etc.). As such, a raw data point may be considered to contain quantitative data regarding a property of interest and metadata (qualitative data) related to the quantitative data. Including such metadata (qualitative data) allows for a greater level of control over the filtering and transformation of the data as part of the curation process.
[0032] The properties included within the one or more raw data points (i.e., the properties of interest) can be set by a user or defined according to a configuration (e.g., a configuration file). For example, a configuration file may include a target and a set of properties for compounds within the plurality of data sources known or predicted to interact with and / or modulate the target. A configuration file comprising the target “HDAC6” and properties “{selectivity toxicity}” would extract selectivity and toxicity data from the plurality of data sources for all compounds which are known or predicted to interact with and / or modulate HDAC6. Alternatively, the properties can be set according to the assay type. For example, a configuration setting the assay type to “{binding}” would result in binding activity data being extracted from the plurality of data sources. The configuration for setting the properties of interest can be included as part of the configuration for the data transformation process described below. Alternatively, the configuration for setting the properties of interest is separate from the data transformationprocess configuration and is included as part of a configuration file provided and / or stored in any suitable format such as JSON, YAML, XML, and the like.
[0033] In one embodiment, the method 100 further comprises, prior to the step of identifying 102 the plurality of data sets, the step of identifying the at least one target. The at least one target can be identified by obtaining a representation of the at least one target (e.g., a name or identifier of the target) from a user. Alternatively, the at least one target can be provided as part of a configuration file (as described above).
[0034] At the step of generating 104, a first configuration for a data transformation process is generated. The data transformation process comprises a sequence of data transformation steps and the first configuration comprises one or more parameter values for configuring one or more steps of the sequence of data transformation steps.
[0035] The data transformation process (alternatively referred to as a data curation process, a curation process, or a data harmonization process) curates the plurality of data sets into a single, harmonized, data set. The data transformation process includes a sequence of configurable transformation steps which each apply a filter, a weight, a conversion, or a transformation to the data within the plurality of data sets in order to generate the curated data set. The first configuration parameterizes, or configures, the data transformation process and so controls the generation of the curated data set from the plurality of (raw) data sets. The first configuration comprises a set of parameters and corresponding parameter values. A parameter of the first configuration is linked to a step of the data transformation process such that the corresponding parameter value is used to configure the step when the data transformation process is executed. A parameter value can be a number, a string, a list of values, and the like. In one embodiment, one or more of the parameter values of the first configuration include a regular expression. For example, a filtering step may be parameterized by the regular expression “F ($ | [Aerml] )” in order to identify, from SMILES representations, compounds containing Fluorine.
[0036] The first configuration can be a file (e.g., a plain text file or the like) comprising the set of parameters and parameter values to configure the data transformation process. In one embodiment, the first configuration is, or is extracted from, a JSON file, a YAML file, an XML file or the like. The first configuration can be generated by loading a configuration for the data transformation process (e.g., from memory or from a persistent storage location). The configuration may be a default configuration comprising an indication of the sequence of datatransformation steps to execute as part of the data transformation process along with one or more default parameter values for configuring one or more steps of the sequence of data transformation steps. Alternatively, the first configuration can be generated by obtaining the one or more parameter values for configuring one or more steps of the sequence of data transformation steps from a user (e.g., as part of an interactive configuration setup process).
[0037] The data transformation process (which is configured by the first configuration) comprises a sequence of transformation steps including one or more filtering steps, one or more conversion steps, one or more weighting steps, and / or one or more transformation steps. The steps are sequential in that the result of applying a first step of the data transformation process is provided as input to the second step of the process. As such, each step within the data transformation process receives an input data set as input and generates an output data set as output (where the output data set is a subset, conversion, weighting, or transformation of the input data set).
[0038] A filtering step of the data transformation process filters (reduces or refines) the data within the input data set according to a filtering criterion such that the output data set corresponds to a subset of the input data set which satisfies the filtering criterion. The filtering criterion can be applied to quantitative and / or qualitative data within the input data set. For example, a filtering criterion may be used to restrict the output data set to compounds within the input data set which are obtained using a specific type of assay. Alternatively, a filtering criterion may be used to restrict the output data set to compounds which satisfy certain substructure filters. As a further example, a filtering criterion may be used to restrict the output data set to data within the input data set which have a value for an IC50 binding activity property of <100pM. A filtering step can be parameterized, or configured, by a parameter value included in the first configuration. Continuing the previous example, the first configuration may include a parameter value for the IC50 binding activity property (e.g., “IC50_binding_activity : 100” where “100” is the configurable / adjustable parameter value for the “IC_50_binding_activity” filtering step parameter).
[0039] A conversion step of the data transformation process applies one or more transformations to data within the input data set to convert the units of one or more properties within the output data set. As such, a conversion step helps to ensure that properties within the curated data set are reported in the same units and are thus directly comparable. For example, a conversion step may be used to convert all activation properties in the input data set (which maybe reported in various units such as nM or / zg / ml) to nM in the output data set. A conversion step can be parameterized, or configured, by a parameter value included in the first configuration. For example, the first configuration may include a parameter value indicating that all concentration properties are to be reported as nM in the curated data set.
[0040] A weighting step of the data transformation process applies a weighting to the input data set such that the output data set comprises a weighted representation of one or more data points within the input data set which satisfy a weighting criterion. As such, a weighting step first identifies any data points within the input data set which satisfy the weighting criterion and then applies a weight to the identified data points. For example, the weighting criterion may specify that the weight should be applied to property values obtained from a specific data source (e.g., experimental data obtained from a higher quality source, such as a peer reviewed journal). A weighting step can be parameterized, or configured, by parameter values included in the first configuration. For example, the first configuration may comprise a “weighting” parameter with a corresponding parameter value for the weight to apply. The first configuration can further comprise a parameter and corresponding parameter value for the weighting criterion such that the weight is applied only to data points within the input data which satisfy the weighting criterion specified by the weighting criterion parameter value.
[0041] A transformation step of the data transformation process applies a transformation or normalization to the input data set such that the output data is a transformed or normalized representation of the input data set. A transformation step can be parameterized, or configured, by one or more parameter values included in the first configuration. For example, the first configuration can include a parameter “normalize” with a corresponding parameter value of “True” such that the output data set corresponds to the normalized input data set (e.g., normalized to the range [0,1 ]). Additionally, or alternatively, the first configuration can include a parameter to restrict the transformation or normalization to one or more parameters identified by one or more parameter values in the first configuration.
[0042] At the step of generating 106, a first curated data set is generated by executing the data transformation process on the plurality of data sets in accordance with the first configuration. The first curated data set comprises a plurality of harmonized data points related to properties of the compounds of interest. The step of generating 106 a first curated data set may be alternatively referred to as curating the first curated data set, creating the first curated data set, or assembling the first curated data set.
[0043] The data transformation process can be a script in a scripting language (e.g., BASH, Python, Ruby, etc.) or a program written in a programming language (e.g., Java, C, C++, Fortran, etc.) comprising the sequence of data transformation steps which themselves may be operations performed as part of the script or program or calls to a further script or program to perform the data transformation step. The first configuration is passed to the data transformation process such that execution of the data transformation process is performed in accordance with the parameter values defined in the first configuration.
[0044] As stated above, the data transformation process takes the plurality of data sets as input, executes the sequence of data transformation steps (in accordance with the parameter values defined in the first configuration) on the plurality of data sets, and generates the first curated data set as output. For example, if the data transformation process consists of the sequence of data transformation steps-> f2-> f3(where each data transformation step takes a data set as input and generates a data set as output), then the curated data set D' for a plurality of data sets D provided as input would be D' = f3(j ifi P ) - Here, the plurality of data sets are provided to the data transformation process as a single data set as a result of being concatenated or otherwise combined into a single data set.
[0045] The first curated data set can take the form of a table, array, or matrix, where each row represents harmonized data obtained from the plurality of data sources for a compound, and each value within the row is a quantitative or qualitative attribute value of the compound. For example, a row of the first curated data set may contain a molecular identifier, a minimum property value, a maximum property value, a property name (e.g., “binding”), a parameter name (e.g., IC50, EC50, etc.), a units identifier, a target name, an assay name and assay type, and an identifier of the data source from which the row of data was extracted. In one embodiment, each row comprises a unique identifier.
[0046] At the step of implementing 108, an error reduction evaluation of the plurality of harmonized data points of the first curated data set is implemented (e.g., executed or performed). The error reduction evaluation is a quantitative evaluation of the plurality of harmonized data points of the first curated data set.
[0047] Whilst the above-described curation process harmonizes the plurality of data sets, the first curated data set may nevertheless include data issues which require remediation in order to improve the quality and consistency of the curated data. To identify a potential data issue withinthe curated data set, an error reduction evaluation (alternatively referred to as a quantitative evaluation) of the curated data set is performed. In general, the error reduction evaluation comprises a set of quantitative analyses which are statistical or predictive analyses of the plurality of harmonized data points which help reveal trends, abnormalities, outliers, or other data issues related to the first curated data set. By identifying such issues, the quality and usability of the first curated data set can be improved by remediating the issues. This in turn helps to improve downstream tasks and operations which utilize the curated data set such as improving predictive models trained on data generated from the curated data set, data visualizations generated from data within the curated data set, etc.
[0048] The quantitative analysis carried out (e.g., executed or performed) as part of the error reduction evaluation can include calculating a correlation of a property of the plurality of harmonized data points. The correlation can be a correlation or comparison between values recorded for the same compound and / or assay from two or more different data sources. For example, a correlation between activity values for oxtriphylline reported in a journal article and a patent publication. Additionally, or alternatively, the correlation can relate to the parameters used to report the property values (e.g., a correlation of EC50 to IC50 activity values for the same or different compounds). Alternatively, the correlation can be a correlation between values recorded for two similar assays expected to be correlated. Alternatively, the correlation can be a correlation or comparison between units or exponents (e.g., a correlation of exponents for data points related to a compound which have the same property value).
[0049] The quantitative analysis carried out (e.g., executed or performed) as part of the error reduction evaluation can include determining a range of values for the plurality of harmonized data points. For example, the quantitative analysis may be used to calculate the maximum and minimum values of each property within the plurality of harmonized data points. The observed range of values (maximum and minimum) may be compared to an acceptable (expected or predetermined) range of values for each property. The difference between the observed and acceptable range of values can then be calculated.
[0050] The quantitative analysis carried out (e.g., executed or performed) as part of the error reduction evaluation can include performing a statistical analysis of a property of the plurality of harmonized data points. For example, a standard deviation or variance of values of the property across the plurality of harmonized data points. As a further example, a mean, median, or modal average of the property may be calculated for the property across the plurality of harmonized datapoints. The statistical analysis therefore provides a property level statistical summary of the plurality harmonized data points. In one embodiment, the statistical analysis is performed on molecular attributes within the plurality of harmonized data points such as molecular weight, number of rotational bonds, and the like.
[0051] The quantitative analysis carried out (e.g., executed or performed) as part of the error reduction evaluation can include generating a distributional representation of the plurality of harmonized data points. The distributional representation can be determined across all properties or for one or more properties. For example, a histogram of values for a property can be calculated from the values of the property across the plurality of harmonized data points. The histogram of values provides a distributional representation of the property across the plurality of harmonized data points. A probability density function can then be fit to the histogram of values to provide a parametric estimation of the distribution of values. The parameters of the distribution can then provide a distributional representation of the plurality of harmonized data points. Alternatively, the distributional representation is a distribution of certain properties or aspects of the plurality of harmonized data points. For example, the distribution of relative contribution of each data source to the curated data set.
[0052] The quantitative analysis carried out (e.g., executed or performed) as part of the error reduction evaluation can include calculating a confidence (or confidence score) of one or more properties of the plurality of harmonized data points. The confidence can be associated with a data source from which a property of a harmonized data point is extracted. For example, if a data source is known or predicted to report inaccurate results, the data points / properties obtained from the data source can be assigned a lower confidence than those obtained from a more reliable data source. Additionally, or alternatively, the confidence can be obtained from predetermined confidence bounds associated with a data source and / or a property. For example, data obtained from a first academic publication known to have a strong peer review process and thus high credibility may be assigned a higher confidence than data obtained from a second academic publication with no peer review process.
[0053] The quantitative analysis carried out (e.g., executed or performed) as part of the error reduction evaluation can include performing a clustering analysis on the plurality of harmonized data points. That is, each of the plurality of harmonized data points are assigned a cluster or group identifier by a clustering algorithm (with the expectation that harmonized data points associated with similar compounds would appear in the same cluster or group). The clustering analysis canbe performed on one or more of the properties of the plurality of harmonized data points. For example, the plurality of harmonized data points can be clustered, or grouped, based on reported values for binding affinity (i.e., one property) or for binding affinity and selectivity (i.e., two properties). Suitable clustering algorithms include k-means clustering, Gaussian Mixture Modelling, spectral clustering, DBSCAN, and the like.
[0054] One or more of the above quantitative analyses can be carried out (e.g., executed or performed) according to groupings defined by the metadata (i.e., qualitative data) of the plurality of harmonized data points. For example, the quantitative analysis can include determining the distribution by year of the data sources used to generate the first curated data set. As a further example, the correlation of property values for IC50 in comparison to property values for EC50 can be determined.
[0055] In one embodiment, all the above quantitative analyses are carried out (e.g., executed or performed) such that all data issues discussed below may be identified. Alternatively, only a subset of the quantitative analyses is carried out thereby restricting the data issues identifiable by the step described below. The set of quantitative analyses carried out may be determined based on an indication of which data issues are to be identified (e.g., as indicated within a configuration file).
[0056] At the step of analyzing 110, the error reduction evaluation is analyzed to detect if a data issue exists in relation to one or more data points within the plurality of data sets and / or the first curated data set.
[0057] The data issue is identified from the output of the reduction evaluation and indicates that there is a quality issue related to data points within the plurality of data sets and / or the first curated data set. The quality issue may hinder or impede the applicability and usefulness of the first curated data set without remediation. For example, a predictive model trained on a data set having a data issue may not provide the same predictive performance as a predictive model trained on a data set where the data issue has been remediated (overcome, reduced, or eliminated).
[0058] The data issue can be an abnormality issue identified by determining that the one or more data points fall outside a range of values. That is, an abnormality issue may be identified by determining that a range of values for a property of the one or more data points is significantly different to an expected range of values for the one or more data points. As such, the difference in ranges indicates that the one or more data points are abnormal in relation to the expected range.
[0059] The data issue can be an outlier issue identified by the one or more data points being outliers with respect to a representative value of the plurality of harmonized data points. An outlier issue can be identified with regard a single property or a plurality of properties. As such, a data point can have an outlier issue with respect to a single property, more than one property, or all properties. In the case of a single property, the representative value is a single numerical value. In the case of a plurality of properties, the representative value is a vector of two or more values. The representative value of the plurality of harmonized data points can be an expected value of the plurality of harmonized data points which is a mean, median, or mode of the plurality of harmonized data points. Alternatively, the expected value is a predetermined value (which may be associated with a predetermined expected value for a property or properties). The one or more data points can be identified as outlier data points if the values for the property or properties of the one or more data points are greater than the representative value for the property or properties by more than a threshold amount. For example, a data point is identified as an outlier data point for a property if the data point’s value for the property is more than 3 standard deviations from the representative value. As a further example, the threshold may be a predefined value related to a property such that a data point is identified as an outlier data point with respect to the property if the data point’s value for the property is beyond the predefined value.
[0060] The data issue can be a transcription issue identified from an exponent or unit-based correlation analysis. A transcription issue can be understood as corresponding to an error within a data set where the units have been incorrectly assigned (e.g., by a human or as a result of document processing such as optical character recognition). For example, an activity value which should have been reported in nM is reported in pM. An exponent or unit-based correlation analysis is used to identify a transcription issue by identifying a pair of data points within the curated data set which relate to the same compound and have identical property values, but the exponent (unit) differs by an order of magnitude.
[0061] The data issue can be a grouping issue based on a clustering analysis where the one or more data points are assigned to an incorrect, unexpected, or erroneous, cluster or group. For example, a clustering analysis (e.g., a k-means clustering) may assign data points for compounds related to a first group to a first cluster but assign a data point for a compound within the same group, which should be assigned to the same cluster or group as other compounds within the first group, to a second different cluster. This then indicates that there is a data issue in relation to the data point for the compound.
[0062] The skilled person will appreciate that the above-described data issues are not necessarily mutually exclusive, and more than one data issue may be identified from the analysis of the output of the error reduction evaluation. As stated previously, which data issues are potentially identified can be determined by a configuration comprising a list of which data issues to look for.
[0063] Upon, or as a result of, detectingthat no data issue exists, the method 100 may terminate or proceed to either or both of paths “1 ” and “2” shown in Figure 1A. Upon, or as a result of, detecting that a data issue exists, the method 100 proceeds to the step of implementing 112 at least one adjustment.
[0064] At the step of implementing 112, at least one adjustment is implemented to the one or more data points and / or the first configuration to reduce or eliminate the data issue and to enhance a downstream process that inputs the enhanced first curated data set as a data input. That is, in consequence of the data issue(s) being reduced or eliminated within a subsequent data set (e.g., a data set generated from the modified or enhanced first curated data set, a second curated data set, etc.), any downstream processes that utilize that data set will be enhanced as a result of the data issue(s) being reduced or eliminated.
[0065] To remediate (overcome, reduce, or eliminate) the data issue and improve the quality, usability, and effectiveness of the curated data set, at least one adjustment is implemented or applied to the one or more data points and / or the first configuration. The adjustment(s) implemented depend on the data issue identified. As such, the at least one adjustment is based on the data issue (or data issues) identified. The at least one adjustment can be made to a property value or plurality of property values, a data point or plurality of data points, and / or a data set or a plurality of data sets.
[0066] The at least one adjustment can include a transformation of the one or more data points within the plurality of harmonized data points of the first curated data set. The at least one adjustment can include removing the one or more data points from the plurality of harmonized data points. The at least one adjustment can include removing data points corresponding to the one or more data points from the plurality of data sets. The at least one adjustment can include weighting the one or more data points based on the error reduction evaluation. The weighting can be applied to the one or more data points in the curated data set or data points corresponding to the one or more data points from or in the plurality of data sets. The at least one adjustment caninclude generating a second configuration for the data transformation process, the second configuration comprising an instruction to exclude or weight the one or more data points from a second curated data set generated from the second configuration.
[0067] As shown in Figure 1A, after the step of implementing 112 the at least one adjustment the method 100 may terminate or proceed to the steps shown in Figure 1 B (path “1 ” in Figure 1A) and / or the steps shown in Figure 1 C (path “2” in Figure 1A). Additionally, or alternatively, the output of the data curation method shown in Figure 1A is output to another method, process, or module.
[0068] Figure 1 B shows further steps performed as part of the method 100 shown in Figure 1 A according to an embodiment of the present disclosure.
[0069] Figure 1 B shows the step of generating 114 a training data set, training 116 a prediction model on the training data set, and the optional step of predicting 118 a property value using the prediction model.
[0070] At the step of generating 114, a training data set is generated from the first curated data set. If the at least one adjustment step (performed at the step of implementing 112) results in the generation of a second curated data set, then the training data set is generated from the second curated data set. Each row of the training data set is associated with a property value of a target reported for a compound. As such, a row of the training data set comprises a molecular representation of a compound and at least one property value for the compound. The training data set is thus generated from the first curated data set by extracting the relevant values — molecular representation and property value(s) — from the first curated data set. For example, a training data set may be generated by extracting a molecular representation value and a corresponding binding activity value from each row of the first curated data set. In instances where property values are represented within the first curated data set as a range of values, a representative value of the range is generated and included in the training data set (e.g., a mean value, a median value, etc.).
[0071] In one embodiment, the molecular identifiers, or molecular representations, in the training data set arefeaturized — i.e., converted ortransformed into a representation usable by and suitable fortrainingthe prediction model — usingany suitable featurization algorithm. For example, each compound within the data set may be identified using a SMILES string or representation. Prior to training the prediction model, or as part of the training process, each SMILES string or representation is converted to a numerical representation such as a one-hot encoded vector suchthat the prediction model is trained to predict a vector of property values from an input one-hot encoded binary vector representing a SMILES string. As a further example, a molecular fingerprinting algorithm such as extended connectivity fingerprints may be used to generate a binary “fingerprint” vector of a compound such that the prediction model is trained to predict a vector of property values from a binary fingerprint vector. Other featurization algorithms include bag-of-bonds, the Coulomb matrix, and symmetry functions.
[0072] At the step of training 1 16, a prediction model is trained on the training data set. The prediction model is trained to predict a property value for an input molecular representation (or a featurization thereof). The prediction model can be a classification model, a regression model, a ranking model, etc. For example, the prediction model can be a regression model trained to predict inhibition of Cyclin-dependent kinase 2 (CDK2) from a training data set extracted from a curated data set specifically generated for CDK2. In this instance, the training data set comprises values related to the inhibition of CDK2 measured using IC50 for a plurality of compounds (extracted from a diverse range of sources as described above) and a regression model is trained to predict, for an input molecular representation of a compound, a value for the inhibition of CDK2 for the compound.
[0073] Examples of suitable prediction models include a linear regression model, a polynomial regression model, a logistic regression model, a Bayesian regression model, a support vector machine (SVM), a multilayer perceptron, a decision tree, a random forest, a boosting model, and the like. The skilled person will appreciate that any suitable approach as is known in the art can be used to train such models on the training set generated at step 114. For example, a standard approach for training a prediction model comprises performing cross-validation to train the prediction model on the training data generated at step 114. Typically, cross-validation involves splitting the training data into K-folds (approximately equal partitions or sets of the training data) and withholding a single fold as a test set and, one by one, using one of the remaining folds as a validation set and the remaining K-2 folds as a training set. The model is then repeatedly trained on the training set using different model hyperparameters and the performance validated on the validation set. Once the best performing hyperparameters are obtained, the model trained according to the best hyperparameters are evaluated on the test set. For training classification models, the cross-validation strategy can be stratified such that the proportion of training instances within each category or class is approximately the same across each fold. Model hyperparameters can be selected using any suitable approach such as grid search or randomizedsearch. Model performance can be estimated using any suitable performance measure and is dependent on the type of model being trained (e.g., mean square error for regression, binary cross entropy for classification, ranking loss for ranking, etc.).
[0074] The trained prediction model may then be used to predict a property value for a previously unseen compound by providing a molecular representation of the previously unseen compound to the prediction model. The trained prediction model can be output and / or used as part of a drug discovery or development process. Advantageously, by identifying and remediating potential data issues in the curated data set from which a training data set is extracted, models trained on the training data set have improved predictive performance which can in turn improve the outcomes of downstream tasks which utilize such predictive models (e.g., improved patient outcomes, improved target identification for drug discovery / development, etc.).
[0075] At the optional step of predicting 118, a property value for a compound is predicted using the (trained) prediction model. That is, a molecular representation of the compound is provided to the prediction model and a predicted property value for the compound is provided as output by the prediction model. In one embodiment, the molecular representation is featurized (converted or transformed) into a representation corresponding to that for which the prediction model was trained for (e.g., converted from a SMILES string to a one-hot encoded vector).
[0076] Figure 1 C shows further steps performed as part of the method 100 shown in Figure 1A according to an embodiment of the present disclosure. Figure 1 C shows the step of generating 120 a report and outputting 122 the report for display to a user. The steps shown in Figure 1 C may be performed in conjunction with, or instead of, the steps shown in Figure 1 B.
[0077] At the step of generating 120, a report comprising a representation of the first curated data set is generated. The report may comprise a representation of the error reduction evaluation of the plurality of harmonized data points.
[0078] In general, the report provides a summary or overview of the data curation process and the corresponding outcome (the curated data set). Due to the size and complexity of the data sources and data sets involved in generating the curated data set, it is often difficult for a user to understand how the curated data was generated and identify the provenance of the data included within the curated data set. Such information can be important when seeking to use the curated data set, or data sets derived therefrom, in real world clinical applications. The report provides a “data audit” for the curated data set which allows a user to understand the hidden state of thecurated data set as well as the process used to generate the curated data set. The report may include one or more of: a list of data sources from which the first curated data set is derived; an indication of data points from the plurality of data sets which are included within the first curated data set; an indication of data points from the plurality of data sets which are excluded from the first curated data set; summary statistics; and one or more visualizations.
[0079] The indication of data points from the plurality of data sets which are included within or excluded from the first curated data set can take the form of a list or table of data points, a summary of the data points, or any other suitable representation. Providing the indication of data points which are included within the first curated data set enables a user to identify data points within the first curated data set which are incorrectly included and should be excluded (e.g., by adjusting the configuration of the data transformation process or by removing the data points from the first curated data set). Conversely, the indication of data points from the plurality of data sets which are excluded from the first curated data set enables a user to identify data points within the plurality of data sets which should be included within the first curated data set (e.g., by adjusting the configuration of the data transformation process or by adding the data points directly to the first curated data set). As such, the report provides explainable insights into the data curation process which allow a user to modify or adapt the data curation process to improve the quality, quantity, and / or effectiveness of curated data generated by the data curation process.
[0080] In one embodiment, the report comprises at least one visualization based on the error reduction evaluation. The report can comprise one or more plots, charts, or graphics which visually convey characteristics of the first curated data set. For example, the report may comprise a plot showing the distribution of values of a property and / or a plot showing the correlation between two or more features. Advantageously, providing at least one visualization as part of the report supports exploratory data analysis and enables a user to gain insights into the structure and distribution of the first curated data set whilst also allowing them to identify potential inconsistencies or discrepancies.
[0081] In one embodiment, the report is a static document such as a PDF file or image. Alternatively, the report is a dynamic document such as an HTML file or the like. Advantageously, this allows for interactive elements to be embedded within the report to allow the user to interact with the features of the report (e.g., data filtering, plot display adjustment, etc.).
[0082] At the step of outputting 122, the report is output for display to a user.
[0083] In one embodiment, outputting the report comprises storing, or saving, the report to a persistent storage such as a non-volatile memory, a non-transitory medium, or the like. Additionally, or alternatively, outputting the report comprises transmitting the report via a network (e.g., a local area network, a wide area network, and the like), or displaying the report, or a portion thereof, for review by a user.
[0084] Figure 2 illustrates an example data curation process to generate a curated data set of properties of compounds which are known to interact with and / or modulate Poly [ADP-ribose] polymerase 1 (PARP1 ).
[0085] The plurality of data sets 202 comprise data sets 204 obtained from a first data source and data sets 206 obtained from a second data source. Data sets 204 include a first data set 204-1 and a second data set 204-2 each comprising raw data points related to properties of compounds of interest and data sets 206 include a third data set 206-1 and a fourth data set 206-2 each comprising raw data points related to properties of compounds of interest.
[0086] Figure 2 further shows a data transformation process 208 comprising a sequence of data transformation steps 210. The sequence of data transformation steps 210 comprise an assay filtering step 212 having a first plurality of parameters 212-1 , a source filtering step 214 having a second parameter 214-1 , a species filtering step 216 having a third parameter 216-1 , a parameter filtering step 218 having a fourth parameter 218-1 , a molecular attribute filtering step 220 having a second plurality of parameters 220-1 , a substructure filtering step 222 having a fifth parameter 222-1 , a conversion step 224 having a sixth parameter 224-1 , and a deduplication step 226. The sequence of data transformation steps 210 are parameterized by a configuration 228 comprising a first plurality of parameter values 230 for configuring the first plurality of parameters 212-1 of the assay filtering step 212, a second parameter value 232 for configuring the second parameter 214-1 of the source filtering step 214, a third parameter value 234 for configuring the third parameter 216-1 of the species filtering step 216, a fourth parameter value 236 for configuring the fourth parameter 218-1 of the parameter filtering step 218, a second plurality of parameter values 238 for configuring the second plurality of parameters 220-1 of the molecular attribute filtering step 220, a fifth parameter value 240 for configuringthe fifth parameter 222-1 of the substructure filtering step 222, and an sixth parameter value 242 for configuring the sixth parameter 224-1 of the conversion step 224.
[0087] The data transformation process 208 generates a first curated data set 244 which is provided to an error reduction evaluation process 246. Based on the evaluation of the first curated data set 244 performed by the error reduction evaluation process 246, an adjustment process 248 causes an adjustment to either the plurality of data sets 202, the first curated data set 244, and / or the configuration 228 to generate a second curated data set 250.
[0088] The plurality of data sets 202 are identified from journal and patent data sources (e.g., as described above in relation to the step of identifying 102 in the method 100 shown in Figure 1 ). The plurality of data sets 202 are obtained from all available data sets by filtering the data sets to return only those related to the target of interest — PARP1. In one embodiment, the target of interest is identified within the configuration 228 such that the plurality of data sets 202 related to PARP1 are identified based on the configuration 228. Data sets 204 of the plurality of data sets 202 comprise 4,319 raw data points across 169 data sets extracted from 169 journal articles (e.g., the first data set 204-1 is extracted from a first journal article, the second data set 204-2 is extracted from a second journal article, etc.) and data sets 206 of the plurality of data sets 202 comprise 6,416 raw data points across 128 data sets extracted from 128 patents (e.g., the third data set 206-1 is extracted from a first published patent or patent application, the fourth data set 206-2 is extracted from a second published patent or patent application, etc.). In total, the plurality of data sets 202 consists of 11 ,329 raw data points comprising properties of interest (activity values and physical properties) of 7,247 unique structures (compounds) which are known to interact with and / or modulate PARP1 .
[0089] The configuration 228 is generated (e.g., as described above in relation to the step of generating 104 in the method 100 shown in Figure 1 ) to configure the data transformation process 208. The configuration 228 is generated from a base configuration for the data transformation process 208 which includes default parameter values (e.g., the first plurality of parameter values 230, etc.) and is then manually updated to specify the parameterization of the data transformation process 208 specific for PARP1 data curation. The configuration 228 is a stored as a JSON file containing parameter values for parameters of the data transformation process 208. A representation of the configuration 228 is provided in TABLE I below.“assay_keyword_denylist" : ["cytotox" , "cell based" , "cellbased" , "whole cell" , "cell protection", "comet" , "viability", "mts assay", "muta", "proliferation" , "growth of", "western" , "migration", "flow cytometry" , "survival" , "colony formation" , "transwell", "invasion", "facs" , "ph domain", "pleckstring homology", "extracellular domain", "kinomescan" , "pubchem", "kinobead" ] , “assay_regex_denylist": [“"] ,“comp_chem_keywords": ["docking" , "QSAR" , "comparative molecular field analysis", "CoMFA" , "quantitative structure" , "topological descriptors", "support vector", "retrospective" , "partition", "prediction" ] ,“species" : ["MISSING" , "human", "homo sapiens","HUMAN" ] ,“parameter" : ["Ki", "KI", "pKi" , "IC50" , "pIC50","PIC50", "KD", "Kd", "DC50" ] , “max_RB": 20,“min_MW": 200,“max_MW": 900,“substructure_filters": [ "C1CCC(O1)OC1CCCO1" ,"C1CCCC(O1)OC1CCCO1" , "C1CCC(O1)OCC1CCCO1" , "C1CCCC (01 )OC1COCCC1" , "NCC(=O)NCC(=O)NCC(=O)NCC=O" ] , “units" “NM"}TABLE 1
[0090] The data transformation process 208 contains a sequence of steps which are executable (in turn) to transform the plurality of data sets 202 into a coherent, harmonized, set of data — i.e., the first curated data set 244. Thus, the first curated data set 244 is generated (e.g., as described above in relation to the step of generating 106 in the method 100 shown in Figure 1 ) by executing the data transformation process 208 on the plurality of data sets 202 in accordance with the configuration 228. That is, each step in the data transformation process 208 is executed in sequence to transform the plurality of data sets 202 into the first curated data set 244.
[0091] Each filtering step within the data transformation process 208 generates an intermediate data set from a data set provided as input (i.e., the assay filtering step 212 receives the plurality of data sets 202 as input and generates a first intermediate data set as output and each of theremaining filtering steps generates an intermediate data set from the intermediate data set generated by the previous filtering step).
[0092] The assay filtering step 212 filters the plurality of data sets 202 such that only those data sets which satisfy the assay requirements as defined by the first plurality of parameter values 230 of the first plurality of parameters 212-1 are include within the intermediate data set output by the assay filtering step 212. In general, the fist plurality of parameters 212-1 define an accept list and a deny list of terms which are applied to values related to assay / assay type to include and / or exclude data from the plurality of data sets 202. The first plurality of parameters 212-1 include an assay keyword acceptlist parameter (“assay_keyword_acceptlist”), an assay regular expression acceptlist parameter (“assay_regex_acceptlist”), an assay keyword denylist parameter (“assay_keyword_denylist”), and an assay regular expression denylist parameter (“assay_regex_denylist”). A parameter value assigned to a keyword list is a list of words or terms which, if found within values related to assay / assay type of a data point cause the data point to be excluded or included. A parameter value assigned to a regex list is a list of regular expressions which are executed to perform the filtering step. The parameters and parameter values for the first plurality of parameters 212-1 are shown in TABLE I.
[0093] The source filtering step 214 removes or filters out any raw data points which originate from sources defined by the second parameter value 232 of the second parameter 214-1 . The second parameter value 232 is a keyword list parameter and the second parameter value 232 contains a list of keywords which, if identified within the bibliographic data of a raw data point (e.g., the title, abstract, etc.), cause the raw data point to be excluded from the intermediate data set generated by the source filtering step 214. In the present instance, the source filtering step 214 removes or filters out any raw data points which originate from computational chemistry sources. The parameter and parameter value for the source filtering step 214 are shown as the “comp_chem_keywords” parameter and corresponding parameter value shown in TABLE I.
[0094] The species filtering step 216 removes of filters out any raw data points which do not relate to species defined accordingto the third parameter value 234 of the third parameter 216-1. For example, the species filtering step 216 can be used to remove raw data points related to nonhuman subject experimental data. The parameter and parameter value for the species filtering step 216 are shown as the “species” parameter and corresponding parameter value shown in TABLE I.
[0095] The parameter filtering step 218 is used to define the set of parameters (e.g., IC50, EC50, etc.) of property values to be included within the first curated data set 244. The fourth parameter value 236 of the fourth parameter 218-1 defines the list of parameters to be included such that any raw data points reporting property values according to a parameter included within the fourth parameter value 236 are included within the intermediate data set generated by the parameter filtering step 218 whilst any raw data points reporting property values according to a parameter not included within the fourth parameter value 236 are excluded from the intermediate data set. The “parameter” and corresponding list of parameter values shown in TABLE I correspond to the fourth parameter 218-1 and fourth parameter value 236 of the parameter filtering step 218.
[0096] The molecular attribute filtering step 220 filters or removes any raw data points related to compounds which do not satisfy the attributes defined by the second plurality of parameter values 238 of the second plurality of parameters 220-1 . Here, attributes of a compound include aspects such as molecular weight, number of rotatable bonds, topological polar surface area, and the like. As shown by the values assigned to the “max_RB”, “min_MW”, and “max_MW” shown in TABLE I, the molecular attribute filtering step 220 removes any raw data points having more than 20 rotatable bond and any data points having a molecular weight outside of the range of 200-900.
[0097] The substructure filtering step 222 filters or removes any raw data points related to compounds which have chemical structures which satisfy the substructure filters defined by the fifth parameter value 240 of the fifth parameter 222-1 . The fifth parameter value 240 contains a list of molecular representations (SMILES strings) defining chemical substructures to be excluded. As such, the intermediate data set output by the substructure filtering step 222 comprises raw data points for compounds which do not satisfy the substructure filters defined by the fifth parameter value 240. The “substructure_filters” parameter and corresponding parameter value shown in TABLE I correspond to the fifth parameter value 240 assigned to the fifth parameter 222-1 of the substructure filtering step 222.
[0098] The conversion step 224 converts all property values to the units defined by the sixth parameter value 242 of the sixth parameter 224-1 . As shown in TABLE I, the sixth parameter 224-1 (“units”) is set to the sixth parameter value 242 (“NM”) such that all property values (i.e., binding activity) are converted to NM. The conversion step 224 comprises processing logic to convert between different units (e.g., processing logic to convert from pM or g / ml into nM). In addition, any conversion failures, such as encountering an unknown unit or unconvertible unitpair (e.g., attempting to convert centimeters to liters), are flagged and subsequently identifiable as part of the error reduction evaluation process 246 and / or the adjustment process 248.
[0099] The de-duplication step 226 is a non-parameterized filtering step which filters or removes any duplicate or redundant raw data points. A duplicate or redundant raw data point is a raw data point which shares the same property value, parameter, and molecular representation as another raw data point (i.e., the duplicate raw data point does not provide any new information).
[0100] The first curated data set 244 (generated by executing the data transformation process 208 on the plurality of data sets 202) contains 11 ,329 data points filtered and / or transformed from the 15,172 data points in the plurality of data sets 202. Although harmonized, the first curated data set 244 may nevertheless include data issues which require remediation in order to improve the quality and consistency of the curated data.
[0101] The error reduction evaluation process 246 analyzes the first curated data set 244 to identify potential data issues within the first curated data set 244 (e.g., as described above in relation to the implementing 108 step of the method 100 shown in Figure 1 ). In the example of Figure 2, the error reduction evaluation process 246 includes generating a first distributional representation of the first curated data set 244 to identify potential outlier based data issues (as illustrated in Figure 3A and described in more detail below); generating a second distributional representation of the first curated data set 244 to identify potential source imbalance issues (as illustrated in Figure 3C and described in more detail below); performing a correlation analysis to identify potential parameter comparability issues (as illustrated in Figure 3D and described in more detail below); and performing an exponent comparison to identify potential transcription issues (as illustrated in Figure 3E). The skilled person will appreciate that this list of quantitative analyses is merely illustrative and further quantitative analyses can be performed on the first curated data set 244.
[0102] Figure 3A shows a plot 302 of a distributional representation of the first curated data set 244. The distributional representation shown in Figure 3A is a histogram of activity values within the first curated data set 244. That is, the histogram is generated from the property values (binding activity) within the first curated data set 244. Based on the distributional representation, an outlier portion 304 can be identified which indicates a potential data issue. The outlier portion 304 corresponds to a set of data points within the first curated data set 244 which fall outside of the distribution of activity values. This is further shown in the plot 306 of Figure 3B whichshows a set of outlier data points 308 which correspond to the data points within the outlier portion 304 of the distributional representation shown in Figure 3A.
[0103] Figure 3C shows a plot 310 of another distributional representation of the first curated data set 244. The distributional representation shown in Figure 3C is a plot of the percentage contribution of each data set (e.g., the first data set 204-1 , the second data set 204-2, etc.) to the first curated data set 244. For example, if a data set has a percentage contribution of 25%, then a quarter of all data points within the first curated data set 244 originate from that data set. An imbalance data issue (or contribution imbalance data issue) can be identified from the distributional representation (i.e., from the percentage contribution of each data set) by identifying any data sets which contribute more than a threshold contribution amount 312 to the first curated data set 244. In the example shown in Figure 3C, the threshold contribution amount 312 is set as 5% such that a set of data sets 314 are identified as causing a potential contribution imbalance.
[0104] Figure 3D shows a correlation analysis in the form of a parameter comparison plot 316 of pairs of data points within the first curated data set 244 where a compound has multiple experimental outcomes. The parameter comparison plot 316 shows the activity value of a compound taken from a first data set / source (A) and the activity value of the compound taken from a second data set / source (B). The parameter comparison plot 316 shows the correlation or comparability of activity values reported using different parameters (e.g., EC50, IC50, etc.). As can be seen from the top-left of the parameter comparison plot 316, there is a high degree of noise, as represented by the number of off-diagonal points within the subplots, for EC50 and IC50 indicatingthat properties reported in EC50 and IC50 are not comparable. In contrast, the number of off-diagonal points in the sub-plots associated with Kd and Ki (bottom-right of the parameter comparison plot 316) indicate that these data points are comparable even across different data sets / sources.
[0105] Figure 3E shows a plot 318 of exponent issues within the first curated data set 244. An exponent issue occurs when two values for the same compound are identical, but the exponent is out by an order of magnitude (e.g., the units have been incorrectly assigned such as reporting an activity in nM when it was actually measured in pM). The plot 318 is a log-log plot of property value pairs such that a compound within the first curated data set 244 having two identical activity values with the same exponent would appear on the diagonal of the plot 318, whereas mismatched exponents would appear as off-diagonal points. Figure 3E shows two groups of data points within the first curated data set 244 which have mismatched exponents — a first group 320where the exponent is out by a factor of 3 (log difference) and a second group 322 where the exponent is out by a factor of 6.
[0106] In summary, four data issues were identified based on the error reduction evaluation process 246 of the first curated data set 244: an outlier data issue (as illustrated in Figure 3A); a source imbalance issue (as illustrated in Figure 3C); a parameter comparability issue (as illustrated in Figure 3D); and a transcription issue (as illustrated in Figure 3E).
[0107] Referring once again to Figure 2, the adjustment process 248 applies adjustments to reduce or eliminate the data issues (e.g., the implementing 112 step of the method 100 shown in Figure 1 ). To reduce or eliminate the outlier data issue, the adjustment process 248 applies an adjustment to remove the outlier data points (e.g., data points within the outlier portion 304 shown in Figure 3A) from the first curated data set 244 and any further subsequently generated curated data sets (e.g., the second curated data set 250). To reduce or eliminate the source imbalance issue, the adjustment process 248 applies an adjustment to apply a weighting to the data points originating from the over-contributing data sets / sources (e.g., the set of data sets 314 shown in Figure 3C) to reduce their contribution to the curated data set. In an alternative implementation, the adjustment could subsample the data points originating from the over-contributing data sets / sources such that only a set number of data points from each of the over-contributing data sets / sources are included in the curated data set. To reduce or eliminate the parameter comparability issue, the adjustment process 248 applies an adjustment to the fourth parameter value 236 to limit the parameters being identified from the plurality of data sets 202 to being Kd and Ki only. That is, the value of the “parameter” parameter shown in TABLE I is updated to the list “["Ki", "KD", "Ki", "KI"]”. A subsequent execution of the data transformation process 208 would thus result in only Kd and Ki values being included in the subsequently created curated data set. To reduce or eliminate the transcription issue, the adjustment process 248 applies an adjustment to remove the data points having mismatched exponents (e.g., data points within the first group 320 and the second group 322 shown in Figure 3) from the first curated data set 244 and any further subsequently generated curated data sets (e.g., the second curated data set 250).
[0108] After the above-described adjustments have been made, and the data transformation process 208 is executed again with the updated configuration 228, the second curated data set 250 comprises harmonized data points which can be subsequently used for several downstream tasks such as generating training data sets for predictive model training and the like.
[0109] Figure 4 shows a portion of a curated data set generated from the example data curation process shown in Figure 2. Figure 4 shows three rows from the second curated data set 250 of Figure 2.
[0110] Figure s shows an example computing system for carrying out the methods of the present disclosure. Specifically, Figure 5 shows a block diagram of an embodiment of a computing system according to example embodiments of the present disclosure. The computing system shown in Figure 5 may correspond to a part, or the whole, of any of the functional units or modules described above.
[0111] Computing system 500 can be configured to perform any of the operations disclosed herein such as, for example, any of the operations discussed with reference to the method steps described in relation to Figures 1A-1 C and / or the functional units described in relation to Figure 2. Computing system includes one or more computing device(s) 502. The one or more computing device(s) 502 of computing system 500 comprise one or more processors 504 and memory 506. One or more processors 504 can be any general purpose processor(s) configured to execute a set of instructions, such as computing instructions including implemented in any one or more programming languages such as Python, Go, C, C++, C#, Java, or the like. For example, one or more processors 504 can be one or more general-purpose processors, one or more field programmable gate array (FPGA), and / or one or more application specific integrated circuits (ASIC). In one embodiment, one or more processors 504 include one processor. Alternatively, one or more processors 504 include a plurality of processors that are operatively connected. One or more processors 504 are communicatively coupled to memory 506 via address bus 508, control bus 510, and data bus 512. Memory 506 can be a random access memory (RAM), a read only memory (ROM), a persistent storage device such as a hard drive, an erasable programmable read only memory (EPROM), and / or the like. The one or more computing device(s) 502 further comprise I / O interface 514 communicatively coupled to address bus 508, control bus 510, and data bus 512.
[0112] Memory 506 can store information that can be accessed by one or more processors 504. For instance, memory 506 (e.g., one or more non-transitory computer-readable storage mediums, memory devices) can include computer-readable instructions (not shown) (e.g., computing instructions) that can be executed by one or more processors 504. The computer-readable instructions can be software written in any suitable programming language or can be implemented in hardware. Additionally, or alternatively, the computer-readable instructions can be executed inlogically and / or virtually separate threads on one or more processors 504. For example, memory 506 can store instructions (not shown), such as computing instructions, that when executed by one or more processors 504 cause one or more processors 504 to perform operations such as any of the operations and functions for which computing system 500 is configured, as described herein. In addition, or alternatively, memory 506 can store data (not shown) that can be obtained, received, accessed, written, manipulated, created, and / or stored. The data can include, for instance, the data and / or information described herein in relation to Figures 1 -4. In some implementations, the one or more computing device(s) 502 can obtain from and / or store data in one or more memory device(s) that are remote from the computing system 500.
[0113] Computing system 500 further comprises storage unit 516, network interface 518, input controller 520, and output controller 522. Storage unit 516, network interface 518, input controller 520, and output controller 522 are communicatively coupled to the central control unit (i.e., the memory 506, the address bus 508, the control bus 510, and the data bus 512) via I / O interface 514.
[0114] Storage unit 516 is a computer readable medium, preferably a non-transitory computer readable medium, comprising one or more programs, the one or more programs comprising computing instructions which when executed by the one or more processors 504 cause computing system 500 to perform the method steps of the present disclosure. Alternatively, storage unit 516 is a transitory computer readable medium. Storage unit 516 can be a persistent storage device such as a hard drive, a cloud storage device, or any other appropriate storage device.
[0115] Network interface 518 can be a Wi-Fi module, a network interface card, a Bluetooth module, and / or any other suitable wired or wireless communication device. In an embodiment, network interface 518 is configured to connect to a network such as a local area network (LAN), or a wide area network (WAN), the Internet, or an intranet.
[0116] The above illustrative examples of various aspects and implementations provide an overview for understanding aspects and implementation of the disclosed method. The figures provided herein depict exemplary aspects of the present system and methods and are not intended to limit the scope of the disclosure.
[0117] Unless otherwise stated, all technical terms used herein have the same meaning as commonly understand by a person skilled in the art. Singular forms “a”, “an” and “the” include plural references unless the context of the disclosure clearly dictates otherwise. The term “or” is intended to encompass “and / or” unless clearly stated otherwise.
[0118] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.
Claims
Attorney Docket: 33719-70815SYSTEMS AND METHODS FOR CURATION OF DIVERSE BIOLOGICAL AND / OR CHEMICAL DATACLAIMSWhat is claimed is:
1. A method for curation of diverse biological and / or chemical data, the method comprising: identifying, by one or more processors, a plurality of data sets derived from one or more data sources, each of the plurality of data sets comprising one or more raw data points related to properties of biological and / or chemical compounds of interest; generating, by the one or more processors, a first configuration for a data transformation process comprising a sequence of data transformation steps, the first configuration comprising one or more parameter values for configuring one or more steps of the sequence of data transformation steps; generating, bythe one or more processors, a first curated data set by executing the data transformation process on the plurality of data sets in accordance with the first configuration, wherein the first curated data set comprises a plurality of harmonized data points related to properties of the biological and / or chemical compounds of interest; implementing, by the one or more processors, an error reduction evaluation of the plurality of harmonized data points of the first curated data set; analyzing, by the one or more processors, output of the error reduction evaluation to detect that a data issue exists in relation to one or more data points within the plurality of data sets and / or the first curated data set; and upon detection that the data issue exists, implementing, by the one or more processors, at least one adjustment to the one or more data points and / or the first configuration to reduce or eliminate the data issue to enhance a downstream process that inputs the enhanced first curated data set as a data input.
2. The method of claim 1 further comprising:training, by the one or more processors, a prediction model on a training data set generated from first curated data set, wherein the prediction model trained to predict a property value for an input molecular representation.
3. The method of claim 1 further comprising: generating, by the one or more processors, a report comprising a representation of the first curated data set.
4. The method of claim 3 wherein the report comprises at least one visualization based on the error reduction evaluation.
5. The method of claim 3 further comprising: outputting, by the one or more processors, the report for display to a user.
6. The method of claim 1 wherein the data issue is an abnormality issue identified by the one or more data points falling outside of a range of values.
7. The method of claim 1 wherein the data issue is an outlier issue identified by the one or more data points being outliers with respect to a representative value of the plurality of harmonized data points.
8. The method of claim 1 wherein the data issue is a transcription issue identified from the one or more data points.
9. The method of claim 1 wherein the at least one adjustment comprises a transformation of the one or more data points within the plurality of harmonized data points of the first curated data set.
10. The method of claim 1 wherein the at least one adjustment comprises removing the one or more data points from the plurality of harmonized data points.11 . The method of claim 1 wherein the at least one adjustment comprises removing data points corresponding to the one or more data points from the plurality of data sets.
12. The method of claim 1 wherein adjusting the one or more data points comprises applying a weighting to the one or more data points, wherein the weighting is based on the error reduction evaluation.
13. The method of claim 1 wherein the at least one adjustment comprises generating a second configuration for the data transformation process, the second configuration comprising an instruction to exclude or weight the one or more data points from a second curated data set generated from the second configuration.
14. The method of claim 1 wherein the error reduction evaluation comprises one or more of: a calculation of a correlation; a determination of a range of values; a statistical analysis; a generation of a distributional representation; a calculation of a confidence score; and / or a clustering analysis.
15. The method of claim 1 wherein the one or more data sources include: experimental biological and / or chemical data sources, publicly and / or commercially available databases, scientific publications, clinical trial results, and / or patent publications.
16. The method of claim 1 wherein the properties included within the one or more raw data points include one or more of a physiological activity; a pharmacological activity including a pharmacokinetic property and / or a pharmacodynamic property; a toxicity or selectivity; a reactivity; a binding affinity.
17. The method of claim 1 wherein the sequence of data transformation steps include one or more of: a filtering step; a conversion step; a weighting step; or a transformation step.
18. The method of claim 1 wherein the one or more raw data points are related to properties of compounds which are known or predicted to interact with and / or modulate at least one target or at least two targets such that the plurality of harmonizeddata points relate to properties of compounds which are known or predicted to interact with and / or modulate the at least one target or at least two targets.
19. A non-transitory computer readable medium comprising instructions which, when executed by one or more processors, cause the one or more processors to: identify a plurality of data sets derived from one or more data sources, each of the plurality of data sets comprising one or more raw data points related to properties of biological and / or chemical compounds of interest; generate a first configuration for a data transformation process comprising a sequence of data transformation steps, the first configuration comprising one or more parameter values for configuring one or more steps of the sequence of data transformation steps; generate a first curated data set by executing the data transformation process on the plurality of data sets in accordance with the first configuration, wherein the first curated data set comprises a plurality of harmonized data points related to properties of the biological and / or chemical compounds of interest; implement an error reduction evaluation of the plurality of harmonized data points of the first curated data set; analyze output of the error reduction evaluation to detect that a data issue exists in relation to one or more data points within the plurality of data sets and / or the first curated data set; and upon detection that the data issue exists, implement at least one adjustment to the one or more data points and / or the first configuration to reduce or eliminate the data issue to enhance a downstream process that inputs the enhanced first curated data set as a data input.
20. A device comprising a memory storing instructions which, when executed by one or more processors, cause the device to:identify a plurality of data sets derived from one or more data sources, each of the plurality of data sets comprising one or more raw data points related to properties of biological and / or chemical compounds of interest; generate a first configuration for a data transformation process comprising a sequence of data transformation steps, the first configuration comprising one or more parameter values for configuring one or more steps of the sequence of data transformation steps; generate a first curated data set by executing the data transformation process on the plurality of data sets in accordance with the first configuration, wherein the first curated data set comprises a plurality of harmonized data points related to properties of the biological and / or chemical compounds of interest; implement an error reduction evaluation of the plurality of harmonized data points of the first curated data set; analyze output of the error reduction evaluation to detect that a data issue exists in relation to one or more data points within the plurality of data sets and / or the first curated data set; and upon detection that the data issue exists, implement at least one adjustment to the one or more data points and / or the first configuration to reduce or eliminate the data issue to enhance a downstream process that inputs the enhanced first curated data set as a data input.
Citation Information
Patent Citations
Visualization and manipulation of biomolecular relationships using graph operators
US20020087275A1
Recommending novel reactants to synthesize chemical products
US20180096100A1
Chemical structure generating device, chemical structure generating program, and chemical structure generating method
US20220172803A1
Systems and methods for predicting outcomes and conditions of chemical reactions with high reliability based on a highly diverse and accurate dataset
US20230131234A1