A software system for online analysis of DNA-encoded compound library screening data visualization
By providing an online analysis software system for visualizing DNA-encoded compound library screening data, the problem of low efficiency of existing tools under large data volumes has been solved. It enables autonomous and customized DEL data analysis, improves analysis efficiency, and promotes early drug development.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- WUXI APPTEC (SHANGHAI) CO LTD
- Filing Date
- 2022-12-27
- Publication Date
- 2026-04-28
AI Technical Summary
Existing DNA-encoded compound library screening data analysis tools cannot meet the needs for efficient, autonomous, and customized analysis, especially in the case of large amounts of data, and are inefficient and cannot meet the complex analytical requirements of the pharmaceutical field.
This software system provides an online visualization and analysis platform for screening data from a DNA-encoded compound library. It includes an analysis logic setting module, a screening data analysis module, a DEL full-length molecule and partial structure display module, a drug-likeness prediction module, and a structural similarity calculation module. It supports custom analysis logic and data filtering, and provides intuitive structural display and preliminary drug-likeness and similarity analysis.
It enables efficient analysis of large datasets, supports autonomous and customized DEL data analysis, improves analysis efficiency, and helps advance early-stage drug development.
Smart Images

Figure CN116189810B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data analysis, and in particular to a software system for online visualization and analysis of DNA-encoded compound library screening data. Background Technology
[0002] DNA-encoded chemical library (DEL) technology is a high-throughput small molecule drug screening platform based on affinity. It rapidly screens up to tens of billions of small molecule compounds to discover those that bind to biological targets, aiding in the discovery, validation, and optimization of lead compounds. Compared to traditional high-throughput screening (HTS), DEL offers advantages such as higher throughput and shorter screening cycles. The construction of the DNA compound library utilizes the principles of combinatorial chemistry. Based on the "Split & Pool" method, multiple reaction cycles allow for geometric growth of the chemical space, constructing a collection of molecules with diverse structures ranging from tens of billions to trillions. A schematic diagram of the typical DEL screening and data analysis workflow is shown below. Figure 1 As shown.
[0003] When using DEL for screening, the DEL is incubated with the target protein. Based on the differences in affinity of molecules for the target protein, compounds with weak affinity are eluted and removed, and the compounds retained on the target protein are detected. Since there is a correspondence between compounds and their DNA coding information, high-throughput sequencing technology is used to decode the signal DNA tag to obtain the corresponding compound information. Data analysis is then used to identify signals with high intensity and low noise levels for experimental verification. Small molecules without DNA tags are then resynthesized and their activity is subsequently verified, thus yielding lead compounds, such as… Figure 1 As shown.
[0004] Conventional DEL data analysis includes filtering data according to analysis logic to reduce false positive data; using line charts, pie charts, etc. to find frequently occurring structural fragments; using 3D scatter plots to find high-confidence "point, line, and surface" features and performing SAR structure-activity relationship analysis. At the same time, combined with chemical information-related analysis, including structural similarity, physicochemical properties, etc., to select promising compounds with "drug-like" and "novel" characteristics. GlaxoSmithKline first used a DEL data analysis method in the literature to analyze "point, line, and surface" features through 3D scatter plots, and then explored the structure-activity relationship within the family (Reference: ClarkMA, AcharyaRA, Arico-MuendelCC, et al. Design, synthesis and selection of DNA-encoded small-molecule libraries[J]. Nature chemical biology, 2009, 5(9): 647-654). Subsequently, data analysis methods were enriched by various researchers. For example, AstraZeneca (References: Goodnow, RA; Dumerin, CE; Keefe, ADDNA-Encoded Chemistry: Enabling the Deeper Sampling of Chemical Space. Nat Rev Drug Discov. 2017 Feb; 16(2):131-147.) summarized some data analysis and visualization methods used in the DEL screening platform in a review, such as 3D bar charts and 3D scatter plots, to explore the relationship between screened molecules and screening results (such as enrichment and copy number).
[0005] Conventional affinity screening experiments yield millions of signals from molecular datasets ranging from hundreds of millions to tens of billions. How to fully utilize the information contained in these signals has become a key challenge for DEL (Drug-Electrical Linkage) technology. Increasingly, scientists are incorporating automated processes and intelligent analysis methods (including cheminformatics and machine learning) into data analysis: combining experimental results and chemical information from massive signal volumes for initial filtering, narrowing the signal count down to a few thousand, before allowing medicinal chemists to perform more refined analysis. This analytical framework fully leverages DEL's high-throughput advantage, making the use of "big data" possible. By rationally planning and utilizing the screening data, the potential information can be fully explored, supporting internal technology development and making the analytical methods more "intelligent."
[0006] In data analysis workflows, the volume of data analyzed is in the tens of thousands, requiring specialized BI tools (Business Intelligence Analytics) to assist in the analysis. Currently, Spotfire is the main BI tool used in DEL scenarios. Compared to other analysis tools, it is often used in scientific research scenarios, supporting highly interactive charts for large datasets; however, it requires a high price, lacks computational plugins for the pharmaceutical field, and only provides basic chemical information analysis, failing to meet the needs of drug-likeness, similarity, structure-enrichment relationship analysis, etc., thus not meeting daily analysis needs. When loading large amounts of data, the interactive charting process experiences significant lag, reducing analysis efficiency. Therefore, a few companies are leveraging Spotfire to further develop and build automated analysis platforms; that is, generating important charts with one click while loading data. For example, Eli Lilly (Reference: J, Román JP, Jessop TC, et al. Design and development of a technology platform for DNA-encoded library production and affinity selection[J].SLAS DISCOVERY: Advancing the Science of Drug Discovery, 2018, 23(5):387-396) reported an internal automation platform based on DEL, encompassing modules such as DEL design, production, and data analysis. The data analysis module provides methods for generating point "line and surface" features, as well as automatically generating 3D scatter plots for visualization analysis. However, this module is built on Spotfire, and the informatics components are not embedded in the analysis process. Summary of the Invention
[0007] The technical problem to be solved by this invention is to provide an efficient DEL screening data visualization and analysis tool, which enables DEL screening experimenters to conduct independent and customized DEL data analysis, and facilitates DEL screening experimenters to independently analyze DEL screening data and search for lead compounds.
[0008] To address the aforementioned technical problems, the present invention provides a software system for online visualization and analysis of DNA-encoded compound library screening data, comprising:
[0009] The analysis logic setting module is used to define the data analysis logic;
[0010] The data analysis module performs statistical analysis on the experimental data based on the data analysis logic, and its analysis process and screening results are based on DNA tag information.
[0011] The DEL full-length molecule and partial structure display module presents the screening results from the screening data analysis module in the form of DEL molecular chemical structures to explore the relationship between DEL molecular chemical structures and screening results. Here, a full-length molecule refers to the actual chemical structure obtained through a complete chemical reaction pathway in library production, where each cycle's building blocks are real. A partial structure refers to a series of full-length molecules containing the same building blocks in one or more cycles, thus sharing a common structure; the partial structure represents the structural pattern of a series of full-length molecules.
[0012] Furthermore, the software system for online visualization and analysis of DNA-encoded compound library screening data also includes:
[0013] The drug-likeness prediction module evaluates the drug-likeness of compounds obtained from screening rules based on certain rules.
[0014] The structural similarity calculation module is used to evaluate the similarity of compounds obtained according to screening rules.
[0015] Furthermore, the analysis logic setting module includes:
[0016] The analysis logic setting component dynamically generates data filters based on the experimental information of the screening experiments, allowing users to customize the analysis logic.
[0017] An analysis logic database can be used to add and / or delete filtered experimental data obtained by analysis logic filtering.
[0018] Data filtered by user-defined analysis logic is stored in the analysis logic database.
[0019] Furthermore, the analysis logic setting component includes:
[0020] The filter information table is used to display the filter settings and related information;
[0021] The logic is simple, and it uses a data filter to set filtering conditions.
[0022] Complex logic settings are used to process filtering conditions obtained from simple logic settings through AND and / or OR;
[0023] Analyze the logic table, which is used to display the logic set by the user.
[0024] Furthermore, the data filtering and analysis module includes:
[0025] The data overview module shows the signal distribution of full-length molecules and partial structures in different libraries under user-defined logic;
[0026] The single-database filtering dashboard module is used to perform statistics and analysis on the data filtered under the analysis logic.
[0027] Furthermore, the DEL full-length molecule and partial structure display module includes:
[0028] The data labeling and display module is used to display molecular information in charts when linked, and allows users to label data;
[0029] A tagged molecule database is used to display data information tagged by users and allows users to add and / or delete tagged data;
[0030] A structure-enrichment relationship exploration table is used to display labeled molecules in the labeled molecule database and allows for the capture of full-length molecules contained in partial structures.
[0031] Furthermore, the drug-like property prediction module includes:
[0032] The Five Rules Analysis module is used to calculate the number of molecules that violate and / or conform to the Five Rules;
[0033] The three-rule analysis module is used to calculate the number of molecules that violate and / or conform to the three rules;
[0034] The MPO analysis module is used to calculate the MPO of molecules.
[0035] Furthermore, the structural similarity calculation module includes:
[0036] Cluster analysis module, which is used for structural cluster analysis of molecules;
[0037] The similarity matrix analysis module is used to calculate the similarity matrix between two groups of molecules;
[0038] The substructure search module is used to search for user-defined structural fragments.
[0039] Furthermore, the clustering analysis module employs the Butina clustering method, K-means clustering method, or K-medoids clustering method for molecular structure analysis.
[0040] This invention provides a software system for online visualization and analysis of DNA-encoded compound library (DEL) screening data. After importing experimental data, users can set logic to analyze the data based on experimental information and actual needs. It supports importing data volumes of up to hundreds of millions for a single project and is specifically designed for DEL, allowing for flexible and varied calculation plugins. It operates smoothly without lag and meets the requirements of DEL data analysis. This design also provides intuitive structural displays and allows users to annotate data according to actual needs, facilitating independent and customized DEL data analysis by DEL screening researchers. Furthermore, this invention assists users in conducting preliminary drug-likeness and similarity analysis, enabling DEL screening researchers to independently analyze DEL screening data and identify lead compounds, thus contributing to the advancement of early drug development. This invention achieves visualization and intelligentization of the DEL analysis process, providing DEL screening data analysts with a reliable data analysis software tool. Attached Figure Description
[0041] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0042] Figure 1 This is a schematic diagram of a standard DEL filtering and data analysis process;
[0043] Figure 2 This is a schematic diagram of the structure of the software system for online visualization and analysis of DEL-filtered data according to the present invention;
[0044] Figure 3 This is a schematic diagram of obtaining the covalent inhibitor 1e of 3CLpro using the present invention;
[0045] Figure 4 This is a schematic diagram of the analysis logic design module interface during the process of obtaining the covalent inhibitor 1e of 3CLpro using this invention; wherein, the "Visualization of selection signals" column is the data overview module of the data screening and analysis module;
[0046] Figure 4a for Figure 4 The Selection Condition Information section;
[0047] Figure 4b for Figure 4 Middle Atom Logic column;
[0048] Figure 4c for Figure 4 The Combined Logic column in the middle;
[0049] Figure 4d for Figure 4 Middle Anlaysis Logic Table column;
[0050] Figure 4e for Figure 4 The "Visualization of selection signals" section is in the middle.
[0051] Figure 5 During the process of obtaining the covalent inhibitor 1e of 3CLpro using this invention, click Figure 4 After selecting Library X, the interface for the data filtering and analysis module is shown in the diagram.
[0052] Figure 5a for Figure 5 The relevant data display and calculation components;
[0053] Figure 5b for Figure 5 Pie chart in the middle;
[0054] Figure 5c for Figure 5 Line chart in the middle;
[0055] Figure 5d for Figure 5 A two-dimensional scatter plot in the image;
[0056] Figure 5e for Figure 5 A 3D scatter plot in the image;
[0057] Figure 6 In the process of obtaining the covalent inhibitor 1e of 3CLpro using this invention, the following applies: Figure 5 A schematic diagram of the interface for enriching signals using data filters;
[0058] Figure 6a for Figure 6 Data filters in the middle;
[0059] Figure 7 This is a schematic diagram of the interface of a structural segment with good enrichment signal performance, which is intelligently displayed by a two-dimensional scatter plot during the process of obtaining the covalent inhibitor 1e of 3CLpro using the present invention.
[0060] Figure 8 In the process of obtaining the covalent inhibitor 1e of 3CLpro using the present invention, the information and screening performance of Series 1 are displayed in a table. Detailed Implementation
[0061] The technical solutions of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0062] Example 1
[0063] Figure 2 A schematic diagram of a software system for online visualization and analysis of DNA-encoded compound library screening data according to the present invention is shown, including:
[0064] The analysis logic setting module is used to define the data analysis logic; the analysis logic set in this module is usually set according to the filtering needs.
[0065] The data analysis module performs statistical analysis on the experimental data based on the data analysis logic, and its analysis process and screening results are based on DNA tag information; this module performs statistical analysis on the data based on DNA tag information.
[0066] The DEL full-length molecule and partial structure display module presents the screening results of the screening data analysis module in the form of DEL molecular chemical structures to explore the relationship between DEL molecular chemical structures and screening results. This module translates the DNA tag information screened by the screening data analysis module into molecular structure information for display. This module allows users to operate according to their needs, such as marking or unmarking partial structures.
[0067] The software system for online visualization and analysis of DNA-encoded compound library screening data of the present invention also includes:
[0068] The drug-likeness prediction module evaluates the drug-likeness of compound molecules obtained according to certain rules based on the screening rules. The purpose of this module is to conduct a preliminary structural drug-likeness assessment of the structures displayed by the DEL full-length molecule and partial structure display module according to the corresponding rules, such as the five rules or the three rules.
[0069] The structural similarity calculation module is used to evaluate the similarity of compounds obtained according to screening rules.
[0070] In one specific implementation, the analysis logic setting module includes an analysis logic setting component and an analysis logic database. The analysis logic setting component dynamically generates data filters based on the experimental information of the screening experiments, allowing users to customize analysis logic. The data obtained from the customized analysis logic filtering is stored in the analysis logic database. The data filters are generated directly by the system when importing experimental information, based on the results of the screening experiments, reading the copy data and enrichment value of each experiment, and calculating the maximum and minimum values of these two values. The analysis logic setting component includes four parts: a screening information table, simple logic settings, complex logic settings, and an analysis logic table. The screening conditions and related information are displayed in the screening information table. Users perform simple logic settings in the screening information table using data filters (selecting copy number and enrichment), and process the simple logic settings using AND and / or OR to obtain filtering conditions to complete complex logic settings. The completed analysis logic is displayed in the analysis logic table. However, there is a limit to the number of analysis logics a single user can generate for a single project, and there is also a limit to the amount of data for a single analysis logic.
[0071] In one specific implementation, the data analysis module includes a data overview module and a single-library data analysis dashboard module. The data overview module displays the signal distribution of full-length molecules and partial structures in different libraries using bar charts under user-defined logic. Users can also name their custom analysis logic for subsequent analysis. The single-library data analysis dashboard module is used for statistical analysis of the data obtained under the analysis logic. It includes data filters, line charts, pie charts, two-dimensional scatter plots, three-dimensional scatter plots, and related data display and calculation components. The functions of each component are as follows:
[0072] Data filters allow users to adjust the range of data;
[0073] Line graphs are used to show the enrichment performance of molecules under different screening conditions;
[0074] Pie charts are used to show the distribution of building blocks;
[0075] Two-dimensional scatter plots are used to intelligently display structural segments of strong signals, facilitating subsequent analysis by users.
[0076] Three-dimensional scatter plots can show the enrichment of full-length molecules, which can be used by users to analyze "point, line and surface" features;
[0077] Among these components, the two-dimensional scatter plot or pie chart will display the corresponding data information, while the line chart will display the current corresponding data, and other charts will highlight the corresponding areas; clicking on the line chart will display the corresponding data information, and other charts will highlight the corresponding areas.
[0078] Additionally, it should be noted that the statistical analysis of the screened data includes, but is not limited to, project information, single-library information, the name of the analysis logic, the number of signals and linkage information for full-length molecules and partial structures.
[0079] In one specific implementation, the DEL full-length molecule and partial structure display module includes a data labeling display module, a labeled molecule database, and a structure-enrichment relationship exploration table, wherein,
[0080] The data labeling and display module is used to display molecular information in charts when linked, and also allows users to label data;
[0081] The tagged molecule database is used to display data information of user-tagged molecules and allows users to add and / or delete tagged data;
[0082] The structure-enrichment relationship exploration table is used to display the labeled molecules in the labeled molecule database. Clicking or selecting a partial structure will display the full-length molecule contained in that part of the structure and the screening experimental signals under various conditions. That is, the structure-activity relationship analysis commonly used in medicinal chemistry is applied to this application to discuss the relationship between changes in the structure of the building blocks and changes in the intensity of the DEL screening signal.
[0083] In one specific implementation, the drug-likeness prediction module includes a five-rule analysis module, a three-rule analysis module, and an MPO analysis module. The five-rule analysis module calculates the number of molecules that violate and / or conform to the five rules; the three-rule analysis module calculates the number of molecules that violate and / or conform to the three rules; and the MPO analysis module calculates the MPO of the molecules. These modules assist users in evaluating the drug-likeness of molecules. It should be noted that the input data for the drug-likeness prediction module is not limited to the dataset corresponding to the analysis logic database, but also includes datasets corresponding to user-labeled molecules and external datasets, further expanding the application scenarios of this invention.
[0084] In one specific implementation, the structural similarity calculation module includes a clustering analysis module, a similarity matrix analysis module, and a substructure search module. The clustering analysis module uses Butina clustering, K-means clustering, or K-medoids clustering methods to perform clustering analysis on the molecular structures. The similarity matrix analysis module calculates the similarity matrix between two groups of molecules. The substructure search module searches for user-defined structural fragments. It should be noted that the input data for the structural similarity calculation module is not limited to the dataset corresponding to the analysis logic database, but also includes datasets corresponding to user-labeled molecules and external datasets, further expanding the application scenarios of this invention.
[0085] Example 2
[0086] This embodiment demonstrates the application of a software system for online visualization and analysis of DNA-encoded compound library screening data in the screening of 3CLpro covalent inhibitors. Using DEL technology, two series of compounds were identified that can interact with SARS-CoV-2M. pro Or 3CLpro undergoes irreversible binding (Reference: Ge R, Shen Z, Yin J, et al. Discovery of SARS-CoV-2 main protease covalent inhibitors from a DNA-encoded library selection[J].SLASDiscovery,2022,27(2):79-85). The following uses one of the series of compounds, Series 1, as an example to illustrate how the system of this invention obtains the linear feature Series 1, and the full-length molecules contained in Series 1. Figure 3 A schematic diagram of the process of using the system of the present invention is shown, as well as the full-length molecules contained therein, such as full-length molecule 1e.
[0087] Initially, when importing experimental data into the system, first enter the analysis logic setting module, such as... Figure 4 As shown, SelectionCondition Information is the selection information table, Atom Logic in Analysis Logic Creation is the simple logic setting, Combined Logic in Analysis Logic Creation is the complex logic setting, and AnalysisLogic Table is the analysis logic table; for clarity, Figures 4a-4d Showing it separately again Figure 4 The various parts are described below. In this embodiment, the analysis logic set according to the analysis target is detailed in the "Target binders" row of the Analysis Logic Table. Based on the established analysis logic, the "Visualization of selection signals" section of the data filtering analysis module is accessed. This section displays the distribution of full-length molecules and partial structures under the current analysis logic. Click... Figure 4 Library X, rich in semaphores, enters the single-database filtering dashboard module.
[0088] Because the signal is too large, data filters are used to further enrich the signal, such as... Figure 6 As shown, the results after further enrichment are as follows: Figure 7As shown, in the Structure Pattern Discovered Across Data (2D scatter plot), the structure segments with better performance are intelligently displayed. In the analysis of "point, line, and surface" features, the two points in the 2D scatter plot correspond to the line features in the 3D scatter plot The Cube, which are highlighted in red. These line features correspond to Series 1 in the reference.
[0089] After labeling it as Series1, further analysis can be performed in the DEL full-length molecule and partial structure display module, such as... Figure 8 As shown in the table, the information and screening performance of Series 1 are displayed. Clicking on SER ANALYSIS will expand the full-length molecules included and related information such as the enrichment of screening experiments, which is used to explore structure-enrichment relationships (only 1e molecules are displayed).
[0090] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A software system for online visualization and analysis of DNA-encoded compound library screening data, characterized in that, include: The analysis logic setting module is used to define the data analysis logic; The screening data analysis module performs statistical analysis on experimental data based on the data analysis logic. Its analysis process and screening results are based on DNA tag information. The screening data analysis module includes a data overview module and a single-library screening data dashboard module. The data overview module displays the signal quantity distribution of full-length molecules and partial structures in different libraries under user-defined logic. The single-library screening data dashboard module performs statistical analysis on the data obtained from screening under the analysis logic. The DEL full-length molecule and partial structure display module presents the screening results of the screening data analysis module in the form of DEL molecular chemical structure, so as to explore the relationship between DEL molecular chemical structure and screening results. The drug-likeness prediction module evaluates the drug-likeness of compounds obtained from screening rules based on certain rules. The structural similarity calculation module is used to evaluate the similarity of compounds obtained according to screening rules.
2. The software system for online visualization and analysis of DNA-encoded compound library screening data as described in claim 1, characterized in that, The analysis logic setting module includes: The analysis logic setting component dynamically generates data filters based on the experimental information of the screening experiments, allowing users to customize the analysis logic. An analysis logic database can be used to add and / or delete filtered experimental data obtained by analysis logic filtering. Data filtered by user-defined analysis logic is stored in the analysis logic database.
3. The software system for online visualization and analysis of DNA-encoded compound library screening data as described in claim 2, characterized in that, The analysis logic setting component includes: The filter information table is used to display the filter settings and related information; The logic is simple, and it uses a data filter to set filtering conditions. Complex logic settings are used to process filtering conditions obtained from simple logic settings through AND and / or OR; Analyze the logic table, which is used to display the logic set by the user.
4. The software system for online visualization and analysis of DNA-encoded compound library screening data as described in claim 1, characterized in that, The DEL full-length molecule and partial structure display module includes: The data labeling and display module is used to display molecular information in charts when linked, and allows users to label data; A tagged molecule database is used to display data information tagged by users and allows users to add and / or delete tagged data; A structure-enrichment relationship exploration table is used to display labeled molecules in the labeled molecule database and allows for the capture of full-length molecules contained in partial structures.
5. The software system for online visualization and analysis of DNA-encoded compound library screening data as described in claim 1, characterized in that, The drug-like property prediction module includes: The Five Rules Analysis module is used to calculate the number of molecules that violate and / or conform to the Five Rules; The three-rule analysis module is used to calculate the number of molecules that violate and / or conform to the three rules; The MPO analysis module is used to calculate the MPO of molecules.
6. The software system for online visualization and analysis of DNA-encoded compound library screening data as described in claim 1, characterized in that, The structural similarity calculation module includes: Cluster analysis module, which is used for structural cluster analysis of molecules; The similarity matrix analysis module is used to calculate the similarity matrix between two groups of molecules; The substructure search module is used to search for user-defined structural fragments.
7. The software system for online visualization and analysis of DNA-encoded compound library screening data as described in claim 6, characterized in that, The clustering analysis module uses the Butina clustering method, K-means clustering method, or K-medoids clustering method for molecular structure analysis.