Machine learning system and method for predicting blood-brain barrier permeability
The system addresses the limitations of current blood-brain barrier permeability prediction methods by using a chi-square test, k-nearest neighbor algorithm, and logistic regression to train an ensemble meta-learner, achieving efficient and interpretable predictions for molecular design improvements.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- LANTERN PHARMA INC
- Filing Date
- 2024-03-14
- Publication Date
- 2026-04-10
AI Technical Summary
Current machine learning methods for predicting blood-brain barrier permeability are limited by black-box models that lack interpretability, inefficient resource usage, and insufficient understanding of molecular correlations, hindering effective drug development.
A system and method utilizing a machine learning model that generates interpretable predictions of blood-brain barrier permeability by employing a chi-square test, k-nearest neighbor algorithm, and logistic regression to balance datasets, and trains an ensemble meta-learner with reduced features, providing insights into molecular structure-permeability correlations.
The system enhances model performance and resource efficiency while offering deeper understanding of molecular design for improved blood-brain barrier permeability, facilitating better drug development.
Smart Images

Figure 2026510927000001_ABST
Abstract
Description
Technical Field
[0001] (Cross - reference to Related Applications) This application claims the priority and benefit of U.S. Provisional Patent Application No. 63 / 452,108, filed on March 14, 2023, the entire content of which is incorporated herein by reference.
[0002] (Field of the Invention) This application relates to artificial intelligence technology, machine learning technology, blood - brain barrier permeability prediction technology, molecular design technology, and data analysis technology. More specifically, it relates to a machine learning system and an associated method for predicting blood - brain barrier permeability.
Background Art
[0003] The blood - brain barrier is a semi - permeable membrane that effectively separates the blood circulating in a person's central nervous system from the extracellular cerebrospinal fluid. Blood - brain barrier permeability is the ability of various substances to pass through the barrier between a person's bloodstream and brain tissue. The various cells of the blood - brain barrier prevent the passage of many types of molecules, such as those harmful to the brain. However, the blood - brain barrier allows the passage of certain substances such as water, oxygen, and lipid - soluble molecules, enabling the passage of essential nutrients. Currently, being able to effectively enhance blood - brain barrier permeability for drug delivery purposes is a highly sought - after goal. For this purpose, various pharmaceutical companies are using technical tools such as software and artificial intelligence systems to determine or predict the blood - brain barrier permeability of specific molecules being considered for drugs.
[0004] While blood-brain barrier permeability predictions using certain existing machine learning and deep learning methods based on molecular structure have been shown to be somewhat accurate, current approaches suffer from several significant shortcomings that reduce their applicability and usefulness. For example, current state-of-the-art methods produce black-box models that cannot provide insight into why molecules are predicted to be permeable or impermeable, thereby making it virtually impossible to use molecular predictions as a tool for improving blood-brain barrier permeability. Based at least on the foregoing, there remains room for substantial enhancements to existing technologies and processes, as well as for the development of new technologies and processes that provide blood-brain barrier permeability prediction capabilities. For example, current technologies can be improved and enhanced to provide improved artificial intelligence model performance on test data, more efficient use of computing resources while generating models and predictions, greater interpretability, and various other advantages. Such enhancements and improvements to methodologies and technologies could provide a deeper understanding of which parts of molecules correlate with blood-brain barrier permeability and, ultimately, which molecules are the best candidates for treating various health conditions. [Overview of the project]
[0005] A system and associated methods for predicting blood-brain barrier permeability are disclosed. In particular, the system and methods involve utilizing a unique process to generate a machine learning model that can effectively predict whether a particular molecule under consideration has blood-brain barrier permeability, while simultaneously utilizing fewer computing resources and features. As a result, the machine learning models generated by utilizing the system and methods are more robust and interpretable. The functionality provided by the system and methods also facilitates an understanding of how the specific chemical structure of the molecule under consideration affects blood-brain barrier permeability and how molecular design can be improved or modified to enhance blood-brain barrier permeability. Furthermore, the system and methods provide a unique model interpretation analysis that advances the chemical engineering of blood-brain barrier permeability therapeutics.
[0006] In one embodiment, a system for predicting blood-brain barrier permeability is provided. In one embodiment, the system may include a memory for storing instructions and a processor configured to execute instructions and perform various operations. In one embodiment, the processor may be configured to generate a plurality of features for one or more structural representations of one or more molecules. In one embodiment, the plurality of features may include one or more molecular fingerprint representations associated with one or more molecules, descriptors, graph embeddings, any other features, or a combination thereof. In one embodiment, the processor may be configured to perform a chi-square test on the plurality of features of one or more structural representations to determine whether blood-brain barrier permeability depends on one or more molecular fingerprint representations. In one embodiment, the processor may be configured to determine the ratio of permeable to impermeable samples associated with one or more molecules and containing one or more molecular fingerprint representations. In one embodiment, the processor may be configured to generate a balanced training dataset by expanding a training dataset containing permeable and impermeable samples based on ratios, by creating a synthetic dataset of a small number of permeable and impermeable classes using the k-nearest neighbor algorithm until the sample counts in the training dataset are balanced between permeable and impermeable samples. In one embodiment, the processor may be configured to reduce the number of features used for the balanced training dataset by using logistic regression with minimum absolute contraction to create a selected set of features for the balanced training dataset. In one embodiment, the processor may be configured to train an ensemble meta-learner to predict blood-brain barrier permeability by utilizing the balanced training dataset having the selected set of features. In one embodiment, the processor may be configured to analyze candidate molecules for blood-brain barrier permeability by utilizing the ensemble meta-learner. In one embodiment, the processor may be configured to generate predictions of whether a candidate molecule is blood-brain barrier permeable by utilizing the ensemble meta-learner.
[0007] In one embodiment, a method for blood-brain barrier permeability is disclosed. The method may include a memory for storing instructions and a processor for executing instructions to perform the functions of the method. In particular, the method may include generating a plurality of features for one or more structural representations of one or more molecules. In one embodiment, the plurality of features may include one or more molecular fingerprint representations associated with one or more molecules. In one embodiment, the method may include performing a chi-square test on the plurality of features of one or more structural representations to determine whether blood-brain barrier permeability depends on one or more molecular fingerprint representations. In one embodiment, the method may include determining the ratio of permeable to impermeable samples that include one or more molecular fingerprint representations associated with one or more molecules. In one embodiment, the method may include, based on the ratio, generating a balanced training dataset by expanding a training dataset containing permeable and impermeable samples by creating a synthetic data of a small number of permeable and impermeable samples using a k-nearest neighbor algorithm until the sample count of the training dataset is balanced between permeable and impermeable samples. In one embodiment, the method may include reducing multiple features used for a balanced training dataset by using logistic regression with minimum absolute contraction to create a selected set of features for a balanced training dataset. In one embodiment, the method may include training an ensemble meta-learner to predict blood-brain barrier permeability by utilizing a balanced training dataset having a selected set of features. In one embodiment, the method may include analyzing candidate molecules for blood-brain barrier permeability by utilizing an ensemble meta-learner. In one embodiment, the method may include generating predictions of whether a candidate molecule is blood-brain barrier permeable by utilizing an ensemble meta-learner. The method may include and / or be modified to include any of the functions of the system and / or any of the functions described herein.
[0008] According to a further embodiment, a computer-readable device, which, when loaded and executed by a processor, includes instructions causing the processor to perform an operation, the operation generating a plurality of features for at least one structural representation of at least one molecule by utilizing instructions from memory executed by the processor, the plurality of features including at least one molecular fingerprint representation associated with at least one molecule, performing a chi-square test on the plurality of features of at least one structural representation to determine whether blood-brain barrier permeability depends on at least one molecular fingerprint representation, determining the ratio of permeable to impermeable samples associated with at least one molecule and including at least one molecular fingerprint representation, and based on the ratio, a training dataset including permeable and impermeable samples, the training dataset samples A computer-readable device comprising: generating a balanced training dataset by expanding it using the k-nearest neighbor algorithm to create a composite data of a small number of permeable and impermeable samples until the count balances between permeable and impermeable samples; reducing multiple features used for the balanced training dataset using logistic regression with minimum absolute contraction to create a selected set of features for the balanced training dataset; training an ensemble meta-learner to predict blood-brain barrier permeability by utilizing the balanced training dataset with the selected set of features; analyzing candidate molecules for blood-brain barrier permeability by utilizing the ensemble meta-learner; and generating a prediction of whether the candidate molecules are blood-brain barrier permeable by utilizing the ensemble meta-learner.
[0009] These and other features of the system and method for predicting blood-brain barrier permeability are described in the following detailed description, drawings, and appended claims. [Brief explanation of the drawing]
[0010] [Figure 1] This is a schematic diagram of a system for predicting blood-brain barrier permeability according to an embodiment of the present disclosure. [Figure 2] A schematic diagram is shown illustrating a system for constructing a machine learning model to predict blood-brain barrier permeability according to an embodiment of the present disclosure. [Figure 3] The table below shows an exemplary feature generation method for generating features that are considered for building a machine learning model to predict blood-brain barrier permeability according to embodiments of this disclosure. [Figure 4] The table shows a reduced set of features used to construct a machine learning model for predicting blood-brain barrier permeability according to embodiments of this disclosure. [Figure 5] This disclosure provides an exemplary process flow diagram for generating features, reducing features, training a machine learning model, and analyzing candidate molecules for blood-brain barrier permeability according to embodiments of this disclosure. [Figure 6] A table is shown illustrating an exemplary mapping of 2D autocorrelation features back to atomic properties according to embodiments of this disclosure. [Figure 7] This embodiment of the present disclosure shows the mapping of fingerprints to a portion of a molecule. [Figure 8] This flowchart illustrates a sample method for predicting the permeability of a molecule under consideration, using a machine learning model constructed according to an embodiment of the disclosure, and utilizing the machine learning model. [Figure 9] A schematic diagram of a machine in the form of a computer system, according to an embodiment of the present disclosure, in which a set of instructions is executed internally, can facilitate the machine in predicting the blood-brain barrier permeability of molecules. [Modes for carrying out the invention]
[0011] Detailed description of the drawing A system 100 and associated methods for predicting blood-brain barrier permeability are disclosed. In particular, the system 100 and methods include utilizing a novel process to generate a machine learning model that can effectively predict the blood-brain barrier permeability of a particular molecule under consideration while utilizing fewer computing resources and features than existing technologies. The functionality provided by the system 100 and methods also generates information indicating how a particular chemical structure of the molecule under consideration affects blood-brain barrier permeability and how the molecular design can be improved or modified to enhance blood-brain barrier permeability. Furthermore, the system 100 and methods provide a unique model interpretation analysis that advances the chemical engineering of blood-brain barrier permeability therapeutics.
[0012] In one embodiment, a system for predicting blood-brain barrier permeability is provided. In one embodiment, the system may include a memory for storing instructions and a processor configured to execute instructions and perform various operations. In one embodiment, the processor may be configured to generate a plurality of features for one or more structural representations of one or more molecules. In one embodiment, the plurality of features may include one or more molecular fingerprint representations associated with one or more molecules. In one embodiment, the processor may be configured to perform a chi-square test on the plurality of features of one or more structural representations to determine whether blood-brain barrier permeability depends on one or more molecular fingerprint representations. In one embodiment, the processor may be configured to determine the ratio of permeable to impermeable samples that include one or more molecular fingerprint representations associated with one or more molecules. In one embodiment, the processor may be configured to generate a balanced training dataset by expanding a training dataset containing permeable and impermeable samples, based on the ratio, by creating a synthetic data set of a small number of permeable and impermeable samples using a k-nearest neighbor algorithm until the sample count of the training dataset is balanced between permeable and impermeable samples. In one embodiment, the processor may be configured to reduce multiple features utilized for a balanced training dataset by using logistic regression with minimum absolute contraction to create a selected set of features for a balanced training dataset. In one embodiment, the processor may be configured to train an ensemble meta-learner to predict blood-brain barrier permeability by utilizing a balanced training dataset having a selected set of features. In one embodiment, the processor may be configured to analyze candidate molecules for blood-brain barrier permeability by utilizing an ensemble meta-learner. In one embodiment, the processor may be configured to generate predictions of whether a candidate molecule is blood-brain barrier permeable by utilizing an ensemble meta-learner.
[0013] In one embodiment, the processor may be configured to generate one or more structural representations of one or more molecules by converting the three-dimensional structures of one or more molecules into a sequence of symbols identifiable by the system. In one embodiment, the processor may be configured to determine that blood-brain barrier permeability depends on one or more molecular fingerprint representations based on one or more molecular fingerprints having a p-value of less than 0.05 (or other desired value).
[0014] In one embodiment, the processor may be configured to classify permeable samples among a group of samples as permeable based on the fingerprint (i.e., fingerprint representation) associated with permeable samples having threshold blood-brain barrier permeability. In another embodiment, the processor may be further configured to classify impermeable samples among a group of samples as impermeable based on the fingerprint associated with impermeable samples having a threshold blood-brain barrier permeability value less than the threshold blood-brain barrier permeability value. In yet another embodiment, the processor may be configured to determine that impermeable samples are a minority class among a group of samples based on the fact that there are more permeable samples than impermeable samples.
[0015] In one embodiment, the processor may be configured to use logistic regression to reduce the coefficients of a feature among multiple features to zero, thereby excluding the feature from being included in a selected set of features used to train a machine learning model to generate predictions about blood-brain barrier permeability or other predictions. In one embodiment, the processor may be further configured to rank the features in the selected set of features in order of importance based on the absolute value of each coefficient of the feature in the selected set of features. In one embodiment, the processor may be configured to generate an ensemble meta-learner from one or more basic learner models trained on a balanced training dataset and utilizing logistic regression, deep neural networks, or a combination thereof. In one embodiment, the processor may be configured to determine the predicted probability of permeability for holdout validation samples not included in the balanced training dataset. In one embodiment, the processor may be further configured to utilize the predicted probability of permeability for the holdout validation samples as input to a logistic regression meta-learner ensemble model. In one embodiment, the processor may be further configured to select the ensemble meta-learner as a combination of basic learner models having the highest area under the receiver operating characteristic curve.
[0016] In one embodiment, a method for blood-brain barrier permeability is disclosed. The method may include a memory for storing instructions and a processor for executing instructions to perform the functions of the method. In particular, the method may include generating a plurality of features for one or more structural representations of one or more molecules. In one embodiment, the plurality of features may include one or more molecular fingerprint representations associated with one or more molecules. In one embodiment, the method may include performing a chi-square test on the plurality of features of one or more structural representations to determine whether blood-brain barrier permeability depends on one or more molecular fingerprint representations. In one embodiment, the method may include determining the ratio of permeable to impermeable samples that include one or more molecular fingerprint representations associated with one or more molecules. In one embodiment, the method may include, based on the ratio, generating a balanced training dataset by expanding a training dataset containing permeable and impermeable samples by creating a synthetic data of a small number of permeable and impermeable samples using a k-nearest neighbor algorithm until the sample count of the training dataset is balanced between permeable and impermeable samples. In one embodiment, the method may include reducing multiple features used for a balanced training dataset by using logistic regression with minimum absolute contraction to create a selected set of features for a balanced training dataset. In one embodiment, the method may include training an ensemble meta-learner to predict blood-brain barrier permeability by utilizing a balanced training dataset having a selected set of features. In one embodiment, the method may include analyzing candidate molecules for blood-brain barrier permeability by utilizing an ensemble meta-learner. In one embodiment, the method may include generating predictions of whether a candidate molecule is blood-brain barrier permeable by utilizing an ensemble meta-learner.
[0017] In one embodiment, the method may include identifying specific parts of candidate molecules that have blood-brain barrier permeability by utilizing an ensemble meta-learner. In one embodiment, the method may include generating an ensemble meta-learner from at least one base learner model trained on a balanced training dataset and using logistic regression, a deep neural network, or a combination thereof. In one embodiment, the method may include stopping the training of one or more base learner models used to generate the ensemble meta-learner at an epoch representing the highest area under the receiver operating characteristic curve on a holdout sample. In one embodiment, the method may include using logistic regression to reduce the coefficients of a feature among multiple features to zero, thereby excluding the feature from being included in a selected set of features. In one embodiment, the method may include generating one or more structural representations of one or more molecules by converting the three-dimensional structures of one or more molecules into sequences of symbols by a system. In one embodiment, the method may include determining the correlation of blood-brain permeability between at least one molecule and at least one candidate molecule.
[0018] According to a further embodiment, a computer-readable device, which, when loaded and executed by a processor, includes instructions causing the processor to perform an operation, the operation generating a plurality of features for at least one structural representation of at least one molecule by utilizing instructions from memory executed by the processor, the plurality of features including at least one molecular fingerprint representation associated with at least one molecule, performing a chi-square test on the plurality of features of at least one structural representation to determine whether blood-brain barrier permeability depends on at least one molecular fingerprint representation, determining the ratio of permeable to impermeable samples associated with at least one molecule and including at least one molecular fingerprint representation, and based on the ratio, a training dataset including permeable and impermeable samples, the training dataset samples A computer-readable device comprising: generating a balanced training dataset by expanding it using the k-nearest neighbor algorithm to create a composite data of a small number of permeable and impermeable samples until the count balances between permeable and impermeable samples; reducing multiple features used for the balanced training dataset using logistic regression with minimum absolute contraction to create a selected set of features for the balanced training dataset; training an ensemble meta-learner to predict blood-brain barrier permeability by utilizing the balanced training dataset with the selected set of features; analyzing candidate molecules for blood-brain barrier permeability by utilizing the ensemble meta-learner; and generating a prediction of whether the candidate molecules are blood-brain barrier permeable by utilizing the ensemble meta-learner.
[0019] As shown in Figure 1, a system for predicting blood-brain barrier permeability according to embodiments of the present disclosure is disclosed. In particular, System 100 may be configured to support, but is not limited to, automation systems, blood-brain barrier prediction systems, data analysis systems and services, data matching and processing systems and services, artificial intelligence services and systems, machine learning services and systems, content distribution services, cloud computing services, satellite services, telephone services, Voice over Internet Protocol (VoIP) services, Software as a Service (SaaS) applications, Platform as a Service (PaaS) applications, social media applications and services, operational management applications and services, productivity applications and services, mobile applications and services, and / or any other computing applications and services. In particular, System 100 may include a first user 101, and the first user 102 may use the first user device 102 to access data, content, and services or perform various other tasks and functions. As an example, the first user 101 may use the first user device 102 to transmit signals to access various online services and content, such as those available on the Internet, on other devices, and / or on various computing systems. As another example, the first user device 102 may be used by the first user 101 to access applications, devices, and / or components of system 100 that provide some or all of the operational functions of system 100. For example, the first user 101 may use the first user device 102 to access an application supported by a machine learning model, which is used to determine whether a particular molecule under consideration or evaluation (e.g., a drug molecule) is permeable to the blood-brain barrier. In some embodiments, the first user 101 may be any type of person, robot, humanoid, program, computer, any type of user, or a combination thereof, which may be located in a particular environment.
[0020] In certain embodiments, the first user 101 may be a person who attempts to determine whether a particular target molecule has blood-brain barrier permeability. In certain embodiments, the first user device 102 may be utilized by the first user to interact with the system 100, other users of the system 100, or combinations thereof. In certain embodiments, the first user device 102 may include a memory 103 that contains instructions and a processor 104 that executes the instructions from the memory 103 to perform various operations that are carried out by the first user device 102. In certain embodiments, the processor 104 may be hardware, software, or a combination thereof. The first user device 102 may also include an interface 105 (e.g., a screen, monitor, graphical user interface, etc.) that enables the first user 101 to interact with various applications executed on the first user device 102 and to interact with the system 100. In certain embodiments, the first user device 102 may be a computer, any type of sensor, laptop, set-top box, tablet device, phablet, server, mobile device, smartphone, smartwatch, and / or any other type of computing device, and / or may include them. Exemplarily, the first user device 102 is shown as a smartphone device in FIG. 1. In certain embodiments, the first user device 102 may be utilized by the first user 101 to control and / or provide some or all of the operating functions of the system 百.
[0021] [ In addition to using the first user device 102, the first user 101 may also utilize and / or access additional user devices. Similar to the first user device 102, the first user 101 may utilize additional user devices to transmit signals for accessing various online services and content. Additional user devices may include memory containing instructions and a processor that executes instructions from memory to perform various operations performed by the additional user device. In some embodiments, the processor of the additional user device may be hardware, software, or a combination thereof. Additional user devices may also include interfaces that may enable the first user 101 to interact with various applications running on the additional user device and to interact with the system 100. In some embodiments, the first user device 102 and / or additional user devices may be, and / or include, computers, any type of sensor, laptops, set-top boxes, tablet devices, phablets, servers, mobile devices, smartphones, smartwatches, and / or any other type of computing device, and / or any combination thereof. The sensors may include, but are not limited to, cameras, motion sensors, acoustic / audio sensors, pressure sensors, temperature sensors, light sensors, heart rate sensors, blood pressure sensors, sweat detection sensors, eye tracking sensors, respiration detection sensors, stress detection sensors, any type of health sensor, humidity sensors, any type of sensor, or combinations thereof.
[0022] The first user device 102 and / or additional user devices may belong to and / or form a communication network. In certain embodiments, the communication network may be a local, mesh, or other network that enables and / or facilitates various aspects of the functionality of system 100. In certain embodiments, the communication network may be formed between the first user device 102 and the additional user devices through the use of any type of wireless or other protocol and / or technology. For example, the user devices may communicate with each other within the communication network by utilizing any protocol and / or wireless technology, satellite, fiber, or any combination thereof. In particular, the communication network may be configured to link and / or communicate with any other network of system 100 and / or communicate with the outside of system 100.
[0023] In one embodiment, the first user device 102 and additional user devices belonging to the communication network can share and exchange data with each other via the communication network. For example, user devices can share with each other information such as: information associated with a molecule (e.g., a molecule under consideration or evaluation by system 100); information about the chemical structure of the molecule; information about whether the molecule is permeable to the blood-brain barrier; information about a machine learning model that predicts the permeability of the molecule to the blood-brain barrier; information about a fingerprint generated for the molecule; information about features generated from the structure and / or fingerprint representation of the molecule; information about features selected to train a machine learning model for generating predictions; information about various components of the user device; information associated with images and / or content accessed by the user of the user device; information identifying the location of the user device; information indicating the type of sensors included in and / or on the user device; information identifying the application being used on the user device; information identifying how the user device is being used by the user; information identifying the user profile of the user of the user device; information identifying the device profile of the user device; information identifying the number of devices in the communication network; information identifying devices being added to or removed from the communication network; any other information; or any combination thereof.
[0024] In addition to the first user 101, the system 100 may also include a second user 110. In one embodiment, the second user 110 may attempt to determine whether different molecules have blood-brain barrier permeability. In one embodiment, the second user 110 may be a patient or another user who may be the target for administration of a candidate molecule to determine the blood-brain barrier permeability of the candidate molecule. In one embodiment, the second user device 111 may be used by the second user 110 to transmit signals requesting various types of content, services, and data provided and / or accessible by the communication network 135 or any other network within the system 100, such as, but not limited to, artificial intelligence and / or machine learning models of the system 100. In one embodiment, the second user device 111 may be used by the second user 110 to perform any operational function of the system 100, or any combination thereof. In a further embodiment, the second user 110 may be a robot, a computer, a vehicle, a humanoid, an animal, any type of user, or any combination thereof. The second user device 111 may include a memory 112 containing instructions and a processor 113 that executes instructions from the memory 112 to perform various operations performed by the second user device 111. In some embodiments, the processor 113 may be hardware, software, or a combination thereof. The second user device 111 may also include an interface 114 (e.g., a screen, monitor, graphical user interface, etc.) that may enable the first user 101 to interact with various applications running on the second user device 111 and, in some embodiments, to interact with the system 100. In some embodiments, the second user device 111 may be a computer, laptop, set-top box, tablet device, phablet, server, mobile device, smartphone, smartwatch, and / or any other type of computing device. Exemplarily, the second user device 111 is shown as a mobile device in Figure 1.In one embodiment, the second user device 111 may also include, but is not limited to, sensors such as a camera, audio sensor, motion sensor, pressure sensor, temperature sensor, light sensor, heart rate sensor, blood pressure sensor, sweat detection sensor, respiration detection sensor, eye tracking sensor, stress detection sensor, any type of health sensor, humidity sensor, any type of sensor, or a combination thereof.
[0025] In one embodiment, the first user device 102, additional user devices, and / or the second user device 111 may have any number of software applications and / or application services stored on and / or accessible thereon. For example, the first user device 102, additional user devices, and / or the second user device 111 may include applications for controlling and / or accessing the operational features and functions of system 100, applications for controlling and / or accessing any device of system 100, applications for generating molecular fingerprint representations, applications for generating machine learning models, applications for generating features from molecular representations, applications for expanding sample sets including sample imbalances, applications for generating blood-brain barrier permeability predictions, cloud-based applications, VoIP applications, other types of phone-based applications, product ordering applications, business applications, e-commerce applications, media streaming applications, content-based applications, media editing applications, database applications, game applications, internet-based applications, browser applications, mobile applications, service-based applications, productivity applications, video applications, music applications, social media applications, any other types of applications, any type of application services, or a combination thereof. In some embodiments, a software application may support functions provided by the system 100 and method described herein. In some embodiments, the software application and service may include one or more graphical user interfaces to enable a first user 101 and / or potentially a second user 110 to interact with the software application.Software applications and services may also be utilized by first and / or potentially second users 101, 110 to interact with any device within System 100, any network within System 100, or any combination thereof. In one embodiment, the first user device 102, additional user devices, and / or potentially second user devices 111 may include associated telephone numbers, device identification information, or any other identifiers to uniquely identify the first user device 102, additional user devices, and / or second user devices 111.
[0026] In one embodiment, for example, a first user 101 can use a first user device 102 to start the operation of the system 100 itself. For example, the first user 101 can start one or more applications that support the functionality of the system 100, activate the operation of one or more machine learning models to generate predictions about the blood-brain barrier permeability of any number of molecules under consideration. In one embodiment, the first user 101 can use the user interface of the first user device 102 to interact with the applications, trigger the training of machine learning models, trigger the generation of features from samples (e.g., labeled or unlabeled samples, depending on the implementation of the system 100), trigger the system 100 to select a subset of the generated features, generate a machine learning model (e.g., a base model), and start the generation of an ensemble machine learning model (e.g., a combination of base models and / or permutations of base models having the highest area under receiver operating characteristics can be selected as an ensemble meta-learner / machine learning model). In one embodiment, the first user 101 may be able to pause or stop the operation of the system 100. In this embodiment, the first user 101 can upload training data for training the model, for example, via the first user device 102 and / or any other device of the system 100.
[0027] System 100 may also include a communication network 135. The communication network 135 may be under the control of a service provider, any designated user, a computer, another network, or a combination thereof. The communication network 135 of System 100 may be configured to link each of the devices within System 100 to one another. For example, the communication network 135 may be used by a first user device 102 to connect to other devices within or outside the communication network 135. Furthermore, the communication network 135 may be configured to transmit, generate, and receive any information and data traversing System 100. In some embodiments, the communication network 135 may include any number of servers, databases, or other components. The communication network 135 also includes and may connect to mesh networks, local networks, cloud computing networks, IMS networks, VoIP networks, security networks, VoLTE networks, wireless networks, Ethernet® networks, satellite networks, broadband networks, cellular networks, private networks, cable networks, the Internet, Internet Protocol networks, MPLS networks, content distribution networks, any network, or any combination thereof. Exemplarily, servers 140, 145, and 150 are shown as being included within the communication network 135. In some embodiments, the communication network 135 may be part of a single autonomous system located in a particular geographical area, or it may be part of a plurality of autonomous systems spanning several geographical areas.
[0028] In particular, the functions of system 100 can be supported and performed by using any combination of servers 140, 145, 150, and 160. Servers 140, 145, and 150 may reside within the communication network 135, but in some embodiments, servers 140, 145, and 150 may reside outside the communication network 135. Servers 140, 145, and 150 may provide and function as server services, performing various operations and functions provided by system 100. In some embodiments, server 140 may include a memory 141 containing instructions and a processor 142 that executes instructions from memory 141 to perform various operations performed by server 140. The processor 142 may be hardware, software, or a combination thereof. Similarly, server 145 may include a memory 146 containing instructions and a processor 147 that executes instructions from memory 146 to perform various operations performed by server 145. Furthermore, the server 150 may include a memory 151 containing instructions and a processor 152 that executes instructions from the memory 151 and performs various operations performed by the server 150. In some embodiments, servers 140, 145, 150, and 160 may be network servers, routers, gateways, switches, media distribution hubs, signaling points, service control points, service switching points, firewalls, routers, edge devices, nodes, computers, mobile devices, or any other suitable computing devices, or any combination thereof. In some embodiments, servers 140, 145, and 150 may be communicably linked to a communication network 135, any network, any device in system 100, or any combination thereof.
[0029] The database 155 of system 100 can be used to store and relay information across system 100, cache content across system 100, store data about each device within system 100, and perform any other typical functions of a database. In some embodiments, the database 155 may be connected to or reside in a communication network 135, any other network, or a combination thereof. In some embodiments, the database 155 may function as a central repository for any information associated with any device and information associated with system 100. Furthermore, the database 155 may include a processor and memory, or be connected to a processor and memory, to perform various operations related to the database 155. In some embodiments, the database 155 may be connected to servers 140, 145, 150, 160, a first user device 102, a second user device 111, additional user devices, any device within system 100, any process in system 100, any program in system 100, any other device, any network, or any combination thereof.
[0030] Database 155 also stores information and metadata obtained from System 100, metadata and other information associated with the first user 101 and the second user 110, artificial intelligence / machine learning models used by System 100 (e.g., base models and ensemble models), sensor data, samples (e.g., samples of permeable molecules, samples of impermeable molecules, any other type of sample, or combinations thereof), simplified molecular input line entry system (SMILE) structures of molecules, transformations or conversions of SMILE structures, features generated from various samples, molecular fingerprint representations, the results of chi-squared tests performed on various features, augmented samples (e.g., when augmenting imbalanced datasets), information identifying which subsets of features from the generated features were selected by System 100, predictions made by System 100 and / or artificial intelligence models, and information and / or content used to train the artificial intelligence models. The system may store user profiles associated with first and second users 101, 110, device profiles associated with any device in the system 100, communications across the system 100, user preferences, information associated with any device or signal in the system 100, information associated with usage patterns related to user devices 102, 111, information obtained from any network in the system 100, historical data associated with first user 101 and second user 110, device characteristics, information about any device related to first user 101 and second user 110, information associated with communication network 135, any information generated and / or processed by the system 100, any information disclosed about any of the operations and functions disclosed about the system 100, any information generated and / or traversed by the system 100, or any combination thereof.Furthermore, the database 155 may be configured to process queries sent to the database by any device within the system 100.
[0031] In particular, as shown in Figure 1, the system 100 may perform any of the operational functions disclosed herein by utilizing the processing power of the server 160, the storage capacity of the database 155, or any other component of the system 100 to perform the operational functions disclosed herein. The server 160 may include one or more processors 162 that can be configured to handle any of the various functions of the system 100. The processors 162 may be software, hardware, or a combination of hardware and software. Furthermore, the server 160 may also include memory 161 that stores instructions that the processors 162 can execute to perform various operations of the system 100.For example, server 160 may assist in handling the loads handled by various devices within system 100, for example, acquiring samples for training one or more machine learning models (e.g., samples of data containing information indicating whether a particular molecule is permeable or impermeable, unlabeled sample data containing information including properties, structural information, samples associated with previous predictions made by the machine learning model, samples showing the accuracy of previous predictions made by the machine learning model, and / or other information associated with the molecule and / or information associated with blood-brain barrier permeability), generating structural representations of molecules (e.g., SMILE representations), generating any number of features from the structural representations of molecules, performing a chi-squared test (or other test) on the features to determine whether blood-brain barrier permeability depends on the molecular fingerprint representation, determining the ratio of permeable to impermeable samples associated with the molecule(s), and impermeable and permeable samples. This includes expanding the training dataset when the sex samples are unbalanced, reducing the amount of features used for a balanced training dataset by employing techniques such as logistic regression to create a set of features selected for a balanced training dataset, training a machine learning model to predict blood-brain barrier permeability, selecting candidate molecules for evaluation by the machine learning model, analyzing the candidate molecules and related information (e.g., structural representation, fingerprint representation), generating predictions about the blood-brain barrier permeability of the candidate molecules, determining the accuracy of the predictions by comparing them with observable results associated with using the candidate molecules on a user (e.g., a first user 101 or a second user 110), retraining the machine learning model based on the comparison and on new and / or updated data sources, and performing any other operations performed in or by the system 100. In one embodiment, multiple servers 160 can be used to process the functions of the system 100.In one embodiment, the server 160 and other devices within the system 100 may utilize the database 155 to store data about the devices within the system 100 or any other information associated with the system 100. In one embodiment, multiple databases 155 may be used to store data in the system 100.
[0032] Figures 1 to 9 show specific exemplary configurations of various components of System 100, but System 100 may include any configuration of components, which may include using more or fewer components. For example, System 100 is shown exemplary as including a first user device 102, a second user device 111, a communication network 135, servers 140, 145, 150, 160, and a database 155. However, System 100 may include multiple first user devices 102, multiple second user devices 111, multiple communication networks 135, multiple servers 140, multiple servers 145, multiple servers 150, multiple servers 160, multiple databases 155, or any number of other components, internal or external to System 100. Furthermore, in some embodiments, a significant portion of the functions and operations of System 100 may be performed by other networks and systems that may be connected to System 100.
[0033] Referring here to Figure 2, an exemplary schematic diagram of a system 200 for building and training a machine learning model to predict blood-brain barrier permeability is provided. In some embodiments, system 200 may be part of system 100 and / or connected to system 100. In some embodiments, the functions and components of system 100 may be combined with the functions and components of system 200. In some embodiments, system 200 may include, but is not limited to, a database 204, a controller 206, a sample creation component 208 (e.g., a module or software process), a sample database 210, a feature generation / selection component 212, a feature database 214, a learner 216 (e.g., a basic learner and / or an ensemble meta-learner machine learning model), a model registry 218, any other components, or combinations thereof. In some embodiments, database 204 may include data (e.g., raw data) obtained from various data sources, such as, but not limited to, cloud computing systems, remote and / or local devices, applications, other databases, or combinations thereof. The data may be data relating to any number of molecules, such as, but not limited to, information describing the molecules, information indicating the properties of the molecules, information indicating the capabilities of the molecules, information indicating the chemical structure of the molecules, the permeability of the molecules (e.g., permeability, partial permeability, impermeability, etc.), any other information, or a combination thereof.
[0034] In one embodiment, a sample generation component 208 can generate samples from raw data from a database 204 and store the samples in a sample database 210 for future retrieval and / or use. Once samples are generated, a feature generation / selection component 212 can extract features from the various samples stored in the sample database 210. The extracted features can be stored in a feature database 214 for future retrieval and / or use. In one embodiment, a learner 216 (e.g., a basic learner and / or an ensemble meta learner) can be trained using features from the feature database 214 to generate predictions about the blood-brain barrier permeability of molecules. The model can be trained using training data from the feature database 214, its performance can be validated using a validation set of data, and its trials can be validated using trial data. The final model can be stored in a model registry 218 and can be invoked by a software process to generate predictions about candidate molecules for which blood-brain barrier permeability predictions are desired. The model generated by system 200 can be adjusted and updated over time as new data, predictions, and / or prediction accuracy are measured over time.
[0035] Operationally, systems 100, 200 can operate and / or perform functions as described and illustrated in Figures 1 to 9, or as otherwise described herein. In some embodiments, the system may use a regularized feature selection technique to select the resulting features for machine learning in order to conserve computational resources by reducing the total number of features generated from a sample set to the most important features and to more effectively train a machine learning model to perform predictions about blood-brain barrier permeability. In some embodiments, such a process can be used to remove noise from the data by reducing the number of features. In some embodiments, the method used may be a penalty or bias-based regularization method in which different coefficients can be assigned to each feature based on its weight when predicting blood-brain barrier permeability. Referring to Figure 3, an exemplary Table 300 showing 6,473 features obtained from a sample is shown, categorized by the number of samples for each category (e.g., Rdk fingerprint, Morgan fingerprint, MACCS fingerprint, Avalon fingerprint, ERG fingerprint, 2D autocorrelation descriptor, 3D autocorrelation descriptor, rules / filters and corresponding attributes, 3D WHIM descriptor, 3D Getaway descriptor, Avalon fingerprint, ERG fingerprint, etc.). As an example, the 6,473 features extracted from the sample can be reduced to 358 features for effective machine learning. Referring here, also to Figure 4, an exemplary Table 400 showing the reduced set of 358 features is shown. In one embodiment, once molecular fingerprint representations of molecules are generated and / or obtained, the similarity between molecular fingerprints can be effectively analyzed.
[0036] Referring here to Figure 5, an exemplary process flow diagram of a process flow 500 for generating features, reducing features, training a machine learning model, and analyzing candidate molecules for blood-brain barrier permeability according to embodiments of the present disclosure is shown. Process flow 500 may be implemented by utilizing system 100, system 200, or a combination thereof. In 502, process flow 500 may include obtaining data from various data sources 502, including but not limited to Therapeutics Data Commons (TDC) data, data obtained via API calls to systems and / or applications, data obtained by other systems, data generated by system 100 and / or system 200 itself, or a combination thereof. In some embodiments, the data may be raw data, samples (e.g., labeled data (e.g., data relating to molecules labeled as permeable or impermeable, chemical structure information of molecules, any other information relating to molecules, etc.) and / or unlabeled data) and / or other types of data. In one embodiment, in 504, process flow 500 may include obtaining and / or generating structural representations of molecules or other structures in the data obtained in 502. The structural representation may be a SMILE structure (i.e., a chemical structure) that can generate features available for model training. In one embodiment, the features may include, but are not limited to, molecular fingerprint representations, descriptors (e.g., feature descriptors, image descriptors, text descriptors, time descriptors, graph descriptors, etc., which include attributes or properties related to molecules in the data), embeddings, and / or any other type of features. In one embodiment, the structural representation may also include identification of the class of the sample, such as permeable or impermeable (or semi-permeable).
[0037] In 506, the molecular structure can be interpreted by the software of systems 100, 200 and converted into a sequence of symbols that can be understood. In 508, 510, 514, 516, 517, 519, 520, 521, 522, 524, and 526, various types of features can be generated and / or extracted from representations. For example, in 508, a 2D graph can be generated from a representation, and in 510, a 3D graph can be generated from a representation. In some embodiments, embeddings can be generated from 2D and 3D graphs. In some embodiments, for example, 2D and 3D descriptors can be generated from embeddings, and 3D descriptors can be generated from embeddings. In some embodiments, 2D autocorrelation features can be generated in 526, and 3D autocorrelation features can be generated in 524 (i.e., autocorrelation may include calculating correlations between pixels and adjacent pixels, and / or correlations within 3D volumes in various directions, to enable machine learning models to identify patterns in the data). In some embodiments, a 3D WHIM descriptor 514 and a 3D Getaway descriptor 516 can also be generated. In some embodiments, the 3D WHIM descriptor 514 may be a 3D structural descriptor obtained from the atomic coordinates of the 3D molecular structure representation of the molecule, and may include information about size, shape, symmetry, atomic distribution, and / or other information related to the molecule. In some embodiments, the 3D Getaway descriptor (e.g., geometry, topology, and atomic-weight assembly) 516 may be a molecular descriptor that matches the atomic relationships by the 3D molecular geometry and topology provided by the molecular influence matrix, and may include, but is not limited to, information such as atomic weights (e.g., atomic mass, polarizability, van der Waals volume, electronegativity, etc.). From the structural representation, various fingerprints such as 517, 519, 520, 521, and 522 can be generated.
[0038] For example, an RDK fingerprint 517 could be a molecular fingerprint used to represent a molecular structure in a format understandable by system 100 and for use in machine learning models of system 100. In one embodiment, an RDK fingerprint 517 could be a circular fingerprint that examines circular neighborhoods around each atom in the molecule. In one embodiment, the fingerprint can be represented as a bit vector where each bit corresponds to the presence or absence of substructures or patterns in the molecule. An RDK fingerprint 517 can also indicate the size of the fingerprint (e.g., the length of the bit vector). Morgan fingerprints can be used to represent molecular features in a compressed format so that they are suitable for processing by systems 100, 200. Similar to an RDK fingerprint 517, a Morgan fingerprint 519 may include information about the chemical environment around each atom in the molecule. Furthermore, the Morgan fingerprint 519 may include the encoding of substructures into the fingerprint (e.g., bond types, atom types, and / or connectivity patterns), the hashing of substructures into a bit vector representation, the size of the fingerprint, and / or any other Morgan fingerprint 519 information. The Molecular Access System (MACCS) fingerprint 520 may be a binary fingerprint representing the molecular structure as a sequence of binary bits. In some embodiments, each bit represents the presence or absence of a particular substructure or molecular pattern. In some embodiments, the MACCS fingerprint 520 may include a fixed-length representation and a predetermined structural key indicating structural features of the molecule, such as but not limited to functional groups, cyclic systems, and / or other characteristic fragments. In some embodiments, an Avalon fingerprint 521 can also be generated. In some embodiments, the Avalon fingerprint 521 may represent a compound related to the molecule and may indicate a chemical substructure or fragment within the molecule.In some embodiments, the Avalon fingerprint 521 can encode the overall and local structural features of the molecule and may include information such as atomic types, ring systems, bonding information, and / or other molecular structural properties, but not limited to these. In some embodiments, the ERG fingerprint (e.g., an expanded-contracted graph) 522 may be graph-based and can capture structural information in binary format (or another desired format). In some embodiments, the ERG fingerprint 522 may include information about substructures within a radius around each atom in the molecule, and the substructures may be hashed into a fixed-length bit sequence, which can be used to generate a binary fingerprint representing the structural features of the molecule indicated by the substructures.
[0039] Once various features (e.g., fingerprint representations, descriptors, autocorrelation, etc.) are aggregated in 528, the process flow 500 can proceed to split the feature data into training data in 530, test data in 531, and test data in 532, which can be used to train, validate, and test machine learning models generated by systems 100 and 200, respectively. In one embodiment, the data can be prepared using a variety of different seeds, which may involve utilizing slightly different molecular compositions for independent training and testing of the models. Using different seed options, the data records can be shuffled in each split. Such shuffling can facilitate the calculation of averaging and standard deviation of performance metrics. In 534, systems 100 and 200 can analyze the training data and determine whether the training data needs to be balanced, such as when the samples are unbalanced in terms of permeable samples versus impermeable samples. Equilibriumization / class resampling can be performed by generating synthetic data of a minority class (i.e., a type of sample that has fewer samples than other types of samples) to equilibrium the permeable and impermeable samples. In 535, feature selection can be performed to select important features (e.g., features known to be important and tagged as such by systems 100, 200, features shown and / or known to correlate with blood-brain barrier permeability, etc.).
[0040] In 538, after feature selection has been performed, the process flow 500 may include further class resampling as needed. In 540, one or more basic learner models and / or ensemble learning models (as described herein) may be trained and / or constructed using training data having a subset of the selected features. In 531, validation data may be used to validate the performance of the generated and / or trained machine learning model(s). In some embodiments, early stopping 536 may be performed to prevent overfitting of the machine learning model during the training process. In some embodiments, early stopping 536 may include monitoring the performance of the machine learning model(s) via a validation dataset and stopping training when the model's performance begins to degrade, instead of waiting for the model to complete all epochs or iterations. After early stopping in 536, the process flow 500 may include class resampling 538, which may be used to train and validate the model(s). In 532, machine learning models may be tested using test data to verify the performance and predictive ability of the model(s) in determining the blood-brain barrier permeability of molecules. In 544, systems 100, 200 can select the optimal model (i.e., the model with the highest predictive performance, best use of computing resources, etc.), and in 546, candidate molecules to be evaluated can be selected. Based on the training in 548, the machine learning model(s) can analyze the candidate molecules and predict their blood-brain barrier permeability. In 550, the machine learning model can map top-level features (e.g., features correlated with blood-brain barrier permeability) to the atomic properties and / or structure / parts of the molecule, as shown in Figure 6. In a specific exemplary test scenario, accuracy >0.912 and AUC >0.928 were achieved for the test dataset, and annotations for 358 features for individual methods are shown in Figure 4. Furthermore, this test accurately demonstrated blood-brain barrier permeability. This information can be used by systems 100 and 200 to understand and modify molecules in order to make them permeable or impermeable to the blood-brain barrier.In some embodiments, higher-level features (e.g., High Shapley Additive Explanations (SHAP) values) can be remapped to atomic properties / parts of the molecule, and the relative states and relative atomic masses of atoms can play a role in determining blood-brain barrier permeability, which belongs to the 2D autocorrelation method (as shown, for example, in Figure 6). In certain scenarios, chemical functional groups can be matched to those correlated with blood-brain barrier permeability. In some embodiments, fingerprints / features can be remapped to parts of the molecular structure. An example showing the mapping of permeability or impermeability to specific parts of a molecule is shown in Figure 7. In certain scenarios, the use of deep neural network algorithms tends to easily overfit, but the way systems 100, 200 combine feature selection and use validation data for early termination is unique in deriving models with high accuracy and ability to identify key features related to blood-brain barrier permeability.
[0041] In particular, System 100 may perform and / or implement functions such as those described in the following methods (one or more). An exemplary method 800 for constructing and utilizing a machine learning model to generate predictions about blood-brain barrier permeability for molecules, drugs, chemicals, or combinations thereof is schematically shown, as in Figure 8. In some embodiments, the method of Figure 8 can be implemented in any of the systems of Figures 1 to 9, and / or other systems, devices, and / or components shown in the figures. In some embodiments, the method of Figure 8 may be implemented by processing logic that includes hardware (e.g., processing devices, circuits, dedicated logic, programmable logic, microcode, device hardware, integrated circuits, etc.), software (e.g., instructions executed on the processing devices), or a combination thereof. In some embodiments, the method of Figure 8 may be implemented at least partially by one or more processing devices (e.g., processors 102, 122, 141, 146, 151, and 161 in Figure 1). Although shown in a specific sequence or order, unless otherwise specified, the order of steps in Method 800 may be modified and / or changed depending on the implementation and purpose. Therefore, the illustrated embodiments should be understood as merely examples, and the illustrated processes may be executed in a different order, and some processes may be executed in parallel. Furthermore, one or more processes may be omitted in various embodiments. Therefore, not all processes are required in all embodiments. Other process flows are also possible.
[0042] In some embodiments, Method 800 and / or functions and features supporting Method 800 may be performed via applications of System 100, machine learning and / or artificial intelligence models of System 100, devices of System 100, processes of System 100, any components of System 100, or combinations thereof. Generally, Method 800 may include the steps of obtaining a sample of data associated with a molecule from various data sources, converting the sample into a structural representation, and generating a number of features from the structural representation, such as a fingerprint representation, descriptor, graph embedding, and / or other representation. The Method includes performing tests on features to determine blood-brain barrier permeability dependence, such as based on a particular fingerprint representation of a molecule. The Method includes analyzing the ratio of permeable to impermeable samples in a sample of data, and if an imbalance between sample types is determined by the System, extending the sample with synthetic data to create a balanced dataset. The Method further includes reducing the features used to train machine learning using techniques such as logistic regression to create a selected set of features for the balanced dataset. This method involves training a machine learning model using a balanced dataset, and then using the machine learning model to predict the blood-brain barrier permeability of candidate molecules.
[0043] In step 802, method 800 may include generating and / or obtaining one or more structural representations of molecules from a plurality of samples. In some embodiments, the samples may be labeled samples (e.g., labeled as permeable or impermeable, or having other labels associated with the molecule) or unlabeled samples (e.g., when reinforcement learning is utilized with the system). In some embodiments, the samples may include data and information obtained from various data sources, such as Therapeutics Data Commons data, having any other type of structure for any number and / or type of molecule, such as SMILE chemical structures, textual chemical structures, visual chemical structures, audio descriptions of chemical structures, audiovisual chemical structures, and / or chemical molecules. In some embodiments, a SMILE structure may represent the structure of a particular molecule and may include a string of characters (e.g., ASCII) that uniquely represents the structure of the molecule. In some embodiments, the structure may include stereochemical and / or connectivity information and may be configured to be human-readable, machine-readable, or a combination thereof. In some embodiments, the structure may include atomic symbols (e.g., O for oxygen), bond symbols (e.g., single bond "-", double bond "=", triple bond "#", etc.), isotopic information, chirality (e.g., using "@" and " / " to represent chiral centers and cis-trans isomerism), hydrogen information (e.g., adding H to the representation for explicit hydrogen), branching and ring structures, and / or any other information. In some embodiments, the generation and / or acquisition of a structural representation from a sample may be performed and / or facilitated by utilizing a first user 101, a second user 110, and / or by utilizing a first user device 102, a second user device 111, servers 140, 145, 150, 160, communication network 135, any component of system 100, any combination thereof, or by utilizing any other suitable program, network, system, or device.In one embodiment, the sample itself can be obtained from data sources such as, but not limited to, data repositories, databases, cloud systems and / or networks, live data feeds, API calls to third-party systems, any other data sources, or combinations thereof.
[0044] In step 804, method 800 may include generating multiple features from one or more structural representations. In one embodiment, the features may be numerical and / or other types of features, one or more fingerprint representations (e.g., Morgan fingerprint, RDKit (RDK) fingerprint, MACCS fingerprint, Avalon fingerprint, ERG fingerprint, and / or other types of fingerprints), descriptors (e.g., 2D and / or 3D autocorrelation descriptor, 3D WHIM descriptor, and / or 3D Features may include Getaway descriptors, graph embeddings (e.g., graph embeddings based on 2D and / or 3D structures), and / or any other type of feature. In some embodiments, numerical features (e.g., some may be binary and some may be non-binary) may be considered as substitutes for different atomic properties such as connectivity (element, number of heavy neighboring atoms, number of hydrogen (H) atoms, charge, isotopes), chemical features (e.g., donor, acceptor, aromatic, halogen, basic, acidic, etc.), bond type, atomic mass, and electrotolpological state. In some embodiments, additional features such as atomic rules (e.g., Lipinski rules, Ghose filters, Veber filters, etc.) and their corresponding features may be included. Attributes such as can also be used. Features can be used to supply machine learning models for training purposes, such as for building and training machine learning models to perform blood-brain barrier permeability prediction of target molecules. In some embodiments, the generation of multiple features can be performed and / or facilitated by using a first user 101, a second user 110, and / or by using a first user device 102, a second user device 111, servers 140, 145, 150, 160, communication network 135, any component of system 100, any combination thereof, or by using any other suitable program, network, system, or device.
[0045] In step 806, method 800 may include performing one or more tests to determine whether permeability depends on a particular fingerprint (and / or other feature) as part of feature engineering by method 800. In one embodiment, for example, a chi-squared statistical test may be performed on all features (e.g., individual fingerprint features) to determine whether permeability depends on the fingerprint (and / or other feature). In one embodiment, fingerprints with a p-value less than 0.05 may be considered significant in the first phase. In one embodiment, in the second phase, the ratio of permeable to impermeable samples may be calculated for each drug molecule sample containing each fingerprint. In one embodiment, fingerprints with a permeability of 50% or less may be classified as fingerprints with a significant negative association, and fingerprints with a permeability of 80% or more may be considered fingerprints with a significant positive association. In one embodiment, the threshold for negative or positive association may be adjusted to a desired value. In one embodiment, a new feature is created that sums the counts of negatively and positively associated fingerprints in each drug molecule sample. In some embodiments, one or more tests can be performed and / or facilitated by utilizing a first user 101, a second user 110, and / or by utilizing a first user device 102, a second user device 111, servers 140, 145, 150, 160, communication network 135, any component of system 100, any combination thereof, or any other suitable program, network, system, or device.
[0046] In step 808, method 800 may include determining the ratio of permeable samples (e.g., drug molecule samples) to impermeable samples. In some embodiments, the samples may include training samples, validation samples, and test samples. Often, there is an imbalance in the samples obtained from a data source. For example, in a particular dataset, there may be significantly more permeable samples than impermeable samples. In some embodiments, the determination of the ratio can be performed and / or facilitated by utilizing a first user 101, a second user 110, and / or by utilizing a first user device 102, a second user device 111, servers 140, 145, 150, 160, communication network 135, any component of system 100, any combination thereof, or by utilizing any other suitable program, network, system, or device. In step 810, method 800 may include determining whether there is an imbalance between permeable and impermeable samples, for example, whether the samples are unequal or whether the other sample is not within a threshold number. In one embodiment, the determination can be performed and / or facilitated by utilizing the first user 101, the second user 110, and / or by utilizing the first user device 102, the second user device 111, the servers 140, 145, 150, 160, the communication network 135, any component of the system 100, any combination thereof, or by utilizing any other suitable program, network, system, or device.
[0047] If an imbalance exists, method 800 can proceed to step 812. Step 812 may include expanding the dataset (e.g., the training dataset) to correct for the sample imbalance. For example, the dataset can be expanded using various techniques, such as using the k-nearest neighbors algorithm on the samples to create synthetic data of a minority class (i.e., samples of a type that has fewer samples than other types) to provide a balanced training dataset. For example, in an exemplary scenario, due to a class imbalance between the majority (75.3%) of drug samples being permeable and the minority (24.7%) of samples being impermeable, the Synthetic Minority Oversampling Technique (SMOTE) can be performed on the training set of data using the k-nearest neighbors algorithm to create synthetic data of a minority class (e.g., samples of a type that has fewer samples than other types) until the sample count balances between permeable and impermeable observations. This expansion to the training dataset can serve as input to the feature selection and model training phases of a workflow related to building a machine learning model for predicting blood-brain barrier permeability. In one embodiment, augmented samples are generated only for samples used to train a machine learning model, and do not need to be generated for validation or test sets of data samples. In one embodiment, augmentation of a dataset can be performed and / or facilitated by using a first user 101, a second user 110, and / or by using a first user device 102, a second user device 111, servers 140, 145, 150, 160, communication network 135, any component of system 100, any combination thereof, or any other suitable program, network, system, or device.
[0048] However, if there is no imbalance in the samples in step 810, method 800 may proceed directly from step 810 to 814, or if the dataset expansion was performed in step 812, method 800 may proceed from step 812 to 814. In step 814, method 800 may include reducing several features (e.g., generated features) that are utilized in a balanced training dataset using any number of techniques, such as performing feature selection from the entire set of features. For example, features can be reduced by using logistic regression with minimum absolute shrinkage and selection operator penalty to shrink the coefficients of the least important features to zero, thereby excluding them from the features selected for use in model training. In one embodiment, 10-fold cross-validation may be performed on the search range of L1 regularization parameter values, which provides the highest mean accuracy across the splits. The L1 regularization value identified as optimal in the search may be used to train a final feature selection model that eliminates features with zero coefficients to be removed, and the remaining features are ranked as most important by the absolute values of their coefficients. In one embodiment, the feature selection process described above is performed once on only a subset of fingerprint features, and all generated types of features can be reviewed again. In one embodiment, two separate feature selection lists can be used in separate models in subsequent phases to generate diversity in model training and prediction. In one embodiment, feature and / or feature selection reduction can be performed and / or facilitated by utilizing the first user 101, the second user 110, and / or by utilizing the first user device 102, the second user device 111, servers 140, 145, 150, 160, communication network 135, any component of system 100, any combination thereof, or by utilizing any other suitable program, network, system, or device.
[0049] In step 816, the method may include training one or more machine learning models to predict blood-brain barrier permeability, for example, by utilizing a balanced training dataset having a selected set of features. In one embodiment, for example, any number of basic learners may be trained and then ultimately used to create and / or train an ensemble meta-learner machine learning model. In one embodiment, various methods may be used for basic learner design and training. The method may include modeling for basic learners that are used to generate a variety of prediction methods that serve as input to the ensemble meta-learner in subsequent phases. In one embodiment, logistic regression may be utilized. In one embodiment, L1 regularization may be used to reduce features during feature selection. In one embodiment, the basic learner model may not include further regularization and may function to provide an easily interpretable model based on the univariate effect of each feature used for training. Another method for training and / or designing a basic learner model may include using a deep neural network. In one embodiment, the neural network design may be generated by searching for an optimal architecture of 2 to 5 fully connected dense hidden layers (or any other desired range). Each hidden dense layer may include L2 regularization and may be followed by a dropout layer. The search further includes the range of neurons to use in each hidden layer. Each iteration of the architectural search may include an early stopping criterion that stops model training at an epoch representing the highest area under the receiver operating characteristic curve (AUC-ROC) on the holdout samples, using a holdout subset from the training data. In one embodiment, the final model may be constructed using the optimal number of layers and neurons identified in the search and may be fully trained up to the early stopping criterion. In one embodiment, each basic learner model may be trained with augmented training data for each set of feature lists identified during the feature selection phase. After training, the predicted probability of transparency can be calculated for holdout validation samples that were not included in the model's training samples.In one embodiment, validation sample predictions can be used when training a subsequent ensemble meta-learner.
[0050] In one embodiment, a trained base learner model may be evaluated by utilizing validation samples to assess the model's performance, such that the model's parameters (e.g., hyperparameters and / or other parameters) can be adjusted and / or tuned to improve its predictive ability to predict blood-brain barrier permeability. Validation sample predictions from the base learner model can be used as feature inputs to a logistic regression meta-learner ensemble model. In one embodiment, all base models can be evaluated as meta-learner inputs, and the combination of base learners having the highest area under the receiver operating characteristic curve can be selected as the final meta-learner used to make blood-brain barrier predictions for candidate molecules. In one embodiment, if a validation set of samples is used to evaluate and tune the model's performance, such as on an iteration basis, test samples can be used to provide an unbiased assessment of the model's performance. In one embodiment, generating predictions on a test set for evaluation is completed in two stages. First, the probability of permeability for each sample can be predicted for each base learner and its associated selected features. Secondly, the base learner probabilities of the test set can be used as input to a meta-learner ensemble model for final prediction.
[0051] In some embodiments, the machine learning / artificial intelligence model may be, include, and / or utilize, a deep convolutional neural network, a one-dimensional convolutional neural network, a two-dimensional convolutional neural network, a long short-term memory network, an autoencoder, a generative adversarial network, a vision transformer, any type of machine learning system, any type of artificial intelligence system, or a combination thereof. In some embodiments, the model may incorporate the use of any type of artificial intelligence and / or machine learning algorithm to facilitate the operation of the artificial intelligence model(s). In particular, system 100 may utilize any number of artificial intelligence models. System 100 may train the artificial intelligence model(s) to infer and learn from the data / information supplied to system 100, thereby enabling the model to generate and / or facilitate predictions about new data and information supplied to system 100 for analysis. For example, a machine learning model(s) may be trained using data samples such as, but are not limited to, images, video content, audio content, text content, augmented reality content, virtual reality content, information about patterns, information about molecules, any type of data, or a combination thereof. The data used to train an artificial intelligence model may be used by the AI model to predict whether a particular molecule is permeable to the blood-brain barrier.
[0052] Once the ensemble meta-learner is created from the basic learner model, Method 800 can proceed to step 818. In step 818, Method 800 may include selecting candidate molecules for evaluation. In one embodiment, the selection of candidate molecules can be performed and / or facilitated by utilizing the first user 101, the second user 110, and / or by utilizing the first user device 102, the second user device 111, servers 140, 145, 150, 160, communication network 135, any component of System 100, any combination thereof, or by utilizing any other suitable program, network, system, or device. In step 820, Method 800 may include analyzing the candidate molecules, for example, by utilizing the ensemble meta-learner machine learning model. In one embodiment, the analysis can be performed and / or facilitated by utilizing the first user 101, the second user 110, and / or by utilizing the first user device 102, the second user device 111, the servers 140, 145, 150, 160, the communication network 135, any component of the system 100, any combination thereof, or by utilizing any other suitable program, network, system, or device.
[0053] In step 822, method 80 may include generating predictions regarding the blood-brain barrier permeability of a candidate. In one embodiment, the predictions can be performed and / or facilitated by utilizing a first user 101, a second user 110, and / or by utilizing a first user device 102, a second user device 111, servers 140, 145, 150, 160, communication network 135, any component of system 100, any combination thereof, or by utilizing any other suitable program, network, system, or device. In step 824, method 800 may include determining the accuracy of the predictions. In one embodiment, for example, determining the accuracy may include comparing the predictions made by the machine learning model with observable results obtained from a user (e.g., second user 110) using molecules during treatment, etc. In one embodiment, the determination of accuracy can be performed and / or facilitated by utilizing the first user 101, the second user 110, and / or by utilizing the first user device 102, the second user device 111, the servers 140, 145, 150, 160, the communication network 135, any component of the system 100, any combination thereof, or by utilizing any other suitable program, network, system, or device. In step 826, method 800 may include training a machine learning model (e.g., an ensemble meta-learner and / or a basic learner) based on the predictions made, the accuracy of the predictions, updated samples of data, new samples of data, or a combination thereof. The process may be repeated as desired so that the predictive ability of the machine learning model(s) improves over time. In particular, method 800 may further incorporate any of the features and functions described for system 100, any other methods disclosed herein, or as otherwise described herein.
[0054] The systems and methods disclosed herein may include further functions and features. For example, the operational functions of System 100 and the methods may be configured to run on a dedicated processor specifically configured to perform the operations provided by System 100 and the methods. In particular, the operational features and functions provided by System 100 and the methods can improve the efficiency of computing devices used to facilitate the functions provided by System 100 and the various methods disclosed herein. For example, by training System 100 over time based on data and / or other information provided and / or generated in System 100, the amount of computer operations that need to be performed by devices within System 100 using the processor and memory of System 100 can be reduced compared to conventional methodologies. In such a situation, less processing power needs to be utilized because the processor and memory do not need to be dedicated to processing. As a result, the use of computer resources is significantly reduced by utilizing the software, techniques, and algorithms provided herein. In some embodiments, the various operational functions of System 100 may be configured to run on one or more graphics processors and / or application-specific integrated processors.
[0055] In particular, in some embodiments, various functions and features of System 100 and the method can operate without human intervention and can be fully performed by computing devices. In some embodiments, for example, a number of computing devices can interact with the devices of System 100 to provide functions supported by System 100. Furthermore, in some embodiments, the computing devices of System 100 can operate continuously and without human intervention to reduce the possibility of errors being introduced into System 100. In some embodiments, System 100 and the method can also provide effective computing resource management by utilizing the features and functions described herein. For example, in one embodiment, a device within System 100 may transmit a signal indicating that only a specific amount of computer processor resources (e.g., processor clock cycles, processor speed, etc.) can be used to train an artificial intelligence model(s), generate features from samples, perform tests on features, rebalance an unbalanced data sample set, reduce the features used to train a model, analyze candidate molecules to predict blood-brain barrier permeability, determine the accuracy of predictions, retrain a machine learning model, and / or perform any other operation performed by System 100, or any combination thereof. For example, the signal may indicate the number of processor cycles of a processor, specify a selected amount of processing power that may be used to update and / or train an artificial intelligence model, and / or dedicated to generating any of the operations performed by System 100. In one embodiment, a signal indicating a specific amount of computer processor resources or computer memory resources used to perform an operation of System 100 may be transmitted from a first user device 102 and / or a second user device 111 to various components of System 100.
[0056] In one embodiment, any device within system 100 may signal a memory device to dedicate only a selected amount of memory resources to various operations of system 100. In another embodiment, system 100 and the method may also include signaling a processor and memory to perform an operational function of system 100 and the method only during periods when the usage of processing resources and / or memory resources within system 100 is a selected value. In yet another embodiment, system 100 and the method may include signaling a memory device used by system 100 to indicate which particular section of memory should be used to store any data used or generated by system 100. In particular, signals sent to the processor and memory may be used to optimize the use of computing resources while performing operations performed by system 100. As a result, such functionality provides substantial operational efficiency and improvement over existing technologies.
[0057] Referring here to Figure 9, at least some of the methodologies and techniques described with respect to exemplary embodiments of System 100 may incorporate, but are not limited to, other computing devices or other machines that, when executed, cause the computer system 900 or a set of instructions to perform any one or more of the methodologies or functions discussed above. The machine may be configured to facilitate various operations performed by System 100. For example, the machine may be configured to assist System 100 by providing processing power to support the processing load experienced in System 100, by providing storage capacity for storing instructions or data that traverse System 100, or by assisting any other operations performed by or within System 100. As another example, computer system 900 may assist in generating models associated with generating predictions related to the permeability of molecules across the blood-brain barrier, any type of prediction generated by System 100, or a combination thereof. As another example, computer system 600 may assist in training the machine learning model of system 100, selecting features for training the machine learning model of system 100, generating molecular fingerprint representations, identifying molecular parts related to blood-brain barrier permeability, providing any other functions provided by system 100, or a combination thereof.
[0058] In some embodiments, the machine may operate as a standalone device. In some embodiments, the machine may be connected (e.g., using a communication network 135, another network, or a combination thereof) to other machines and systems, such as a first user device 102, a second user device 111, servers 140, 145, 150, a database 155, a server 160, any other systems, programs, and / or devices, or any combination thereof, and may assist in operations performed by them. The machine may be connected to any component within system 100. In a networked deployment, the machine may operate as a server or client user machine in a server-client user network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine may include a server computer, a client user computer, a personal computer (PC), a tablet PC, a laptop computer, a desktop computer, a control system, a network router, a switch or bridge, or any machine capable of executing a set of instructions (sequential or otherwise) specifying actions to be performed by that machine. Furthermore, although a single machine is shown, the term “machine” should also be interpreted to include any set of machines that individually or collectively execute a set (or set) of instructions in order to perform any one or more of the methodologies described herein.
[0059] The computer system 900 may include a processor 902 (e.g., a central processing unit (CPU), a graphics processing unit (GPU, or both)), main memory 904, and static memory 906, which communicate with each other via a bus 908. The computer system 900 may further include a video display unit 910, which may be, but is not limited to, a liquid crystal display (LCD), a flat panel, a solid-state display, or a cathode ray tube (CRT). The computer system 900 may also include an input device 912, such as a keyboard, a cursor control device 914, such as a mouse, a disk drive unit 916, a signal generating device 918, such as a speaker or remote control, and a network interface device 920.
[0060] In some embodiments, the disk drive unit 916 may include a machine-readable medium 922 that stores one or more instruction sets 924, such as software that embodies any one or more of the methodologies or functions described herein, including, but not limited to, the methods described above. The instructions 924 may also be entirely or at least partially present in the main memory 904, static memory 906, or processor 902, or a combination thereof, during their execution by the computer system 900. In some embodiments, the main memory 904 and processor 902 may also constitute the machine-readable medium.
[0061] Dedicated hardware implementations, including but not limited to application-specific integrated circuits, programmable logic arrays, and other hardware devices, can also be constructed to implement the methods described herein. Applications, which may include devices and systems of various embodiments, broadly encompass a variety of electronic and computer systems. Some embodiments implement functionality in two or more specific interconnected hardware modules or devices, or as part of an application-specific integrated circuit, where relevant control and data signals are communicated between and through the modules. Thus, exemplary systems are applicable to software, firmware, and hardware implementations.
[0062] According to various embodiments of this disclosure, the methods described herein are intended to operate as software programs running on a computer processor. Furthermore, software implementations may include, but are not limited to, distributed processing or component / object distributed processing, parallel processing, or virtual machine processing, and can also be constructed to implement the methods described herein.
[0063] This disclosure envisions a machine-readable medium 922 including instructions 924, which enables a device connected to a communication network 135, another network, or a combination thereof to transmit or receive voice, video, or data using instructions and to communicate through the communication network 135, another network, or a combination thereof. The instructions 924 may further be transmitted or received through the communication network 135, another network, or a combination thereof via a network interface device 920.
[0064] Although the machine-readable medium 922 is shown as a single medium in exemplary embodiments, the term “machine-readable medium” should be interpreted to include a single or multiple mediums (e.g., a centralized or distributed database, and / or associated caches and servers) that store one or more sets of instructions. The term “machine-readable medium” should also be interpreted to include any medium capable of storing, encoding, or carrying a set of instructions for machine execution, which causes a machine to execute any one or more of the methodologies of this disclosure.
[0065] Accordingly, the terms “machine-readable medium,” “machine-readable device,” or “computer-readable device” shall be interpreted as including, but not limited to, memory devices, solid-state memory such as memory cards, or magneto-optical or optical media such as other packages, disks or tapes containing one or more read-only (non-volatile) memories, random-access memories, or other rewritable (volatile) memories, or other self-contained information archives or sets of archives that are considered equivalent distribution media to tangible storage media. In some embodiments, “machine-readable medium,” “machine-readable device,” or “computer-readable device” may not be transient, and in some embodiments may not include waves or signals themselves. Accordingly, this disclosure shall be deemed to include any one or more machine-readable mediums or distribution media, including equivalents and successor media recognized in the Art that are enumerated herein and on which software implementations herein are stored.
[0066] The examples of configurations described herein are intended to provide a general understanding of the structures of various embodiments and are not intended to serve as a complete description of all elements and features of devices and systems that may utilize the configurations described herein. Other configurations may be used and derived therefrom, and as a result, structural and logical substitutions and modifications may be made without departing from the scope of this disclosure. The figures are also merely illustrative and may not be drawn to scale. Certain proportions may be exaggerated and other proportions may be minimized. Therefore, this specification and the drawings should be considered illustrative rather than restrictive.
[0067] Therefore, while specific configurations are illustrated and described herein, it should be understood that any configuration calculated to achieve the same objective can be used in place of any specific illustrated configuration. This disclosure is intended to encompass any and all adaptations or variations of various embodiments and configurations of the present invention. Combinations of the above configurations, and other configurations not specifically described herein, will be apparent to those skilled in the art by considering the above description. Thus, this disclosure is not limited to any specific configuration(s) disclosed as the best form contemplated for carrying out this disclosure, and the present invention is intended to encompass all embodiments and configurations included in the appended claims.
[0068] The above is provided for the purpose of illustrating, explaining, and describing embodiments of the present invention. Modifications and adaptations to these embodiments will be obvious to those skilled in the art and can be made without departing from the scope or spirit of the invention. Considering the embodiments described above, it will be obvious to those skilled in the art that the embodiments can be modified, narrowed, or expanded without departing from the scope and spirit of the claims described below.
Claims
1. It is a system, Memory for storing instructions, A processor configured to execute instructions, wherein the instructions cause the processor to Generate a plurality of features for at least one structural representation of at least one molecule, wherein the plurality of features include at least one molecular fingerprint representation associated with the at least one molecule. To determine whether blood-brain barrier permeability depends on the at least one molecular fingerprint representation, a chi-square test is performed on the plurality of features of the at least one structural representation. Determine the ratio of permeable samples to impermeable samples that are associated with the at least one molecule and include the at least one molecular fingerprint representation. Based on the aforementioned ratio, the training dataset, which includes the transparent and impermeable samples, is expanded by creating a composite data set of a small number of transparent and impermeable samples using the k-nearest neighbor algorithm until the sample count of the training dataset is balanced between the transparent and impermeable samples, thereby generating a balanced training dataset. To create a selected set of features for the balanced training dataset, logistic regression with minimum absolute contraction is used to reduce the number of features available for the balanced training dataset. By utilizing the balanced training dataset having the selected set of features, an ensemble meta-learner is trained to predict blood-brain barrier permeability. By utilizing the aforementioned ensemble meta-learner, candidate molecules for blood-brain barrier permeability are analyzed. A system configured to use the aforementioned ensemble meta-learner to generate predictions as to whether the candidate molecule has blood-brain barrier permeability.
2. The system according to claim 1, wherein the processor is further configured to generate the at least one structural representation of the at least one molecule by converting the three-dimensional structure of the at least one molecule into a sequence of symbols identifiable by the system.
3. The system according to claim 1, wherein the aforementioned features further include descriptors, graph embeddings, or a combination thereof.
4. The system according to claim 1, wherein the processor is further configured to determine that blood-brain barrier permeability depends on the representation of the at least one molecular fingerprint, based on the at least one molecular fingerprint having a p-value of less than 0.
05.
5. The system according to claim 1, wherein the processor is further configured to classify the permeable samples among the plurality of samples as permeable based on a fingerprint associated with the permeable samples having threshold blood-brain barrier permeability, and the processor is further configured to classify the impermeable samples among the plurality of samples as impermeable based on a fingerprint associated with the impermeable samples having less than threshold blood-brain barrier permeability.
6. The system according to claim 1, wherein the processor is further configured to determine that the impermeable sample is the minority class among the plurality of samples, based on the fact that the number of permeable samples is greater than the number of impermeable samples.
7. The system according to claim 1, wherein the processor is further configured to use logistic regression to reduce the coefficients of the features among the plurality of features to zero, thereby excluding the feature from being included in the selected set of features.
8. The system according to claim 1, wherein the processor is further configured to rank the features in the selected set of features in order of importance based on the absolute value of each coefficient of the features in the selected set of features.
9. The system according to claim 1, wherein the processor is further configured to generate the ensemble meta-learner from at least one basic learner model trained on the balanced training dataset and utilizing logistic regression, a deep neural network, or a combination thereof.
10. The system according to claim 1, wherein the processor is further configured to determine the predicted probability of penetration for holdout validation samples not included in the balanced training dataset.
11. The system according to claim 10, wherein the processor is further configured to use the predicted probability of the transparency of the holdout validation sample as input to a logistic regression meta-learner ensemble model.
12. The system according to claim 1, wherein the processor is further configured to select the ensemble meta-learner as a combination of basic learner models having the highest area under the receiver operating characteristic curve.
13. It is a method, By utilizing instructions from memory executed by a processor, a plurality of features are generated for at least one structural representation of at least one molecule, wherein the plurality of features include at least one molecular fingerprint representation associated with the at least one molecule. By utilizing the instructions from the memory executed by the processor, a chi-square test is performed on the plurality of features of the at least one structural representation in order to determine whether blood-brain barrier permeability depends on the at least one molecular fingerprint representation, Determining the ratio of permeable samples to impermeable samples that are associated with the at least one molecule and include the at least one molecular fingerprint representation, Based on the aforementioned ratio, the training dataset, which includes the transparent and impermeable samples, is expanded by creating a composite data of a small number of transparent and impermeable samples using the k-nearest neighbor algorithm until the sample count of the training dataset is balanced between the transparent and impermeable samples, thereby generating a balanced training dataset. To create a selected set of features for the balanced training dataset, logistic regression with minimum absolute contraction is used to reduce the number of features used for the balanced training dataset. By utilizing the balanced training dataset having the selected set of features, an ensemble meta-learner is trained to predict blood-brain barrier permeability, By utilizing the aforementioned ensemble meta-learner, candidate molecules for blood-brain barrier permeability can be analyzed, By utilizing the ensemble meta-learner and the instructions from the memory executed by the processor, a prediction is generated as to whether the candidate molecule has blood-brain barrier permeability. Methods that include...
14. The method according to claim 13, further comprising using the ensemble meta-learner to identify a specific portion of the candidate molecule having blood-brain barrier permeability.
15. The method according to claim 13, further comprising generating the ensemble meta-learner from at least one basic learner model trained on the balanced training dataset and utilizing logistic regression, a deep neural network, or a combination thereof.
16. The method according to claim 13, further comprising stopping the training of at least one basic learner model used to generate the ensemble meta learner at an epoch representing the highest area under the receiver operating characteristic curve on the holdout sample.
17. The method according to claim 13, further comprising using the logistic regression to reduce the coefficient of one of the features to zero, thereby excluding the feature from being included in the selected set of features.
18. The method according to claim 13, further comprising generating the at least one structural representation of the at least one molecule by converting the three-dimensional structure of the at least one molecule into a sequence of symbols.
19. The method according to claim 13, further comprising determining the correlation of blood-brain permeability between the at least one molecule and the at least one candidate molecule.
20. A non-temporary computer-readable device comprising instructions, wherein, when the instructions are loaded and executed by a processor, the processor, Generate a plurality of features for at least one structural representation of at least one molecule, wherein the plurality of features include at least one molecular fingerprint representation associated with the at least one molecule. To determine whether blood-brain barrier permeability depends on the at least one molecular fingerprint representation, a chi-square test is performed on the plurality of features of the at least one structural representation. Determine the ratio of permeable samples to impermeable samples that are associated with the at least one molecule and include the at least one molecular fingerprint representation. Based on the aforementioned ratio, the training dataset, which includes the transparent and impermeable samples, is expanded by creating a composite data set of a small number of transparent and impermeable samples using the k-nearest neighbor algorithm until the sample count of the training dataset is balanced between the transparent and impermeable samples, thereby generating a balanced training dataset. To create a selected set of features for the balanced training dataset, logistic regression with minimum absolute contraction is used to reduce the number of features available for the balanced training dataset. By utilizing the balanced training dataset having the selected set of features, an ensemble meta-learner is trained to predict blood-brain barrier permeability. By utilizing the aforementioned ensemble meta-learner, candidate molecules for blood-brain barrier permeability are analyzed. A non-temporal computer-readable device configured to generate a prediction of whether the candidate molecule has blood-brain barrier permeability using the aforementioned ensemble meta-learner.