Systems and methods for a machine learning model that includes a self-attention layer, a long short-term memory layer, a dropout layer, and a one-dimensional convolutional neural network layer
Patent Information
- Application Number
- PCT/US2026/021169
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-27
- Filing Date
- 2026-03-27
- Publication Date
- 2026-10-01
Smart Images

Figure US2026021169_01102026_PF_FP_ABST
Abstract
Description
Attorney Docket No.: BSAI-OOl / OIWO 357321-2002SYSTEMS AND METHODS FOR A MACHINE LEARNING MODEL THAT INCLUDES A SELF-ATTENTION LAYER, A LONG SHORT-TERM MEMORY LAYER, A DROPOUT LAYER, AND A ONE-DIMENSIONAL CONVOLUTIONAL NEURAL NETWORK LAYERCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims priority to and benefit of U.S. Provisional Application No. 63 / 778,684, titled "MACHINE LEARNING INFERENCE FROM GENE EXPRESSION DATA FOR THE DETECTION OF NEUROLOGICAL AND DEVELOPMENTAL DISORDERS,” filed March 27, 2025, the entire disclosure of which is incorporated herein by reference in its entirety.FIELD
[0002] One or more embodiments relate to systems and methods for a machine learning model that includes a self-attention layer, a long short-term memory layer, a dropout layer, and a one-dimensional convolutional neural network layer. One or more embodiments relate to a machine learning (ML) model(s) for the detection of a neurological and developmental disorder(s) from gene expression data. Specifically, given an indication of the ribonucleic acid (RNA) abundance from a blood sample, the machine learning model(s) is configured to predict a likelihood the blood sample originated from an individual with the specified neurological or developmental disorder(s).BACKGROUND
[0003] Genetics and environmental exposure during development and throughout one’s life can be key determinants to susceptibility of that individual to disorders and diseases. While disease susceptibility for chromosomal disorders such as down syndrome and monogenic diseases such as sickle-cell anemia and many metabolic diseases is governed almost entirely by a Mendelian allele (enabling highly predictive diagnostic from deoxyribonucleic acid (DNA) sequences as early as embryonic stages), many diseases and disorders (e.g., human diseases and disorders) are influenced by contributions from many genes and environmental factors that knowledge of DNA mutation alone is insufficient to confer a definitive diagnostic.Attorney Docket No.: BSAI-OOl / OIWO 357321-2002
[0004] Gene products such as RNA, proteins, and their metabolites are genetic outputs that reflect complex interplay between genetic constitution of an individual and that individual’s development and environmental exposure. Therefore, the presence and abundance of gene products and metabolites can serve as diagnostic and prognostic biomarkers for predicting susceptibility, severity and remission of diseases with defined or unknown genetic contributions.
[0005] Neurological disorders such as autism spectrum disorder (ASD) can be challenging to diagnose due to involvement of hundreds of genes. The genetic contributions of these genes to the eventual manifestation of ASD are relatively small, preventing effective early diagnostic of ASD based solely on an individual’s DNA information. Current known diagnosis of ASD relies on behavioral and cognitive assessment tests administered by behavioral specialists. As such, the diagnosis is limited by age due to requirement of a child reaching a certain age for the observed behavior to be manifested, faces a long bottleneck due to shortage of staff and behavioral specialists, and time-consuming as definitive diagnosis can require multiple visits and behavioral assessment sessions.Description of the Related Art:
[0006] Three known approaches have been employed to speed up ASD diagnosis: 1) behavioral assessments aided by artificial intelligence (Al) in the form of machine learning (ML); 2) the use of genes, gene products or metabolites as diagnostic biomarkers for ASD; (3) the use / deployment of a ML model trained on large datasets of genes, gene products, or metabolites to aid diagnostic assessment.However, these approaches are limited by the use of biomarkers that represent only a subset of ASD or insufficient small dataset.
[0007] ML is increasingly used to aid doctors in medical diagnostic care. To address the bottleneck in ASD diagnosis, the company, Cognoa™, has developed Canvas Dx™, approved by the Food and Drug Administration (FDA) in 2021 as a software as a medical device (SaMD) to aid ASD diagnosis (Wall, et al “Optimizing a de novo artificial intelligence-based medical device under a predetermined change control plan: Improved ability to detect or rule out pediatric autism 2023” Intelligence-Based Medicine, 2023, doi.org / 10.1016 / j.ibmed.2023.100102; US20220344030A1). Using ML trained on video clips of kids with and without ASD, the platform serves as an aid to clinicians in their diagnostic determination of at-risk individuals. Similarly,Attorney Docket No.: BSAI-001 / 01WO 357321-2002EarliPoint™ by EarliTec Diagnostics™ approved in 2023 by the FDA 510(k) authorization via Cognoa’s Canvas Dx™ as a market predicate utilizes ML to aid ASD diagnosis (Jones, et al, “Development and Replication of Objective Measurements of Social Visual Engagement to Aid in Early Diagnosis and Assessment of Autism” JAMA Network Open, 2023, doi:10.1001 / jamanetworkopen.2023.30145; AU2023283390A1). EarliPoint™ trained on eye movements from kids with and without ASD while watching short video scenes assists clinicians in their assessment of at-risk individuals for ASD.
[0008] Rapid advances in genotyping have led to the use of array and sequencing technologies in genetic diagnostics from blood samples. For instance, Agilent™ utilizes their high-density oligo DNA tiling array technology to gain FDA 510(k) clearance for their GenetiSure Dx Postnatal Assay™ diagnostic platform to detect genetic anomalies such as chromosomal abnormalities or copy number variations linked to neurodevelopmental and dysmorphic disorders in pediatric and adult populations, and GeneDx relies on exome sequencing to identify gene variants known to associate with neurological disorders including ASD (see, e.g., U.S. Patent No. 10,450,611B2).
[0009] Alternative to DNA sequencing of genetic variants contributing to ASD includes the use of biomarker screening. NeuroQure™ offers a skin test for ASD that uses biomarkers associated with abnormal calcium (Ca2+) signaling in sporadic ASD to measure agonist-evoked Ca2+ signals in human primary skin fibroblasts (see, e.g., U.S. Patent Application Publication No. 20190079078 Al; Schmunk et al, “High-throughput screen detects calcium signaling dysfunction in typical sporadic autism spectrum disorder,” 2017 Feb. 1, Sci Rep. 7, 40740; doi: 10.1038 / srep40740. PMID: 28145469; PMCID: PMC5286408).
[0010] High throughput acquisition of molecular data has overwhelmed the ability of human experts in the respective fields to adequately make use of the data to diagnose a condition. ML has been deployed to train on large datasets and the trained model is subsequently utilized to aid clinicians and researchers in diagnostic assessment. Linus Biotechnology™ relied on ML model trained on metabolites such as metal concentration in tooth and hair of individuals to aid diagnostic determination in ASD (see, e.g., U.S. Patent Application Publication No. 20250029689-Al).SUMMARYAttorney Docket No.: BSAI-OOl / OIWO 357321-2002
[0011] In an embodiment, a training dataset associated with a first class having a condition and a second class not having the condition is received. A machine learning (ML) model is trained based on the training data to generate a trained ML model. The trained ML model includes a plurality of sets of layers, each set of layers from the plurality of sets of layers including a self-attention layer, a long short-term memory layer, a dropout layer, and a one-dimensional convolutional neural network (CNN) layer.BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 illustrates a diagram of used RNA-sequence datasets and processing workflow for machine learning training, according to an embodiment.
[0013] Figure 2 illustrates a diagram of feature selection methodology used for machine learning training, according to an embodiment.
[0014] Figure 3 illustrates equations used to transform transcripts per kilobase million (TPM) in ASD sample to binary itemsets, according to an embodiment.
[0015] Figure 4 illustrates a diagram of machine learning model, according to an embodiment.
[0016] Figure 5 illustrates an example of an output of a positive ASD sample with likelihood of association to known ASD phenotypes, according to an embodiment.
[0017] Figure 6 illustrates boxplots of relative normalized count (TPM) of a selected gene-set from a positive ASD sample versus non-ASD samples, according to an embodiment.
[0018] Figure 7 illustrates a system block diagram of a diagnosis compute device, according to an embodiment.DETAILED DESCRIPTION
[0019] Described herein are various examples. It is not intended to be limiting to a specific embodiment. One embodiment is related to a compute device(s) that utilizes (e.g., receives, stores, etc.) transcript abundance from two disparate datasets. One dataset is a large dataset of RNA abundance from blood samples of siblings with and without ASD and the other dataset is a smaller dataset of RNA abundance from various brain tissues diagnosed with neurological or degenerative disorders other than ASD. In some implementations, a ML model(s) is first trained using a subset of samples as to which are ASD positive or negative. The ML model(s) will then beAttorney Docket No.: BSAI-OOl / OIWO 357321-2002presented with a subset of samples not seen before and infer as to whether the sample is positive or negative for ASD. In some implementations, in response to determining whether the sample is positive or negative for ASD, a remedial action can occur. Examples of remedial actions include generating an alert / output indicating the positive or negative result, generating an alert / output recommending further testing, generating a personalized treatment and / or intervention plan, and / or the like. In some implementations, the remedial action includes displaying (e.g., at diagnosis compute device 700 or a compute device not shown in Fig. 7) output of upregulated and / or downregulated genes in a negative sample showing that those genes fall within a normal and / or predetermined range of the non-ASD population.
[0020] Some implementations relate to methods, systems, software, devices and / or platforms to assess individuals at-risk for neurological, developmental or behavioral impairments. For example, some implementations relate to the use of a ML model(s) upon being trained on a given training dataset containing transcript abundance of individuals with and without the tested condition (e.g., ASD), where the ML model(s) can generate an output indicating determine whether a given biological sample not present in the training dataset is associated with (e.g., has or is predicted to have) the tested condition.
[0021] In one embodiment, the biological source of transcript abundance could be RNA extracted from blood, tissue culture or any accessible tissue of choice. In another embodiment, transcript abundance is obtained or determined by DNA or RNA sequencing (RNA-seq) techniques or any other methods that insofar the output quantity of which maintains the correlational relationship with the level of RNA in the tested sample as measured by a molecular technique. In another embodiment, transcript abundance can correspond to total RNA or a specific subset or type such as messenger RNAs (mRNAs), or a subset of RNA in the tested sample. In another embodiment, transcript abundance can correspond to total RNA or a specific subset or type such as messenger RNAs (mRNAs), or a subset of RNA in the tested sample that originates from individuals known or at-risk for ASD.
[0022] RNA sequencing (RNA-seq) is a genomic tool to analyze the quantity and genetic sequences of RNA in a given biological sample. One of the outputs of this process is the quantification of transcript abundance. After assigning to an annotated DNA element (e.g., protein-coding genes, noncoding RNA genes, regulatory RNAs, etc.) in a genome, transcript abundance is first measured in absolute counts and laterAttorney Docket No.: BSAI-OOl / OIWO 357321-2002normalized into transcripts per kilobase million (TPM), reads per kilobase per million mapped reads (RPKM) or Fragments Per Kilobase per Million mapped fragments (FPKM). The most granular and lowest level of this quantification is transcript count. To represent transcript abundance, a median normalized count can be computed for each assigned DNA element. In some implementations, two datasets are used for training a ML model(s): 1) a first dataset that includes genetic samples from families that each include a child (e.g., between the ages of 2 and 13) with ASD and a child (e.g., between the ages of 2 and 13) without ASD (e.g., SFARI’s Simons Simplex Collection (SSC)) and 2) a second dataset that include RNA transcripts from various regions and structures of postmortem brain of non-ASD individuals at various ages (e.g., the Developmental Transcriptome Dataset (DTD) from BrainSpan).
[0023] The first dataset can include RNA transcripts from siblings with and without ASD (e.g., a SFARI SSC dataset from https: / / www.sfari.org / resource / simons-simplex-collection / ). The second dataset can include RNA transcripts from various regions and structures of postmortem brain of non-ASD individuals at various ages (e.g., BrainSpan DTD dataset from https: / / www.brainspan.org / ). Fig. 1 describes a diagram that summarizes the extraction methodology using SFRARI’s SSC dataset and BrainSpans’ DTD dataset.
[0024] Data preprocessing takes the extracted data a step further by implementing the feature selection methodology described herein. In some implementations, feature selection can be defined as the process of identifying the subset of influential (e.g., the most influential) input features — such as genes — from the first and / or second datasets (e.g., the SSC and DTD datasets). Prior to feature selection, a feature pruning methodology, such as Variance thresholding, can be applied. Variance thresholding can prune features which have a variance less than a predetermined value (e.g., 0.1, 0.2, 0.3, and / or the like) in the training set; this is useful because some genes provide little to no information which helps the ML model separate the ASD from the non-ASD group given that the TPMs for this gene are relatively constant throughout all samples. Therefore, in some implementations, the first and second datasets are preprocessed (e.g., to identify those genes having a variance less than a predetermined value), and the ML model is trained using the preprocessed first and second datasets (e.g., and not the first and second datasets prior to preprocessing).
[0025] In some implementations, the feature selection methodology starts with a defined binary transformation and uses the result from this transformation in anAttorney Docket No.: BSAI-001 / 01WO 357321-2002association rule mining model to identify frequent itemsets (e.g., frequent sets of genes in the ASD group). The feature selection model is applied to the training data set, and the chosen set of features are generalized to the testing set as well. Diagram in Fig. 2 provides a visual summary of the process. In Fig. 2, the combined training dataset can be a combination of a first dataset (e.g., SSC data) and a second dataset (e.g., BrainSpan data), and feature selection can be performed to generate the frequent gene-set, akin to a preprocessing step to prepare data for training and / or testing. The frequent gene-set can be, for example, a set of features (e.g., list of gene names) from the combined training dataset that is determined (e.g., by a compute device or a human) to be relevant (e.g., TPM outside a predetermined acceptable range, correlation with a target variable, mutual information with the target, low redundancy, increases model complexity, etc.) and / or determined to be the most relevant, and used to train the model (e.g., instead of training the model using the entire combined training dataset). By training a model using relevant (e.g., the most relevant) features from a dataset rather than the entire dataset, a performance of the model can be improved (e.g., reduced overfitting, faster convergence, reduced memory requirements, etc.).
[0026] In some implementations, any machine learning model described herein can be trained using a feedback loop. For example, errors (e.g., detected by a human and / or a separate machine learning model) produced by a model during training can be fed back into the model for re-training (e.g., until an accuracy rate of the machine learning model is above a predefined threshold), resulting in a more accurate model.
[0027] In some implementations, to mitigate catastrophic forgetting during training, one or more of memory -based replay, parameter regularization, elastic weight consolidation, knowledge distillation, or selective parameter freezing can be applied. For example, historical training samples, compressed representations, or intermediate feature vectors can be stored in a replay buffer and periodically reintroduced during retraining. In some implementations, penalty terms can be applied to a loss function to constrain updates to previously learned parameters.
[0028] Given a certain train-test split, the non-ASD sample is separated from the ASD. For each gene in the extracted data’s gene pool, the non-ASD sample is used to identify the median and interquartile range (IQR) for that gene. The function described in Fig. 3 is used to transform the TPMs in the ASD sample to binary itemsets. For the train-test split of a dataset, some implementations split the datasetAttorney Docket No.: BSAI-001 / 01WO 357321-2002into a predetermined ratio (e.g., 80% for training and 20% for testing). In some implementations, cross-validation is performed where different train-test splits (e.g., different 80% training and 20% testing splits) are taken such that every sample in the dataset has been used at least once in training and at least once in testing (e.g., where the final performance metrics are the average of the metrics for each of the different 80% training 20% testing splits). Cross validation is done to ensure that different training and testing sets generate nearly the same results, which further validates model performance and ensures generalization of learnt patterns.
[0029] Fig. 3 depicts an equation to determine a gene’s TPM for a sample, p and IQR are the median and interquartile range, respectively, of the gene which is derived from the non-ASD sample, a is a scaling parameter to be defined by the user (e.g., or can default to a predetermined value like 0.5). After this binary transformation, two datasets are generated: a) a dataset containing a binary response for upregulation per gene for each sample, and b) a dataset containing a binary response for downregulation per gene for each sample. In some implementations, genes which have more than 90% of their response in all ASD samples as 0 are disregarded (e.g., deleted) in a false percentage filtering step. Thus, two datasets are generated as a result with genes that are mostly neither upregulated nor downregulated in the ASD sample. This allows for the use of association rule mining algorithms in which the frequency pattern (FP)-Growth model is used to identify frequently upregulated and frequently downregulated genes in the ASD group. Both are combined into one geneset which will be the selected set with the most influential features. Robust Scaling is then applied to remove the effect of outlier TPMs. The data preprocessing and feature selection can be implemented using, for example, scikit-learn and PySpark.Machine Learning Model(s)
[0030] Some implementations use an ML model(s) that revolves around the model’s modular block consisted of several layers of convolution neural network (CNN) or recurrent neural network (RNN) described in Fig. 4. As described herein, training dataset can be first generated as a sequence of genes which are considered to be frequently upregulated or downregulated in the ASD sample. Thus, training can be primarily trying to capture relationships between elements of this sequence in relation to the response variable. Hence, in some implementations the following block can be used: self-attention layer, long short-term memory (LSTM) layer, dropout layer, and aAttorney Docket No.: BSAI-OOl / OIWO 357321-2002one-dimensional CNN layer. This block can be repeated n times (where n is a positive integer) followed by a classification head composed of a one-dimensional Average Pooling layer followed by two dense layers for classification (e.g., a dense layer and classifier layer). In some implementations, a given block includes a self-attention layer, followed by a LSTM layer, followed by a dropout layer, followed by a onedimensional CNN layer, with no additional and / or intervening layers. In some implementations, a given blocks includes a self-attention layer, a LSTM layer, a dropout layer, and a one-dimensional CNN layer, where (1) the dropout layer is not the first layer in the block (e.g., the dropout layer follows at least one of the selfattention layer, the LSTM layer, or the one-dimensional CNN layer in the block) and (2) the self-attention layer, the LSTM layer, and the one-dimensional CNN layer can be in any sequence / order (e.g., the self-attention layer can be the first, second, third, or fourth layer in the block, the LSTM layer can be the first, second, third, or fourth layer in the block, and the one-dimensional CNN layer can be the first, second, third, or fourth layer in the block). In some implementations, a ML model includes one or more modular blocks (e.g., each block including its own attention layer, LSTM layer, dropout layer, and CNN layer) followed by a classification head that includes a global average pooling layer (e.g., that (1) performs an averaging step to have an appropriate input size for the following dense layer and (2) is not trained or trainable), a dense layer with x units (where x can be any positive integer), and a dense layer with one unit. The model diagram is shown in Fig. 4.
[0031] First, the self-attention layer is a CNN that allows for the capture of dependencies and relationships within the input sequence. Said differently, the selfattention layer is configured to determine how much importance / weight to give different features. The LSTM layer is a type of RNN suitable for capturing information latent in sequential data. The RNN can capture relationships between the output of the self-attention layer. The dropout layer is used for regularization and to prevent the overfitting of the two layers that precede it. Afterwards, a CNN layer is used to aggregate information coming out of the previous layers and capture patterns provided in the sequential output of the dropout layer. This block is repeated n times, and finally, a classification head (e.g., one-dimensional global average pooling layer, dense layer, and classifier) consisting of or including the layers mentioned above is used for transforming the output of the blocks into binary classification for ASD. In short, this model can capture dependencies and relationships among frequentlyAttorney Docket No.: BSAI-OOl / OIWO 357321-2002upregulated and downregulated genes in the ASD sample of the training data for the purpose of the binary classification at hand. The model’s size is appropriate to the dataset size (which can be very small compared to existing SOTA models which often overfit the data). The model and training can be implemented using, for example, Tensorflow, and the hyperparameters used in the entire pipeline are, in an example, as follows:• Binary Transformationo a : 0.5o FP-Growth minimum confidence: 0.9o FP-Growth minimum support: 0.7• Feature Pruningo Variance Threshold: 0.1o False Percentage Threshold: 0.9• Scaling: Robust Scaler• Modelo Number of blocks: 1o LSTM units: 16o Dropout rate: 0.01o Convolution number of filters: 10o Convolution kernel size: 3o Convolution activation: Rectified Linear Unit (Relu) o Dense layer units: 64o Dense layer activation: Reluo Classification layer activation: Reluo Loss function: Binary Crossentropyo Optimizer: Adamo Training Epochs: 20o Training Batches: 32Training-testing split:• Train: 3241o Autistic: 1527o non- Autistic: 1714• Test: 811Attorney Docket No.: BSAI-OOl / OIWO 357321-2002o Autistic: 382o non-Autistic: 429The result on the test split of our dataset are as follows:• TP: 333• TN: 278• FP: 151• FN: 49• Test accuracy: 0.75• Test precision: 0.69• Test recall: 0.87Processing Results
[0032] To bridge transcript abundance or gene expression with ASD phenotypes which are of potential relevance to patients and physicians, a methodology was devised to generate phenotypic information from sample data and model inference. Alongside outputs generated by the ML model, phenotypic analysis can be provided at the level of behavioral and biological processes. At least two tools alongside the dataset can be used: Research identified ASD-Risk genes and Human Phenotype Ontology. After prediction, the upregulated and downregulated gene-sets are augmented by adding ASD-risk genes from the associated ASD-Risk gene-sets which are also upregulated and downregulated given the same threshold used for the binary transformation mentioned earlier. Those augmented gene-sets are fed into the Human Phenotype Ontology tool, and Gene Enrichment analysis is performed to provide phenotypes which are significant in the pathways within the gene-set. An example of the output of this procedure is shown in Fig. 5.
[0033] Finally, box-plots are created to showcase the relative normalized count (TPM) of a certain gene from the identified gene-sets for the test sample compared to the TPMs of the non-ASD sample such as the boxplots depicted in Fig. 6.
[0034] In some implementations, the trained ML model (e.g., trained ML model 706 of FIG. 7) can be validated using one or more datasets (e.g., publicly available datasets) that were not included in the training dataset (e.g., training data 708 of FIG.7). The validation data can include, for example, datasets from ASD studies such as those associated with discordant siblings (e.g., Tomaiuolo P, Piras IS, Sain SB, Picinelli C, Baccarin M, Castronovo P, Morelli MJ, Lazarevic D, Scattoni ML, TononAttorney Docket No.: BSAI-001 / 01WO 357321-2002G, Persico AM. RNA sequencing of blood from sex- and age-matched discordant siblings supports immune and transcriptional dysregulation in autism spectrum disorder. S ci Rep. 2023 Jan 16; 13(l):807. doi: 10.1038 / s41598-023-27378-w. PMID: 36646776; PMCID: PMC9842630). These validation data results can include sensitivity, specificity, precision and accuracy.
[0035] Some embodiments identify and exclude near-perfect gender predictors, such as those with an Area Under the Curve (AUC) score > 0.99 and those with loud gender signals as a consequence of the skewed gender ratio within the preprocessed training set; expression of these genetic elements can be muffled or treated as null from the training set. Said similarly, in some embodiments, the feature selection and training process is configured to mitigate demographic bias, such as bias associated with biological sex, that may otherwise impair model generalization. Certain genes may exhibit expression patterns that correlate almost exclusively with gender rather than with the neurological or developmental condition of interest. When training data includes a skewed gender distribution, such features could artificially inflate model performance by acting as proxies for gender rather than disease state. Accordingly, genes exhibiting unusually strong gender-predictive characteristics are identified and their contribution to training is reduced or eliminated, thereby encouraging the machine learning model to learn condition-relevant biological patterns instead of demographic confounders.
[0036] Some embodiments apply paired T-test between affected and non-affected sibling or close relative to identify biomarkers in which the null hypothesis assumes the mean difference between the transcript per millions (TPMs) of affected and unaffected samples within the same family is zero; identified biomarkers can serve as selection features during model training. Said similarly, in some embodiments, biomarker identification leverages genetically and environmentally controlled sample pairings, such as affected and unaffected siblings or close relatives, to improve sensitivity to condition-specific expression differences. By comparing transcript abundance values within related individuals rather than across unrelated populations, some implementations reduce background variation attributable to shared heredity and environment. Statistical testing can applied to these paired samples to identify genes whose expression consistently deviates between affected and unaffected relatives. Genes satisfying this criterion can be treated as informative biomarkers and incorporated as selected features during model training, thereby increasing theAttorney Docket No.: BSAI-001 / 01WO 357321-2002likelihood that the trained model captures biological signals attributable to the condition rather than inter-individual variability.
[0037] Some embodiments will utilize Cohen’s d to set effect size threshold (e.g., > 1.5, 1.4, 1.3, and / or the like) to select high-precision gene markers; these gene markers can include all or combination thereof ADGRA 3, IGFBP3, CHI3L2, FBLN2, and CR2. Said similarly, in some embodiments, candidate biomarkers are further filtered based on the magnitude of their expression differences between classes, as quantified by an effect size metric. Statistical significance alone may result in the inclusion of features whose differences, while consistent, are too small to contribute meaningfully to classification accuracy. By applying a minimum effect size threshold, some embodiments preferentially select genes exhibiting large and reliable expression shifts, which are more likely to support precise and stable inference by the trained model. This approach improves robustness and interpretability of feature selection, and supports the use of high-impact gene markers either individually or in combination during machine learning training and inference.
[0038] Figure 7 illustrates a system block diagram of a diagnosis compute device, according to an embodiment. Diagnosis compute device 700 can be any type of compute device, such as a server, desktop, laptop, tablet, phone, edge device, and / or the like. Diagnosis compute device 700 includes processor 702 operatively coupled to memory 704 (e.g., via a system bus). Optionally, processor 702 can also be coupled to a communicator (not shown in Figure 7).
[0039] The memory 704 can be, for example, a random-access memory (RAM), a memory buffer, a hard drive, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), and / or the like. The memory 704 can be configured to store, for example, data. In some instances, the memory 704 can store, for example, one or more software programs and / or code that can include instructions to cause the processor 702 to perform one or more processes, functions, and / or the like. In some embodiments, the memory 704 can include extendable storage units that can be added and used incrementally. In some implementations, the memory 704 can be a portable memory (for example, a flash drive, a portable hard disk, and / or the like) that can be operatively coupled to the processor 702. In some instances, the memory 704 can be remotely operatively coupled with the compute device. For example, a remote database device can serve as a memory and be operatively coupled to the compute device.Attorney Docket No.: BSAI-OOl / OIWO 357321-2002
[0040] The communicator can be a hardware device operatively coupled to the processor 702 and memory 704 and / or software stored in the memory 704 executed by the processor 702. The communicator can be, for example, a network interface card (NIC), a Wi-FiTM module, a Bluetooth® module and / or any other suitable wired and / or wireless communication device. Furthermore, the communicator can include a switch, a router, a hub and / or any other network device. The communicator can be configured to connect diagnosis compute device 700 to a communication network (not shown in Figure 7). In some instances, the communicator can be configured to connect to a communication network such as, for example, the Internet, an intranet, a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a worldwide interoperability for microwave access network (WiMAX®), an optical fiber (or fiber optic)-based network, a Bluetooth® network, a virtual network, and / or any combination thereof. In some instances, the communicator can facilitate receiving and / or transmitting data or files through a communication network. In some instances, received data and / or a received file can be processed by the processor 702 and / or stored in the memory 704.
[0041] The processor 702 can be, for example, a hardware based integrated circuit (IC) or any other suitable processing device configured to run and / or execute a set of instructions or code. For example, the processor 702 can be a general-purpose processor, a central processing unit (CPU), an accelerated processing unit (APU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a programmable logic array (PLA), a complex programmable logic device (CPLD), a programmable logic controller (PLC) and / or the like. The processor 702 can be operatively coupled to the memory 704 through a system bus (for example, address bus, data bus and / or control bus).
[0042] Memory 704 can include (e.g., store) trained ML model 706 and training data 708. A ML model can be trained based on training data 708 to generate trained ML model 706. Training data 708 can include, for example, any of the training data disclosed herein. Trained ML model 706 can include multiple layers. In some implementations, trained ML model 706 has the architecture / structure illustrated and described with respect to Figure 4. Trained ML model 706 can receive a representation of a biological sample, and generate an output indicating whether (or a probability that) the biological sample has or is predicted to have a condition (e.g., autism).Attorney Docket No.: BSAI-OOl / OIWO 357321-2002
[0043] Although Figure 7 illustrates a single compute device, in other implementations, additional compute devices can be used. For example, in some implementations, a first compute device trains a ML model to generate trained ML model 706, and a second compute device receives and provides an input to trained ML model 706 to generate an output.
[0044] Some implementations are related to determining whether a biological sample is associated with a tested condition (e.g., autism). A dataset of x size (e.g., where x is at least 100, 500, 1000, 2000, 4000, etc.) containing two classes corresponding to with and without a tested condition are received. The elements of the class can be numerical output corresponding to certain types of biological molecule. The counts of the elements can be normalized and aggregated. Feature selection is performed to guide the machine learning model. The machine learning model post-training determines whether a numerical output of a biological sample (e.g., a biological sample not previously seen by the ML model) belongs to one of the two classes. Some implementations further include generating a digital output of the identity and abundance of biological molecules and relevant phenotypes associated with the tested condition. In some implementations, the biological molecule is a ribonucleic acid, deoxyribonucleic acid or polypeptide. In some implementations, the numerical output preserves the measured quantity of the biological molecule from the biological sample.
[0045] In a first embodiment (e.g., which can be executed by processor 702), a training dataset (e.g., training data 708) associated with a first class having a condition and a second class not having the condition is received. A machine learning (ML) model is trained based on the training dataset to generate a trained ML model (e.g., trained ML model 706). The trained ML model includes a plurality of sets of layers. Each set of layers from the plurality of sets of layers includes a self-attention layer, a long short-term memory layer, a dropout layer, and a one-dimensional convolutional neural network (CNN) layer.
[0046] In some implementations of the first embodiment, the trained ML model further includes a one-dimensional average pooling layer following the plurality of sets of layers.
[0047] In some implementations of the first embodiment, the trained ML model further includes a one-dimensional average pooling layer following the plurality ofAttorney Docket No.: BSAI-OOl / OIWO 357321-2002sets of layers, and a plurality of dense layers following the one-dimensional average pooling layer.
[0048] Some implementations of the first embodiment further include receiving a representation of a biological sample, and inputting the representation of the biological sample to the trained ML model. The trained ML model is configured to generate an output predicting whether the biological sample belongs to the first class or the second class.
[0049] In some implementations of the first embodiment, the condition is autism.
[0050] Some implementations of the first embodiment further include receiving a preliminary training dataset, and preprocessing the preliminary training dataset to generate the training dataset.
[0051] Some implementations of the first embodiment further include receiving a preliminary training dataset associated with a plurality of features, identifying features from the plurality of features associated with a variance less than a predetermined threshold, and removing the features from the plurality of features associated with the variance less than the predetermined threshold to generate the training dataset.
[0052] In a second embodiment (e.g., which can be executed by processor 702), a representation of a biological sample is received. The representation of the biological sample is input to a trained ML model (e.g., trained ML model 706) to generate an output predicting whether the biological sample is associated with a condition. The trained ML model includes a plurality of sets of layers. Each set of layers from the plurality of sets of layers includes a self-attention layer, a long short-term memory layer, a dropout layer, and a one-dimensional convolutional neural network (CNN) layer.
[0053] In some implementations of the second embodiment, a ML model is trained using a training dataset (e.g., training data 708) associated with a first class having the condition and a second class not having the condition to generate the trained ML model.
[0054] In some implementations of the second embodiment, the trained ML model further includes a one-dimensional average pooling layer following the plurality of sets of layers.
[0055] In some implementations of the second embodiment, the trained ML model further includes a one-dimensional average pooling layer following the plurality ofAttorney Docket No.: BSAI-OOl / OIWO 357321-2002sets of layers, and a plurality of dense layers following the one-dimensional average pooling layer.
[0056] In some implementations of the second embodiment, the condition is autism.
[0057] In a third embodiment (e.g., which can be executed by processor 702), a training dataset (e.g., training data 708) associated with a first class having a condition and a second class not having the condition is received. A machine learning (ML) model is trained based on the training dataset to generate a trained ML model (e.g., trained ML model 706). The trained ML model includes a plurality of sets of layers. Each set of layers from the plurality of sets of layers includes a self-attention layer, a long short-term memory layer, a dropout layer, and a one-dimensional convolutional neural network (CNN) layer. A representation of a biological sample not included in the training dataset is received. The representation of the biological sample is input to the trained ML model. The trained ML model is configured to generate an output predicting whether the biological sample belongs to the first class or the second class.
[0058] In some implementations of the third embodiment, the trained ML model further includes a one-dimensional average pooling layer following the plurality of sets of layers.
[0059] In some implementations of the third embodiment, the trained ML model further includes a one-dimensional average pooling layer following the plurality of sets of layers, and a plurality of dense layers following the one-dimensional average pooling layer.
[0060] Some implementations of the third embodiment further include receiving a preliminary training dataset, identifying a representation of gene in the preliminary dataset having a likelihood of indicating gender outside a predetermined acceptable range, and modifying, to generate the training dataset, the preliminary training dataset by at least one or removing or modifying the representation of the gene in the preliminary training dataset.
[0061] Some implementations of the third embodiment further include receiving a first transcript count associated with a first user having the condition and a second transcript count associated with a second user not having the condition, the first user being a relative (e.g., parent, sibling, grandparent, grandchild, cousin, etc.) of the second user, comparing the first transcript count and the second transcript count to identify a biomarker, and updating, to generate the training dataset, a preliminary training dataset to include the biomarker.Attorney Docket No.: BSAI-001 / 01WO 357321-2002
[0062] Some implementations of the third embodiment further include determining, for each gene from a plurality of genes, an effect size for that gene based on a statistical comparison between samples of the first class and samples of the second class, and updating, to generate the training dataset, a preliminary dataset by (1) including genes from the plurality of genes having the effect size exceeding a predetermined threshold and (2) omitting genes from the plurality of genes having the size effect not exceed the predetermined threshold.
[0063] In some implementations of the third embodiment, determining the effect size comprises applying at least one of a paired statistical test or an effect size measure to identify genes having a statistically meaningful (e.g., being outside predefined range) difference in expression between the first class and the second class.
[0064] In some implementations of the third embodiment, the training data includes the following genes: ADGRA3, IGFBP3, CHI3L2, FBLN2, and CR2.
[0065] While various embodiments have been described above, it should be understood that they have been presented by way of example only, and not limitation. Where methods and / or schematics described above indicate certain events and / or flow patterns occurring in certain order, the ordering of certain events and / or flow patterns may be modified. While the embodiments have been particularly shown and described, it will be understood that various changes in form and details may be made.
[0066] Although various embodiments have been described as having particular features and / or combinations of components, other embodiments are possible having a combination of any features and / or components from any of embodiments as discussed above.
[0067] Some embodiments described herein relate to a computer storage product with a non-transitory computer-readable medium (also can be referred to as a non-transitory processor-readable medium) having instructions or computer code thereon for performing various computer-implemented operations. The computer-readable medium (or processor-readable medium) is non-transitory in the sense that it does not include transitory propagating signals per se (e.g., a propagating electromagnetic wave carrying information on a transmission medium such as space or a cable). The media and computer code (also can be referred to as code) may be those designed and constructed for the specific purpose or purposes. Examples of non-transitory computer-readable media include, but are not limited to, magnetic storage media such as hard disks, floppy disks, and magnetic tape; optical storage media such as CompactAttorney Docket No.: BSAI-OOl / OIWO 357321-2002Disc / Digital Video Discs (CD / DVDs), Compact Disc-Read Only Memories (CD-ROMs), and holographic devices; magneto-optical storage media such as optical disks; carrier wave signal processing modules; and hardware devices that are specially configured to store and execute program code, such as Application-Specific Integrated Circuits (ASICs), Programmable Logic Devices (PLDs), Read-Only Memory (ROM) and Random-Access Memory (RAM) devices. Other embodiments described herein relate to a computer program product, which can include, for example, the instructions and / or computer code discussed herein.
[0068] Some embodiments and / or methods described herein can be performed by software (executed on hardware), hardware, or a combination thereof. Hardware modules may include, for example, a general-purpose processor, a field programmable gate array (FPGA), and / or an application specific integrated circuit (ASIC). Software modules (executed on hardware) can be expressed in a variety of software languages (e.g., computer code), including C, C++, Java™, Ruby, Visual Basic™, and / or other object-oriented, procedural, or other programming language and development tools. Examples of computer code include, but are not limited to, microcode or micro-instructions, machine instructions, such as produced by a compiler, code used to produce a web service, and files containing higher-level instructions that are executed by a computer using an interpreter. For example, embodiments may be implemented using imperative programming languages (e.g., C, Fortran, etc.), functional programming languages (Haskell, Erlang, etc.), logical programming languages (e.g., Prolog), object-oriented programming languages (e.g., Java, C++, etc.) or other suitable programming languages and / or development tools. Additional examples of computer code include, but are not limited to, control signals, encrypted code, and compressed code.
Claims
Attorney Docket No.: BSAI-OOl / OIWO 357321-2002CLAIMSWhat is claimed is:
1. A method, comprising:receiving a training dataset associated with a first class having a condition and a second class not having the condition; andtraining a machine learning (ML) model based on the training dataset to generate a trained ML model, the trained ML model including a plurality of sets of layers, each set of layers from the plurality of sets of layers including a self-attention layer, a long short-term memory layer, a dropout layer, and a one-dimensional convolutional neural network (CNN) layer.
2. The method of claim 1, wherein the trained ML model further includes a onedimensional average pooling layer following the plurality of sets of layers.
3. The method of claim 1, wherein the trained ML model further includes a onedimensional average pooling layer following the plurality of sets of layers, and a plurality of dense layers following the one-dimensional average pooling layer.
4. The method of claim 1, further comprising:receiving a representation of a biological sample; andinputting the representation of the biological sample to the trained ML model, the trained ML model configured to generate an output predicting whether the biological sample belongs to the first class or the second class.
5. The method of claim 1, wherein the condition is autism.
6. The method of claim 1, further comprising:receiving a preliminary training dataset; andpreprocessing the preliminary training dataset to generate the training dataset.
7. The method of claim 1, further comprising:receiving a preliminary training dataset associated with a plurality of features;Attorney Docket No.: BSAI-OOl / OIWO 357321-2002identifying features from the plurality of features associated with a variance less than a predetermined threshold; andremoving the features from the plurality of features associated with the variance less than the predetermined threshold to generate the training dataset.
8. An apparatus, comprising:a memory; anda processor operatively coupled to the memory, the processor configured to:receive a representation of a biological sample; andinput the representation of the biological sample to a trained ML model to generate an output predicting whether the biological sample is associated with a condition, the trained ML model including a plurality of sets of layers, each set of layers from the plurality of sets of layers including a self-attention layer, a long short-term memory layer, a dropout layer, and a one-dimensional convolutional neural network (CNN) layer.
9. The apparatus of claim 8, wherein a ML model is trained using a training dataset associated with a first class having the condition and a second class not having the condition to generate the trained ML model.
10. The apparatus of claim 8, wherein the trained ML model further includes a one-dimensional average pooling layer following the plurality of sets of layers.
11. The apparatus of claim 8, wherein the trained ML model further includes a one-dimensional average pooling layer following the plurality of sets of layers, and a plurality of dense layers following the one-dimensional average pooling layer.
12. The apparatus of claim 8, wherein the condition is autism.
13. A non-transitory, processor-readable medium storing instructions that, when executed by a processor, cause the processor to:receive a training dataset associated with a first class having a condition and a second class not having the condition;Attorney Docket No.: BSAI-OOl / OIWO 357321-2002train a machine learning (ML) model based on the training dataset to generate a trained ML model, the trained ML model including a plurality of sets of layers, each set of layers from the plurality of sets of layers including a self-attention layer, a long short-term memory layer, a dropout layer, and a one-dimensional convolutional neural network (CNN) layer;receive a representation of a biological sample not included in the training dataset; andinput the representation of the biological sample to the trained ML model, the trained ML model configured to generate an output predicting whether the biological sample belongs to the first class or the second class.
14. The non-transitory, processor-readable medium of claim 13, wherein the trained ML model further includes a one-dimensional average pooling layer following the plurality of sets of layers.
15. The non-transitory, processor-readable medium of claim 13, wherein the trained ML model further includes a one-dimensional average pooling layer following the plurality of sets of layers, and a plurality of dense layers following the onedimensional average pooling layer.
16. The non-transitory, processor-readable medium of claim 13, wherein the non-transitory, processor-readable medium further stores instructions that, when executed by the processor, cause the processor to:receive a preliminary training dataset;identify a representation of a gene in the preliminary training dataset having a likelihood of indicating gender outside a predetermined acceptable range; and modify, to generate the training dataset, the preliminary training dataset by at least one or removing or modifying the representation of the gene in the preliminary training dataset.
17. The non-transitory, processor-readable medium of claim 13, wherein the non-transitory, processor-readable medium further stores instructions that, when executed by the processor, cause the processor to:Attorney Docket No.: BSAI-OOl / OIWO 357321-2002receive a first transcript count associated with a first user having the condition and a second transcript count associated with a second user not having the condition, the first user being a relative of the second user;compare the first transcript count and the second transcript count to identify a biomarker; andupdate, to generate the training dataset, a preliminary training dataset to include the biomarker.
18. The non-transitory, processor-readable medium of claim 13, wherein the non-transitory, processor-readable medium further stores instructions that, when executed by the processor, cause the processor to:determine, for each gene from a plurality of genes, an effect size for that gene based on a statistical comparison between samples of the first class and samples of the second class; andupdate, to generate the training dataset, a preliminary dataset by (1) including genes from the plurality of genes having the effect size exceeding a predetermined threshold and (2) omitting genes from the plurality of genes having the effect size not exceeding the predetermined threshold.
19. The non-transitory, processor-readable medium of claim 18, wherein determining the effect size comprises applying at least one of a paired statistical test or an effect size measure to identify genes having a statistically meaningful difference in expression between the first class and the second class.
20. The non-transitory, processor-readable medium of claim 18, wherein the training dataset includes ADGRA3, IGFBP3, CHI3L2, FBLN2, and CR2.