Method for identifying characteristics of a gene sequence, computer program, computer system and computer-readable storage medium

A machine learning-based method classifies gene sequences to predict expression patterns, allowing for precise gene editing and agricultural improvements by identifying and modifying causal features in gene sequences.

JP7764091B2Active Publication Date: 2025-11-05INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021181752
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-11-19
Filing Date
2021-11-08
Publication Date
2025-11-05
Estimated Expiration
2041-11-08

AI Technical Summary

Technical Problem

Current methods for predicting gene expression patterns rely heavily on experimental data and prior knowledge, lacking a comprehensive approach to classify gene sequences based on sequence features using machine learning without requiring such data.

Method used

A method utilizing a trained machine learning model to classify gene sequences by generating feature sets, modifying causal features, and determining target features for gene editing, enabling classification and manipulation of gene expression patterns.

Benefits of technology

Enables accurate classification of gene sequences associated with expression patterns, facilitating gene editing and improving agriculture through precise manipulation of gene expression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007764091000001
    Figure 0007764091000001
  • Figure 0007764091000002
    Figure 0007764091000002
  • Figure 0007764091000003
    Figure 0007764091000003
Patent Text Reader

Abstract

To classify genetic sequences according to sequence features associated with gene expression.SOLUTION: Genetic sequences are classified according to sequence features associated with gene expression by: receiving genetic sequence data; determining a genetic sequence feature set; determining a first classification for the genetic sequence feature set according to a machine learning model; determining causal features associated with the first classification for the genetic sequence according to the machine learning model; altering the causal feature set for the genetic sequence to yield an altered causal feature set; determining a second classification for the altered causal feature set according to the machine learning model, where the second classification differs from the first classification; and determining a set of target features, where the target features include causal features the altered causal feature set.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure relates generally to detecting and identifying gene sequence expression profiles. This disclosure relates specifically to identifying gene sequence features associated with gene expression. [Background technology]

[0002] Understanding gene expression (also known as the transcriptome) is essential for understanding the biological development and disease of organisms. Machine learning (ML) is used to predict transcriptome profiles using DNA sequence and / or epigenetic data. DNA sequence data typically contains transcription factor binding sites (TFBSs) and / or enhancers. These attributes are thought to contribute to the regulation of gene expression, and attributes such as DNA sequence features can be identified from existing resources that are widespread and publicly available for many species. Current methods utilize experimental gene expression data and / or prior knowledge of gene expression regulatory elements. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] US Patent Application Publication No. 2016 / 364522 [Patent Document 2] US Patent Application Publication No. 2019 / 114390 [Patent Document 3] US Patent Application Publication No. 2020 / 294627 [Patent Document 4] International Publication No. 2020 / 41204 [Non-patent literature]

[0004] [Non-Patent Document 1] AGARWAL et al., "Predicting mRNA Abundance Directly from Genomic Sequence Using Deep Convolutional Neural Networks", 2020, Cell Reports 31, 107663, May 19, 2020, 18 pages, Internet<URL : https: / / doi.org / 10.1016 / j.celrep.2020.107663> [Non-patent document 2] BEER et al., "Predicting Gene Expression from Sequence", Cell, Vol. 117, pp. 185-198, April 16, 2004 [Non-patent document 3] HAFEZ et al., "McEnhancer: predicting gene expression via semi-supervised assignment of enhancers to target genes", Genome Biology (2017) 18:199, 21 pages, DOI 10.1186 / s13059-017-1316-x [Non-patent document 4] MELL et al., "The NIST Definition of Cloud Computing", Recommendations of the National Institute of Standards and Technology, Special Publication 800-145, September 2011, 7 pages [Non-Patent Document 5] NATARAJAN et al., "Predicting cell-type-specific gene expression from regions of open chromatin," Genome Research, downloaded from genome.cshlp.org on October 7, 2020, pp. 1711-1722, Internet<URL: http: / / www.genome.org / cgi / doi / 10.1101 / gr.135129.111> . [Non-patent document 6] SINGH et al., "DeepChrome: deep-learning for predicting gene expression from histone modifications", Bioinformatics, 32, 2016, pp. i639-i648, doi: 10.1093 / bioinformatics / btw427, ECCB 2016 [Non-Patent Document 7] WILCZYNSKI et al., "Predicting Spatial and Temporal Gene Expression Using an Integrative Model of Transcription Factor Occupancy and Chromatin State", (2012), PLoS Computational Biology, December 2012, Volume 8, Issue 12, e1002798, 11 pages, doi:10.1371 / journal.pcbi.1002798 [Non-patent document 8] WU et al., "MetaCycle: an integrated R package to evaluate periodicity in large scale data", Bioinformatics, 32(21), 2016, pp. 3351-3353, doi: 10.1093 / bioinformatics / btw405, Advance Access Publication Date: 4 July 2016, Applications Note Summary of the Invention [Problem to be solved by the invention]

[0005] A method, a computer program, a computer system and a computer-readable storage medium for classifying gene sequences according to sequence features associated with gene expression are provided. [Means for solving the problem]

[0006] The following presents a summary to provide a basic understanding of one or more embodiments of the present disclosure. This summary is not intended to identify key or critical elements or to delineate the full scope of particular embodiments or the scope of the claims. Its sole purpose is to present concepts of the invention in a simplified form as a prelude to the more detailed description that is presented later. In one or more embodiments described herein, devices, methods, systems, computer-implemented methods, apparatus, or computer program products, or combinations thereof, are described that enable classification of gene sequence data for complex patterns of gene expression.

[0007] Aspects of the present invention disclose methods, systems, and computer-readable media related to classifying genetic sequences according to sequence features related to gene expression by receiving genetic sequence data, determining a genetic sequence feature set, determining a first classification for the genetic sequence feature set according to a machine learning model, determining a causal feature set associated with the first classification for the genetic sequence feature set according to the machine learning model, modifying the causal feature set for the genetic sequence to generate an modified causal feature set, determining a second classification for the modified causal feature set according to the machine learning model, wherein the second classification differs from the first classification, and defining a set of target features, wherein the target features include causal features from the modified causal feature set. [Brief explanation of the drawings]

[0008] The above and other objects, features, and advantages of the present disclosure will become more apparent through a more detailed description of several embodiments of the present disclosure in the accompanying drawings, in which like reference numerals generally refer to like components in embodiments of the present disclosure.

[0009] [Figure 1] 1 provides a schematic diagram of a computing environment according to an embodiment of the present invention; [Figure 2] 1 provides a flowchart illustrating a sequence of operations according to an embodiment of the present invention. [Figure 3] 1 illustrates a cloud computing environment according to an embodiment of the present invention. [Figure 4] 1 illustrates abstraction model layers according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0010] Some embodiments will now be described in more detail with reference to the accompanying drawings, in which embodiments of the present disclosure are shown. However, the present disclosure can be embodied in various ways and therefore should not be construed as limited to the embodiments disclosed herein.

[0011] In embodiments, one or more components of the system may utilize hardware or software, or both, to solve problems that are highly technical in nature (e.g., determining a set of gene sequence features; determining a first classification for the set of gene sequence features according to a machine learning model; determining a set of causal features for the gene sequence according to the machine learning model; modifying the set of causal features for the gene sequence to generate a modified causal feature set; determining a second classification for the modified causal feature set according to the machine learning model, where the second classification differs from the first classification; and determining a set of target features, etc.). These solutions are not abstract and cannot be performed as a set of mental acts by a human, for example, due to the processing power required to facilitate classification of gene sequences. Furthermore, some of the processing performed may be performed by a special-purpose computer to perform defined tasks related to classification of gene sequences. For example, a special-purpose computer may be utilized to perform tasks related to classification of gene sequences, etc.

[0012] Accurate classification of gene sequences provides an understanding of gene sequence attributes related to gene expression patterns. Identifying sequences associated with gene expression patterns over the course of a day (circadian rhythm) allows for the control and manipulation of these expression patterns through gene editing using tools such as Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR / Cas9). Applications include gene expression therapy and improved agriculture. The disclosed embodiments enable classification of gene sequences related to patterns of gene expression.

[0013] In embodiments, a method utilizes a trained machine learning (ML) model to classify gene sequences. The method trains the model according to the desired classification properties. As an example, to classify gene sequences or associated gene promoter sequences as either circadian or non-circadian, the method utilizes labeled data containing gene sequences known to be either circadian or non-circadian in their expression as training and test data for developing an ML classification model.

[0014] The method evaluates time-series transcriptome data for a set of genes and associated gene promoters. In embodiments, the method collects associated promoter sequences for an input gene as a set of base pairs immediately upstream from the base pair sequence of the gene. For example, the method collects 1500 base pairs upstream from the gene as the promoter sequence of the gene. The transcriptome includes messenger RNA data associated with gene / gene promoter activity. The time-series transcriptome data provides data related to changes in messenger RNA of the gene / gene promoter over an observed period of time. The transcriptome changes over time indicate changes in gene / promoter activity or gene / promoter expression over an observed period of time.

[0015] In one embodiment, transcriptome analysis of individual genes / promoters in a gene / promoter set was performed every two hours over the entire 48-hour observation period. The gene / promoter sequences used included known and publicly available gene / promoter sequences. Circadian genes exhibit regular, cyclical changes in expression over a 24-hour period, along with accompanying changes in transcriptome data. Non-circadian gene expression lacks such regular, cyclical changes in expression. This analysis resulted in a training dataset of 50,000 genes / promoters, 25,000 of which were labeled as circadian due to changes in the transcriptome data over the observed period, and an additional 25,000 were labeled as non-circadian based on the time-series transcriptome data. The method labeled the genes / promoters in the training set according to the expression data observed in the time-series transcriptome data. Genes / promoters with time-series data containing cyclic expression patterns over a 24-hour period were labeled as circadian, while genes / promoters without such cyclic expression patterns were labeled as non-circadian. Similarly, the method can be adapted to classify and label training datasets for other complex expression patterns using time-series transcriptome data. Once categorized and labeled, there is no need to regenerate the training gene sequence set.

[0016] After generating a training dataset using a time series transcriptome analysis of available gene sequences, the method processes each gene in the 50,000 gene training dataset. The method generates a set of nucleotide subsequences, or k-mers, of the gene. In an embodiment, the method utilizes k-mers of nucleotides that are 6 in length. Other k-mer lengths, e.g., 4, 8, 10, 12, and more, can be selected and used. In k-mers, the method generates a set of all possible combinations for the nucleotide choices A, T, G, and C (adenine, thymine, guanine, and cytosine). For k-mers, there are a total of 4096 possible combinations for the four nucleotide bases in the six sets.

[0017] For each possible k-mer combination, the method analyzes a training set of genes to determine the number of occurrences of the k-mer in each gene in the training data set. In embodiments, this analysis results in a matrix indicating the number of occurrences of each k-mer in each of the genes. For each gene, the matrix entries constitute a gene feature.

[0018] In an embodiment, the method counts feature occurrences across the base pair sequence of a gene and also counts feature occurrences across the base pair sequence of an associated gene promoter. The matrix contains the distribution of feature count values ​​for each gene and gene promoter. In this embodiment, the total number of possible features is doubled to 8192: 4096 possible features for genes and 4096 possible features for gene promoters.

[0019] In an embodiment, the method counts the occurrence of features across the combined sequences of genes and gene promoters. In this embodiment, the matrix contains feature count values ​​for each of the 4096 possible features.

[0020] In an embodiment, the method reduces the number of features for each gene from a possible 4096 to a smaller number, such as 100 features. As an example, the method can use a chi-squared test to identify the top 100 features from the entire set of features in the matrix.

[0021] In an embodiment, the method utilizes a classification algorithm to predict the classification of labeled data in a training set. Exemplary classification algorithms include logistic regression, Random Forest, XGBoost, decision trees, K-NN (K-nearest neighbor), Gaussian Process, LightGBM (gradient boosting), and SVM (support vector machine). The method splits the training data set using 80% of the data for training and 20% of the data for testing the developed algorithm. In this embodiment, the method utilizes a k-nearest neighbor algorithm to achieve 77% accuracy in classifying labeled training data using a k value of 2. The method can utilize other k values ​​depending on the desired accuracy in fitting and predicting the training data. The developed model relies solely on the distribution of k-mers within the training set sequences, without experimental data related to gene sequences. For example, the trained model classifies a feature set derived from input data sequences as either circadian or non-circadian. The dichotomy of classification arises from the nature of the training data set. By analogy, labeled training data associated with other complex gene expression patterns yields a model adapted to classify feature sets from input sequences as either matching or not matching the complex gene expression pattern.

[0022] In practice, the method receives genetic sequence data, processes the sequence data as described above to generate a feature set for the sequence, and passes the feature set to a classification model for analysis. The model returns the feature set and the associated genetic sequence classification.

[0023] In embodiments, a user interface, such as a graphical user interface (GUI), provides user access to the disclosed methods. The methods receive gene sequence data from a user. The user can download or otherwise provide publicly available genomic (and epigenetic, if available) resources for the species of interest, or can use private user-defined datasets. In embodiments, the methods provide links to publicly available genome databases using application program interfaces (APIs) associated with such databases. The provided gene sequence resources are in the form of genome sequences with gene annotations and / or DNA methylation and / or histone modifications, etc.

[0024] The method processes the provided sequence data and analyzes the provided data to count the occurrence count of each of the 4096 possible k-mer AGTCs, each nucleotide combination of a k-mer with six bases. In embodiments, the method utilizes epigenetic data to ignore known heavily methylated transcription factor binding sites (TFBSs) from the set of features captured in the feature matrix. Ignoring these sites reduces the number of matrix values ​​and limits the feature matrix to features / attributes associated with sequence differences associated with differential expression. TFBSs serve a practical function for expression, not as gene attributes. The method captures each feature count as a value in the matrix associated with each gene analyzed.

[0025] The method provides a matrix of features to a trained ML model for classification. The method can reduce the number of matrix values ​​from a maximum of 4096 to a smaller number, such as 100, before passing the feature set to the ML model for classification. An ML model, such as a k-nearest neighbor model, classifies each input feature set. The method provides an explanation for the classification in the form of a feature vector of the input feature set and its nearest neighbors that leads to the classification. The method compares the input feature vector with the feature vectors of the nearest neighbors, and from the comparison, identifies a set of candidate causal features (i.e., features from the input feature set that are most likely to result in the classification of the input as the final classification assigned to the input).

[0026] In an embodiment, the method uses data from a comparison of the input feature vector with the k-nearest neighbor feature vector to rank the features of the candidate causal feature set.

[0027] In an embodiment, the method selectively evolves an input gene "in silico." For each feature in the candidate causal feature set, the method selectively edits the input gene sequence and removes the candidate feature from the sequence and the sequence's Feature Set. The method then classifies the edited Feature Set. The method classifies edited features that result in a change in classification, e.g., a feature that changes a sequence from circadian to non-circadian, as members of the target Feature Set. The method compiles a complete target Feature Set of all candidate causal features that result in a change in classification after editing. The complete target Feature Set provides candidates for actual gene editing to alter the pattern of gene expression of the original input gene. By selectively removing the candidate target features using means such as CRISPR / Cas9, the gene's expression pattern will change as indicated by a change in classification of the edited and evolved sequence.

[0028] In embodiments, the final target feature set provides a means to identify genetic homologs to an input gene sequence from a first species in closely related species. As one example, a user of the method can apply classification results associated with bread wheat, Triticum aestivum, to related wheat species such as Triticum durum, or related cereal species such as barley or oat species. As another example, a user can apply gene expression classification results associated with a first subject's genome to the genomes of other subjects of the same species. Application of the disclosed embodiments to human genetic sequences presupposes that the human donor has consented or otherwise opted in to the use of their genetic sequence data by users of the disclosed methods and systems.

[0029] In an embodiment, the method maintains a candidate causal feature set for each classification of the model. In this embodiment, the method selects features from the candidate causal feature set for a first classification and adds them to input genetic sequences identified by the model as a different classification through in silico evolution. Similarly, the method selects features from the candidate causal feature set for a classification and removes them from input genetic sequences identified by the model as that classification through in silico evolution.

[0030] In an embodiment, the method begins the in silico evolution of the input sequence with the highest ranked candidate causal feature and proceeds from this highest ranked candidate to the lowest ranked candidate. In this embodiment, the method stops the in silico evolution of the candidate causal features after a threshold number of consecutively ranked candidate causal features fail to result in a change in classification, e.g., the method stops the in silico evolution of the input gene sequence with the candidate causal features after each of 10 consecutively ranked candidates fails to result in a change in classification.

[0031] FIG. 1 provides a schematic diagram of exemplary network resources relevant to implementing the disclosed invention. The invention can be implemented in any of the disclosed element processors processing instruction streams. As shown, a networked client device 110 wirelessly connects to a server subsystem 102. A client device 104 wirelessly connects to the server subsystem 102 via a network 114. Client devices 104 and 110 include a gene sequence classification program (not shown) along with sufficient computing resources (processor, memory, network communication hardware) to execute the program. Client devices 104 and 110 function as user interface devices, allowing users to provide input gene sequence and epigenetic data to the disclosed methods and systems. Client devices 104 and 110 also function as output devices, allowing the disclosed embodiments to provide output data to a user.

[0032] As shown in Figure 1, server subsystem 102 includes server computer 150. Figure 1 illustrates a block diagram of components of server computer 150 within networked computer system 1000 in accordance with an embodiment of the present invention. It should be understood that Figure 1 is illustrative of only one implementation and is not intended to imply any limitation with respect to the environments in which different embodiments may be implemented. Many modifications to the depicted environments may be made.

[0033] Server computer 150 may include processor 154, memory 158, persistent storage 170, communication unit 152, input / output (I / O) interface 156, and communication fabric 140. Communication fabric 140 provides communication between cache 162, memory 158, persistent storage 170, communication unit 152, and input / output (I / O) interface 156. Communication fabric 140 may be implemented with any architecture designed to pass data or control information, or both, between a processor (e.g., a microprocessor, communication and network processor, etc.), system memory, peripherals, and any other hardware components in the system. For example, communication fabric 140 may be implemented with one or more buses.

[0034] Memory 158 and persistent memory 170 are computer-readable storage media. In this embodiment, memory 158 includes random access memory (RAM) 160. In general, memory 158 may include any suitable volatile or non-volatile computer-readable storage medium. Cache 162 is a high-speed memory that improves the performance of processor 154 by retaining recently and nearly recently accessed data from memory 158.

[0035] Program instructions and data used to implement embodiments of the present invention, such as a gene sequence classification program 175, are stored in persistent storage 170 for execution and / or access by one or more of the respective processors 154 of server computer 150 via cache 162. In this embodiment, persistent storage 170 includes a magnetic hard disk drive. Alternatively, or in addition to a magnetic hard disk drive, persistent storage 170 may include a solid-state hard drive, a semiconductor storage device, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, or any other computer-readable storage medium capable of storing program instructions or digital information.

[0036] The media used by persistent storage 170 may also be removable. For example, a removable hard drive may be used for persistent storage 170. Other examples include optical and magnetic disks, thumb drives, and smart cards that are inserted into a drive for transfer to another computer-readable storage medium that is also part of persistent storage 170.

[0037] Communications unit 152 provides for communication with other data processing systems or devices, including resources of client computing devices 104 and 110, in these examples. In these examples, communications unit 152 includes one or more network interface cards. Communications unit 152 may provide communications using either or both physical and wireless communications links. Software distribution programs, as well as other programs and data used to implement the present invention, may be downloaded to persistent storage 170 of server computer 150 through communications unit 152.

[0038] The I / O interface 156 allows for the input and output of data with other devices that may be connected to the server computer 150. For example, the I / O interface 156 may provide a connection to an external device 190, such as a keyboard, keypad, touchscreen, microphone, digital camera, or any other suitable input device, or combination thereof. The external device 190 may also include portable computer-readable storage media, such as thumb drives, portable optical or magnetic disks, and memory cards. Software and data used to implement embodiments of the present invention, such as the gene sequence classification program 175 on the server computer 150, may be stored on such portable computer-readable storage media and loaded into the persistent storage 170 via the I / O interface 156. The I / O interface 156 is also connected to a display 180.

[0039] Display 180 provides a mechanism for displaying data to a user and may be, for example, a computer monitor, or may function as a touch screen, such as the display of a tablet computer.

[0040] Figure 2 provides a flowchart 200 illustrating exemplary activities associated with practicing the present disclosure. After the program begins, a user provides genetic sequence data obtained from public sources, private sources, or a combination of public and private sources to a genetic sequence classification program 175. The input data includes genomic sequence data 214, as well as gene annotation and DNA methylation or histone modification data, or both. The input data can further include prior domain knowledge of the genomic sequence, e.g., epigenetic data 218, such as heavily methylated TFBS sites in the sequence.

[0041] At 220, the method of genetic sequence classification program 175 processes the input genetic data 214 to generate a matrix of sequence features for the input data. The sequence features include data regarding the distribution of possible six-base k-mers within the genomic sequences of the input data 214.

[0042] At 230, the method of gene sequence classification program 175 optionally utilizes epigenetic data 218 to reduce the number of entries in the feature matrix from 220. The method removes features associated with known heavily methylated TFBS sites from the matrix or reduces the associated matrix entry value to zero.

[0043] At 240, the method of gene sequence classification program 175 classifies or predicts a classification for either the input gene sequence feature set from 220 or the epigenetic information modified feature set from 230. The method utilizes a machine learning model trained to classify gene sequences using a training dataset of labeled gene sequence data associated with the desired classification. As an example, a machine learning model trained using labeled gene sequences associated with each of circadian and non-circadian gene sequences provides a prediction of either circadian or non-circadian for the provided input feature set.

[0044] At 250, the method of genetic sequence classification program 175 uses the classification model description for classification to generate a candidate causal feature set. This set includes sequence features of the input genetic sequence that most likely led to the model's classification of that input sequence. In an embodiment, the method ranks the members of the candidate feature set from most likely to least likely.

[0045] At 260, the method of genetic sequence classification program 175 selectively compiles the input genetic sequence and associated input sequence feature set from either 220 or 230. For each member of the candidate causal feature set, the method removes the feature from the input genetic sequence and associated input sequence feature set.

[0046] At 270, the genetic sequence classification program 175 method uses the trained machine learning model to predict or classify the compiled input Feature Set. The method passes 280 input features whose removal would change the classification to the target Feature Set. The method returns to 260 and edits each candidate causal feature in turn, editing the input sequence and associated Feature Set with only a single candidate causal feature each iteration.

[0047] In an embodiment, the method provides a general candidate causal feature set for each possible classification of the machine learning model. In this embodiment, at 260, the method removes candidate causal features from the input sequence and input features from the general candidate causal feature set for classification of the input sequence, or adds candidate causal features from the general candidate causal feature set for a different classification. As an example, for an input sequence classified as circadian, the method adds candidate causal features from the general candidate causal feature set for non-circadian sequences, or removes candidate causal features from the candidate causal feature set for the input sequence and input feature set. In this embodiment, the method refines a target feature set for each possible classification of the machine learning classification model. (Features added from the general candidate causal feature set that result in a change in classification are added to the associated target feature set for that classification; e.g., the method adds features from the general candidate causal feature set added to a circadian sequence that result in reclassification of that sequence to non-circadian to the target feature set for non-circadian sequences.)

[0048] The method provides the user, via user interface 210, with a set of target features from 280. The user can utilize the target features to selectively edit actual gene sequences for gene therapy related to altering gene expression patterns, or to modify gene expression in plant species to improve agricultural production.

[0049] In embodiments, execution of the disclosed methods requires computational resources beyond those locally available to the user, and in this embodiment, the user connects to networked resources, including edge cloud and cloud resources, to enable timely execution of the methods.

[0050] Although this disclosure includes detailed descriptions of cloud computing, it should be understood that implementation of the teachings described herein is not limited to cloud computing environments. Rather, embodiments of the invention may be practiced in connection with any other type of computing environment now known or later developed.

[0051] Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal administrative effort or interaction with a service provider. This cloud model includes at least five characteristics, at least three service models, and at least four deployment models.

[0052] The features are as follows:

[0053] On-Demand Self-Service: Cloud consumers can automatically and unilaterally provision computing capacity, such as server time and network storage, as needed, without the need for human interaction with the service provider.

[0054] Broad network access: Functionality is available over the network and accessed through standard mechanisms that facilitate use by heterogeneous thin or thick client platforms (e.g., cell phones, laptops, and PDAs).

[0055] Resource Pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, dynamically allocating and reallocating different physical and virtual resources on demand. Consumers are location-independent in that they generally have no control or knowledge of the exact location of the resources provided, although they may be able to identify a location at a higher level of abstraction (e.g., country, state, or data center).

[0056] Rapid Elasticity: Capabilities can be rapidly provisioned and rapidly scaled out, and rapidly released and rapidly scaled in, rapidly and elastically, in some cases automatically. To the consumer, these capabilities available for provisioning often appear unlimited, available for purchase in any quantity at any time.

[0057] Service Metering: Cloud systems automatically control and optimize resource usage by using metering capabilities at some level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both providers and consumers of the services used.

[0058] The service model is as follows:

[0059] Software as a Service (SaaS): The ability to offer consumers the ability to use a provider's applications running on a cloud infrastructure. These applications are accessible from a variety of client devices through a thin-client interface such as a web browser (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even individual application functions, with the possible exception of limited user-specific application configuration settings.

[0060] Platform as a Service (PaaS): The capability offered to consumers to deploy consumer-created or acquired applications, created using programming languages ​​and tools supported by the provider, onto a cloud infrastructure. The consumer does not manage or control the underlying cloud infrastructure, such as the network, servers, operating systems, or storage, but does have control over the deployed applications and, in some cases, the application-hosting environment configuration.

[0061] Infrastructure as a Service (IaaS): The capability offered to consumers to provision processing, storage, network, and other basic computing resources on which they can deploy and run any software, which may include operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but do have control over the operating systems, storage, deployed applications, and possibly limited control over the selection of network components (e.g., host firewalls).

[0062] The deployment model is as follows:

[0063] Private Cloud: Cloud infrastructure is operated exclusively for an organization. This cloud infrastructure can be managed by that organization or a third party and can be on-premise or off-premise.

[0064] Community Cloud: Cloud infrastructure is shared by several organizations to support a specific community with common concerns (e.g., mission, security requirements, policies, and compliance considerations). The cloud infrastructure can be managed by the organizations or a third party and can reside on-premises or off-premises.

[0065] Public Cloud: Cloud infrastructure is available to the general public or large industry groups and is owned by organizations that sell cloud services.

[0066] Hybrid Cloud: A cloud infrastructure is a blend of two or more clouds (private, community, or public) that remain unique entities but are tied together by standardized or proprietary technologies that enable data and application portability (e.g., cloud bursting for load balancing between clouds).

[0067] A cloud computing environment is a service oriented environment that focuses on statelessness, low coupling, modularity, and semantic interoperability. At the heart of cloud computing is an infrastructure that includes a network of interconnected nodes.

[0068] Referring now to FIG. 3, an exemplary cloud computing environment 50 is shown. As shown, the cloud computing environment 50 includes one or more cloud computing nodes 10 that can communicate with local computing devices used by cloud consumers, such as, for example, a personal digital assistant (PDA) or mobile phone 54A, a desktop computer 54B, a laptop computer 54C, or a computer system 54N, or any combination thereof. The nodes 10 can communicate with each other. These nodes can be physically or virtually grouped in one or more networks (not shown), such as a private cloud, a community cloud, a public cloud, or a hybrid cloud, or any combination thereof, as described above. This enables the cloud computing environment 50 to provide infrastructure as a service, a platform as a service, or software as a service, or any combination thereof, without requiring the cloud consumer to maintain resources on their local computing device. It will be understood that the types of computing devices 54A-N shown in FIG. 3 are intended to be exemplary only, and that computing node 10 and cloud computing environment 50 can communicate with any type of computerized device over any type of network or over a network-addressable connection (e.g., using a web browser), or both.

[0069] Referring now to Figure 4, a set of functional abstraction layers provided by cloud computing environment 50 (Figure 3) is shown. It should be understood in advance that the components, layers, and functions shown in Figure 4 are intended to be merely exemplary, and embodiments of the present invention are not limited thereto. As shown, the following layers and corresponding functions are provided:

[0070] Hardware and software layer 60 includes hardware and software components. Examples of hardware components include mainframe 61, RISC (Reduced Instruction Set Computer) architecture-based server 62, server 63, blade server 64, storage device 65, and network and network components 66. In some embodiments, software components include network application server software 67 and database software 68.

[0071] The virtualization layer 70 provides an abstraction layer through which the following examples of virtual entities can be provided: virtual servers 71, virtual storage 72, virtual networks including virtual private networks 73, virtual applications and operating systems 74, and virtual clients 75.

[0072] In one example, the management layer 80 can provide the following functions: Resource provisioning 81 provides dynamic procurement of computing and other resources utilized to execute tasks within the cloud computing environment. Metering and pricing 82 provides cost tracking as resources are utilized within the cloud computing environment and billing or invoicing for the consumption of these resources. In one example, these resources can include application software licenses. Security provides identity verification for cloud consumers and tasks and protection for data and other resources. User portal 83 provides access to the cloud computing environment for consumers and system administrators. Service level management 84 provides allocation and management of cloud computing resources to ensure required service levels are met. Service level agreement (SLA) planning and fulfillment 85 provides pre-provisioning and procurement of cloud computing resources to anticipate future requirements according to SLAs.

[0073] The workload tier 90 provides examples of functions that can utilize a cloud computing environment. Examples of workloads and functions that can be provided from this tier include mapping and navigation 91, software development and lifecycle management 92, virtual classroom instruction delivery 93, data analysis processing 94, transaction processing 95, and gene sequence classification programs 175.

[0074] The present invention may be embodied as a system, method, or computer program product, or combination thereof, at any possible level of technical detail. The present invention may be advantageously implemented in any system, single or parallel, that processes instruction streams. The computer program product may include computer-readable storage medium(s) having computer-readable program instructions for causing a processor to perform aspects of the present invention.

[0075] A computer-readable storage medium may be any tangible device capable of holding and storing instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, or semiconductor storage device, or any suitable combination of the above. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile discs (DVDs), memory sticks, mechanically encoded devices such as punch cards or ridge structures in grooves having instructions recorded thereon, and any suitable combination of the above. As used herein, computer-readable storage media is not to be construed as transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses through a fiber optic cable), or electrical signals sent through wires.

[0076] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or storage device over a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network can include copper cables, optical fibers, wireless networks, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in a computer-readable storage medium within the respective computing / processing device.

[0077] The computer-readable program instructions for carrying out the operations of the present invention may be source or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or object-oriented programming languages ​​such as Smalltalk, C++, and conventional procedural programming languages ​​such as the "C" programming language or similar programming languages. The computer-readable program instructions may run entirely on the user's computer, partially on the user's computer as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) can execute computer readable program instructions to individualize the electronic circuitry by utilizing state information in the computer readable program instructions to implement aspects of the present invention.

[0078] Aspects of the present invention will be described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0079] These computer-readable program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine, whereby the instructions, executed by the processor of the computer or other programmable data processing apparatus, create means for performing the functions / operations specified in one or more blocks of the flowcharts and / or block diagrams. These computer program instructions can also be stored in a computer-readable medium that can direct a computer, programmable data processing apparatus, or other device, or combination thereof, to function in a particular manner, whereby the instructions stored in the computer-readable medium can include an article of manufacture containing instructions that implement aspects of the functions / operations specified in one or more blocks of the flowcharts and / or block diagrams.

[0080] The computer program instructions may be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to generate a computer-implemented process, whereby the instructions running on the computer, other programmable apparatus, or other device perform the functions / operations specified in one or more blocks of the flowcharts and / or block diagrams.

[0081] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions shown in the blocks may occur in a different order than that shown in the figures. For example, two blocks shown in succession may in fact be executed substantially simultaneously, or the blocks may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or a combination of dedicated hardware and computer instructions.

[0082] References herein to "one embodiment," "one embodiment," "exemplary embodiment," and the like indicate that the described embodiment may include one or more particular features, structures, or characteristics; however, it should be understood that such particular features, structures, or characteristics may or may not be common to every single embodiment of the invention disclosed herein. Moreover, such phrases do not necessarily refer to any one particular embodiment per se. Thus, if one or more particular features, structures, or characteristics are described in connection with one embodiment, it is believed to be within the knowledge of one of ordinary skill in the art to affect such one or more features, structures, or characteristics in connection with other embodiments, where applicable, whether or not explicitly stated.

[0083] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. Furthermore, it will be understood that the terms "comprise" and / or "comprising," when used herein, specify the presence of stated features, integers, steps, operations, elements, or components, or combinations thereof, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, or groups, or combinations thereof.

[0084] The description of various embodiments of the present invention has been presented for purposes of explanation, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the present invention. The terms used herein have been selected to best explain the principles, practical applications, or technical improvements of the embodiments over those found in the market, or to enable those skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for classifying gene sequences according to sequence features associated with gene expression by computer information processing, comprising: receiving, by one or more computer processors, genetic sequence data; determining, by the one or more computer processors, a gene sequence feature set based on the received gene sequence data, the gene sequence feature set being a set of gene sequence features based on the number of occurrences of k-mers in each gene; determining, by the one or more computer processors, a first classification for the gene sequence feature set according to a machine learning model; determining, by the one or more computer processors, a causal feature set associated with the first classification for the gene sequence feature set by identifying features in the gene sequence feature set that are most likely to cause the first classification obtained by the machine learning model; generating, by the one or more computer processors, a modified Causal Feature Set by removing the Causal Feature Set from the Gene Sequence Feature Set; determining, by the one or more computer processors, a second classification for the modified causal feature set according to the machine learning model, the second classification being different from the first classification; and determining, by the one or more computer processors, a set of target features, the target features including causal features that are features in the modified set of causal features that caused the change from the first classification to the second classification; A method comprising:

2. The method of claim 1 , wherein determining the set of gene sequence features comprises determining the set of gene sequence features according to epigenetic data.

3. Determining the gene sequence feature set comprises: determining a set of all possible combinations of gene sequence features; determining the distribution of each possible combination of gene sequence features within said gene sequence; 3. The method of claim 1 or claim 2, comprising:

4. 4. The method of claim 1, wherein determining a first classification for the gene sequence feature set according to the machine learning model comprises determining a circadian / non-circadian classification for the gene sequence.

5. 5. The method of claim 1, further comprising identifying, by the one or more computer processors, genetic homologies to the gene sequence in related species according to the set of target features.

6. 6. The method of any one of claims 1 to 5, further comprising identifying, by the one or more computer processors, editing candidates within the gene sequence according to the set of target features, wherein the editing candidates are associated with altering expression of the gene sequence.

7. 7. The method of claim 1, further comprising ranking the set of target features according to their prediction of gene sequence expression.

8. 1. A computer program for classifying genetic sequences according to sequence features associated with the genetic sequences, comprising: program instructions for receiving genetic sequence data; program instructions for determining, based on the received gene sequence data, a gene sequence feature set, the set of gene sequence features being based on the number of occurrences of k-mers in each gene; program instructions for determining a first classification for the gene sequence feature set according to a machine learning model; program instructions for determining a causal feature set associated with the first classification for the gene sequence feature set by identifying features in the gene sequence feature set that are most likely to result in the first classification obtained by the machine learning model; program instructions for generating a modified causal Feature Set by removing the causal Feature Set from the gene sequence Feature Set; program instructions for determining a second classification for the modified causal feature set according to the machine learning model, the second classification being different from the first classification; and program instructions for determining a set of target features, the target features including causal features that caused a change from the first classification to the second classification in the changed set of causal features; a computer program comprising:

9. 9. The computer program product of claim 8, wherein the program instructions for determining the set of gene sequence features comprise program instructions for determining the set of gene sequence features according to epigenetic data.

10. The program instructions for determining the gene sequence feature set include: program instructions for determining a set of all possible combinations of genetic sequence features; program instructions for determining the distribution of each possible combination of genetic sequence features within said genetic sequence; 10. A computer program according to claim 8 or claim 9, comprising:

11. 11. The computer program product of claim 8, wherein the program instructions for determining a first classification for the gene sequence feature set according to the machine learning model comprise program instructions for determining a circadian / non-circadian classification for the gene sequence.

12. 12. The computer program of claim 8, further comprising program instructions for identifying genetic homologies to the gene sequence in related species according to the set of target features.

13. 13. The computer program of claim 8, further comprising program instructions for identifying candidate editing sites within the gene sequence according to the set of target features, the candidate editing sites being associated with altered expression of the gene sequence.

14. 14. The computer program of claim 8, further comprising program instructions for ranking the set of target features according to their prediction of gene sequence expression.

15. 1. A computer system for classifying genetic sequences according to genetic sequence characteristics associated with the genetic sequences, comprising: one or more computer processors; one or more computer readable storage devices; stored program instructions on said one or more readable storage devices for execution by said one or more computer processors; wherein the stored program instructions include: program instructions for receiving genetic sequence data; program instructions for determining, based on the received gene sequence data, a gene sequence feature set, the set of gene sequence features being based on the number of occurrences of k-mers in each gene; program instructions for determining a first classification for the gene sequence feature set according to a machine learning model; program instructions for determining a causal feature set associated with the first classification for the gene sequence feature set by identifying features in the gene sequence feature set that are most likely to result in the first classification obtained by the machine learning model; program instructions for generating a modified causal Feature Set by removing the causal Feature Set from the gene sequence Feature Set; program instructions for determining a second classification for the modified causal feature set according to the machine learning model, the second classification being different from the first classification; and program instructions for determining a set of target features, the target features including causal features that caused a change from the first classification to the second classification in the changed set of causal features; 1. A computer system comprising:

16. 16. The computer system of claim 15, wherein the program instructions for determining the set of gene sequence features comprise program instructions for determining the set of gene sequence features according to epigenetic data.

17. The program instructions for determining the gene sequence feature set include: program instructions for determining a set of all possible combinations of genetic sequence features; program instructions for determining the distribution of each possible combination of genetic sequence features within said genetic sequence; 17. A computer system according to claim 15 or claim 16, comprising:

18. 18. The computer system of claim 15, wherein the program instructions for determining a first classification for the gene sequence feature set according to the machine learning model include program instructions for determining a circadian / non-circadian classification for the gene sequence.

19. 19. The computer system of claim 15, wherein the stored program instructions further comprise program instructions for identifying genetic homologies to the gene sequence in related species according to the set of target features.

20. 20. The computer system of claim 15, wherein the stored program instructions further comprise program instructions for identifying candidate editing sites within the gene sequence according to the set of target features, the candidate editing sites being associated with altered expression of the gene sequence.

21. A computer-readable storage medium having stored thereon a computer program according to any one of claims 8 to 14.

Citation Information

Patent Citations

  • Methods and machine learning for disease diagnosis

    CA3117218A1

  • Systems and methods for classifying, prioritizing and interpreting genetic variants and therapies using a deep neural network

    US20160364522A1

  • Drug repurposing based on deep embeddings of gene expression profiles

    US20190114390A1

  • Optimization of Gene Sequences for Protein Expression

    US20200294627A1

  • Artificial intelligence analysis of RNA transcriptome for drug discovery

    WO2020041204A1