ARCHITECTURE FOR ANALYSIS OF GENOME DATA
Patent Information
- Application Number
- DE602016092251
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2015-12-18
- Filing Date
- 2016-12-16
- Publication Date
- 2025-05-14
- Estimated Expiration
- 2036-12-16
AI Technical Summary
Current genomic data analysis tools are inadequate for efficiently processing and interpreting the vast amounts of data produced by next-generation sequencing (NGS) technologies, which are constrained by the size of DNA fragments and plagued by errors.
A collaborative architecture for genomic data analysis that utilizes a multi-tenant, multi-instance system to facilitate calculations on genomic data, allowing for precise and interpreted genomic analyses while ensuring user confidentiality and autonomy.
The architecture enables efficient exploitation of genomic data from NGS technologies, providing a collaborative environment for precise genomic analyses, simplified access to complex results, and enhanced user confidentiality and autonomy.
Description
TECHNICAL FIELD OF THE INVENTION
[0001] One aspect of the present invention relates to an architecture for analyzing genomic data, such as short sequences of ribonucleic acid (RNA) or deoxyribonucleic acid (DNA). TECHNOLOGICAL BACKGROUND OF THE INVENTION
[0002] DNA sequencers emerged in the 1970s with the mission of digitizing the DNA of living organisms' cells. These sequencers notably enabled the sequencing of the first human genome during the 1990s. At that time, this sequencing required more than 10 years of work and approximately $1 billion. Recently, a major technological advance has resulted in the emergence of a new generation of very high-throughput sequencers (also known by the acronym NGS), and has changed the landscape of cellular and molecular biology. The price of sequencing has become very accessible, and the dizzying increase in the production capacity of these sequencers has made it possible to sequence entire genomes in just a few days.
[0003] Today, this technological revolution is revolutionizing biology and medicine. In cancerology, for example, it is now possible to very quickly sequence all the information contained in a patient's cancer cells, with the aim of searching for molecular markers that will be used for diagnosis, prognosis, monitoring of residual disease, or response to treatment. However, sequencers have always been limited, and continue to be limited, by the size of the sequenced DNA fragments. Preparing a sample involves cutting it into short strands of DNA (of only a few hundred nucleotides) so that a sequencer can very quickly produce hundreds of millions of short DNA (or RNA) sequences in text form: these are called reads. These reads cover the entire starting genetic material (for example, the genome or transcriptome of an individual).But the information they produce is cut into a puzzle of millions of completely disorganized reads, some of which contain errors produced during sequencing or sample preparation. In other words, the sequences produced by NGS are constrained and do not lead to easily usable results.
[0004] Biologists and clinicians are faced with a deluge of data (Big Data) that must be stored, sorted, structured, and interpreted. However, biologists, or even clinicians, currently lack any tools adapted to exploiting this data. GENERAL DESCRIPTION OF THE INVENTION
[0005] In this context, the present invention aims to provide an architecture for the analysis of genomic data making it possible to efficiently exploit genomic data obtained using a very high throughput sequencer.
[0006] To this end, the invention relates to an architecture for the analysis of genomic data as defined in independent claim 1.
[0007] By means of the invention, a user can perform, via a calculation system, calculations on genomic data contained for example in the knowledge base. The results of the calculations carried out can be stored in the knowledge base and be accessible to other users, who can for example enrich these results by comparing them with their own.
[0008] This architecture therefore forms a collaborative working environment for data analysis from new very high resolution sequencing (NGS) technologies. This collaborative working environment allows for precise and interpreted genomic analyses of a pathology or a patient, in relation to a biological and / or medical question, while providing simplified, fully autonomous access to complex results.
[0009] Furthermore, it should be noted that the architecture is based on: a multi-tenant system because multiple users can use a single compute node of a single compute system, and a multi-instance system because a user can use multiple compute nodes of a compute system.
[0010] This hybrid architecture makes it possible to get the most out of both systems and thus guarantee confidentiality for users who consume the most computing resources.
[0011] In a non-limiting implementation, the architecture comprises a plurality of shared spaces, each of the shared spaces being constructed and arranged: to contain metadata and results from collaborative calculations carried out by the calculation systems of an authenticated entity comprising 1 to n users, and to be accessible in read mode only to users of the authenticated entities and benefiting from a right to consult said space.
[0012] In a non-limiting implementation, the architecture comprises a common space, said common space being accessible to all users of the authenticated entities of the architecture.
[0013] In a non-limiting implementation, the computational system is constructed and arranged to classify short digital sequences of biological material.
[0014] In a non-limiting implementation, the computing system is constructed and arranged to aggregate short digital sequences of biological material of the same class.
[0015] In a non-limiting implementation, the computational system is constructed and arranged to identify a biomarker sequence answering a biological question.
[0016] In a non-limiting implementation, the architecture comprises a plurality of resources, said resources being formed by: A system for identifying an entity and the elements it manipulates, A system for identifying a user, A system for identifying a calculation associated with at least one user and one entity, A system for correlating calculations carried out, A system for logically grouping calculation identifiers associated with a user and an entity, A system for associating meta-information with data in the knowledge base.
[0017] In a non-limiting implementation, the knowledge base is formed by a database ring.
[0018] In a non-limiting implementation, the interface for accessing the knowledge base is a REST API. BRIEF DESCRIPTION OF THE FIGURE
[0019] Other characteristics and advantages of the invention will emerge clearly from the description given below, for information purposes only and in no way limiting, with reference to the attached figure 1 illustrating, in a schematic manner. DETAILED DESCRIPTION OF AT LEAST ONE EMBODIMENT OF THE INVENTION
[0020] Figure 1 illustrates an exemplary embodiment of an architecture 100 for the analysis of genomic data according to a non-limiting aspect of the invention.
[0021] This architecture 100 comprises a plurality of application nodes 101. Each of the application nodes 101 comprises a calculation system 102 comprising at least one calculation node, the calculation system 102 being constructed and arranged to carry out genomic calculations.
[0022] When the computing system 102 is requested by a user to perform a computation requiring significant resources, the computing system 102 can implement several computing nodes. It is recalled that when an action is assigned to a computing node, this is in fact performed by a microprocessor of the computing node controlled by instruction codes stored in a memory of the computing node. If an action is assigned to an application, this is in fact performed by a microprocessor of the computing node into a memory from which the instruction codes corresponding to the application are loaded.
[0023] To perform these genomic calculations, the computing system 102 is constructed and arranged to load bioinformatics tools for organizing and interpreting biomarker sequences. In other words, calculation means in particular organizing and interpreting biomarker sequences. The bioinformatics tools that can be loaded by the computing system 102 can, for example, be formed by different types of software such as the CRAC software, the CracTools software, or the CRAC & the CT software suite.
[0024] In a non-limiting implementation, the computing system 102 can be constructed and arranged to classify short digital sequences of biological material. To do this, the computing system 102 can, for example, implement the software called CRAC which is specialized in the processing of RNA-Seq sequences and makes it possible to classify the reads. RNA-Seq are obtained by very high-throughput sequencing of RNA.
[0025] In a non-limiting implementation, the calculation system 102 is further constructed and arranged to aggregate short digital sequences of biological material of the same class. To do this, the calculation system 102 can for example implement the software called CracTools which is specialized in the post-processing of reads classified by means of the CRAC software and makes it possible to aggregate the reads of the same class to identify a biomarker sequence.
[0026] In another non-limiting implementation, the computing system 102 is further constructed and arranged to identify a biomarker sequence answering a biological question. To do this, the computing system 102 may for example implement the software suite called CRAC & the CT which connects the CRAC and CracTools software in order to identify biomarker sequences answering a specific biological question.
[0027] Furthermore, each of the application nodes 101 comprises a human-machine interface 103 constructed and arranged to communicate with the computing system 102 of the application node 101. This human-machine interface 103 can be presented in the form of a SaaS type web application. This web application makes it possible to interface the user with a computing system 102 making it possible to organize and interpret the biomarker sequences but also to be able to consult and edit a knowledge base.
[0028] In order to be able to consult and edit the knowledge base, the architecture 100 also comprises an access interface 104 to a knowledge base 105, the access interface 104 to the knowledge base 105 is for example formed by a REST API.
[0029] The knowledge base 105 is constructed and arranged to communicate with the plurality of application nodes 101 by means of the access interfaces 104. The knowledge base 105 notably comprises genomic data.
[0030] Furthermore, the knowledge base 105 comprises a plurality of private spaces 106, each of the private spaces 106 being constructed and arranged: to contain metadata and results from a calculation carried out by the calculation system 102 of an authenticated entity comprising 1 to n user(s), and to be accessible in reading and writing only to users of the authenticated entity.
[0031] Each entity comprising one or more users therefore has a private space 106 in which the resources (i.e. the calculation results) are hidden from the other entities of the architecture 100 in order to meet the need for confidentiality that several users would like. It is therefore not possible for other users of other entities to consult them.
[0032] This private base 106 contains all the work data of the user or users of an authenticated entity. For this purpose, the data of the private space 106 is encrypted by the encryption key of the user's entity. Thus, only the users of the entity can decrypt the content of their private space 106 in the knowledge base 105.
[0033] The knowledge base 105 also comprises a plurality of shared spaces 107, each of the shared spaces 107 being constructed and arranged: to contain metadata and results from collaborative calculations carried out by the calculation systems 102 of an authenticated entity comprising 1 to n user(s), and to be accessible in reading and writing only to users of the authenticated entities and benefiting from a right to consult said space.
[0034] Each entity therefore has a shared space 107 to which it is possible to invite other users from one or more other entities in order to carry out private but collaborative work. The shared space 107 contains all the common data within the framework of a collaborative project between several users from one or more other entities and thus allows data exchanges between users belonging to different entities but working on the same project.
[0035] In order to ensure the security of the stored data, the data of a shared space 107 is stored encrypted in the knowledge base 105. Consequently, the storage is accessible only from the computing systems 102 authenticated to this shared space 107.
[0036] In one implementation, the user encrypts his data in the (common) knowledge base via the human-machine interface used. In this case, security is increased by the fact that only he has access to his encryption key and therefore only he is technically able to read the encrypted data.
[0037] The knowledge base 105 also includes a common space 108, the common space 108 being accessible to all users of the authenticated entities of the architecture 100.
[0038] In other words, the knowledge base 105 includes a public base 108 which contains all the public data accessible by all users registered in the architecture 100. It should be emphasized that any third party not authenticated with the architecture 100 does not have access to this common space 108, nor even to the architecture 100.
[0039] It follows that the architecture 100 offers users the possibility of grouping the data into a private space 106 (creation of a group by affinity), into a shared space 107 (creation of a larger group by affinity) or into a common space 108. A space 106, 107 or 108 constitutes a set of data gathered according to one or more criteria.
[0040] A space 106, 107 or 108 may constitute, for example, a set of data gathered according to a criterion of confidentiality (private, shared or common) and ownership (the group of users having inserted this data).
[0041] Furthermore, a user belonging to an entity can, at any time, move his data from one space to another within the knowledge base 105 using his human-machine interface 103. He thus has total control over his data and can share it, at any time, with the user network of his choice. He can only modify the confidentiality criterion, not the ownership criterion.
[0042] For example, a user A shares some of his data so that it can be accessed by a user B. He then moves his data from his private space 106 to a shared space 107 of his entity using his decryption key, then he associates a reading right to this shared space 107 with user B. The latter can therefore consult the biomarker sequences present in the shared database 107 of user A.
[0043] According to another example, a user C shares some of his data so that it is accessible by all users registered in the architecture 100. He then moves his data from his private space 106, by means of his human-machine interface 103, to the common space 108 using his decryption key. All users authenticated to the architecture 100 then have access to the decrypted data.
[0044] In a non-limiting embodiment, the biomarker sequences in the knowledge base are all associated with keywords. If a user is working on a particular type of data, they can share their work using the associated keywords. In this case, other users working on that particular type of data will be immediately notified, and the association is either common or private.
[0045] Furthermore, as already mentioned, as soon as a calculation carried out by a user via at least one calculation system 102 is completed, the results are transmitted to the knowledge base 105 via an access interface 104 which makes it possible to link software code to the knowledge base 105, for example via a REST API type interface. The result of each calculation corresponds to a list of biomarker sequences of different types depending on the bioinformatics software implemented.
[0046] The human-machine interface 103 allows the user to deposit NGS data files and to interface the user with a calculation system 102 making it possible to organize and interpret the biomarker sequences, but also to be able to consult and edit the knowledge base 105.
[0047] This knowledge base 105 is thus capable of storing information of different types, of being able to connect all this information together and of associating metadata with biomarker sequences, then of sharing them between several users with a previously defined security level.
[0048] In other words, this knowledge base 105 contains public data and is continually fed by users' calculations (or analyses), then corrected and updated. Indeed, each user can, at the same time, consolidate their results in relation to the set of biomarkers that they have identified, then in turn complete the content of the knowledge base by associating new metadata with biomarkers.
[0049] For example, a clinician can exchange, for each patient or clinical protocol, his information and interpretations with other centers (e.g. network between several university hospitals), which will in turn allow him to enrich the data on the patient's pathology, better understand and manage the treatment, and thus be able to integrate the information listed in the knowledge base 105.
[0050] This modular 100 architecture offers flexibility to easily integrate with future medical tools, which will make it easy to complete the range of offers with new applications and also guarantee the most accurate results on the market. This type of 100 architecture allows the user to directly access the applications via the human-machine interface 103 and to consult or download their interpreted results online.
[0051] In a non-limiting implementation, the architecture 100 comprises a set of resources 109 making it possible to organize the biomarker sequences present in the knowledge base 105.
[0052] These resources are all implemented within Knowledge Base 105.
[0053] For example, the architecture 100 includes a system for identifying an entity and the elements it manipulates (for example, reads). As previously mentioned, an entity is formed by a grouping of users and the elements they manipulate. Consequently, the entity can designate one or more user(s) but can also identify a service of the user or users (in the latter case, this service is completely independent of the other services of this same user).
[0054] For example, an analysis has its own identifier but also ownership information (user and entity) and confidentiality (common, shared or private space).
[0055] The architecture 100 also includes a user identification and authentication system. Therefore, the user identification and authentication system represents a user of the architecture, for example a researcher or a clinician. In other words, a person who creates and consults the data resulting from very high-throughput sequencing is in no way an analyzed patient. The user is identified with a unique name (the login) and authenticates with a password that only he or she knows. The system identifies the user if said user exists in the identification database. The system uses a “digest” (checksum of the password) of the authentication data to perform a comparison and validate the authentication. The identification database contains for each user the encryption key of his or her associated entity.This key is itself encrypted with the user's password. During authentication, the system uses the user's password to release the entity's encryption key. This key is then encrypted with an asymmetric system before being returned to the user. With this architecture, we obtain these security features: . Passwords are not stored, encryption keys are only accessible with a valid password from an authorized user, the user cannot view their own encryption key, the user carries with them a certificate ensuring that they have been properly identified and authenticated.
[0056] The architecture 100 may also include a system for identifying a calculation associated with at least one user and one entity. In other words, this system makes it possible to designate the results of an analysis uniquely associated with a user who is the initiator of the analysis as well as with the entity linked to the initiator of the analysis. For example, the system for identifying a calculation associated with at least one user and one entity may contain information obtained with the CRAC & the CT software suite. The resulting data may be used for the detection of biomarker sequences and feed the knowledge base.
[0057] The architecture 100 may also include a system for correlating the results obtained.
[0058] For example, if an analysis generates a biologically erroneous biomarker, it is possible to search for analyses that had the same previous result and thus avoid unnecessary biological tests.
[0059] The architecture 100 may also include a system for logically grouping calculation identifiers associated with a user and an entity. This grouping system makes it possible to improve the understanding of a biological problem, the prognosis or the diagnosis of a patient since it makes it possible to have an overall view of the calculations (e.g., the biomarker sequences) which are correlated with each other by applying a defined rule.
[0060] For example, it is possible to group together analyses from patients with the same pathology to determine common points or statistical information.
[0061] The architecture 100 may also include a system for associating meta-information with data in the knowledge base. This system may be used by users to associate a text message commenting on the result of a calculation. For example, it is possible to generate an inter-entity discussion on a problem identified by an analysis group or a specific marker. This discussion allows for rapid exchange among the different specialists on specific questions.
[0062] In a non-limiting implementation, the knowledge base 105 is formed by a database ring.
[0063] In other words, data storage is based on a key / column database ring knowledge base. This gives all users real-time read and write access to the knowledge base. The knowledge base consists of resources that can be added / removed based on capacity and performance needs, thus expanding the storage ring. This allows for the benefit of distributed, unlimited, and highly available storage since the data is replicated in three different locations.
[0064] The location of these replications is automatically calculated according to criteria (machine and geographical) so that it covers all possible failures. For example. data A is stored on machine X with two other copies on machines Y and Z. All machines X, Y and Z constitute the entire knowledge base.
[0065] It is recalled that when an action is attributed to the knowledge base, this is in fact carried out by a microprocessor of the knowledge base controlled by instruction codes stored in a memory of the knowledge base. If an action is attributed to an application, this is in fact carried out by a microprocessor of the knowledge base in a memory of which the instruction codes corresponding to the application are stored.
[0066] Generally speaking, in a classic database, there is the "CAP" theorem: C stands for "Consistency", A for "Availability" and P for "Partition" (fault tolerance and capacity). According to the invention, the knowledge base is focused on A to always guarantee a response and on P to allow massive storage ("big data"). As for C, it is sufficiently optimal to meet the needs of consistency in the context of use. No current database can meet the three conditions (CAP). In addition, a second C is added to the architecture according to the invention for "Confidentiality" because it allows to guarantee a higher level of security than the standard in force and that a classic database could meet.
[0067] Furthermore, in the implementation of the architecture according to the invention, for each analysis (or calculation) elements of the knowledge base (genomic biomarkers) are generated. Each of these elements contains a reference to the analysis(s) that generated it. As a result, it is possible to group together analyses sharing the same elements (biomarkers) and, by extrapolation, users of the analyses can be put in touch and can put themselves in touch.
[0068] The architecture according to the invention makes it possible to completely democratize the analysis so that a clinician or biologist is completely autonomous in his analysis (no more need for bioinformatics and biostatistics specialists between him and his NGS data), with a chosen level of security, defined sharing, guaranteed quality, adapted visual tools and delimited sharing.
Claims
1. Architecture (100) for genomic data analysis comprising: - a plurality of applicative nodes (101), each applicative node comprising: - a computing system (102) comprising at least one computing node, the computing system (102) being constructed and configured to perform genomic computations, - a human-machine interface (103) constructed and configured to communicate with the computing system (102), - an access interface (104) to a knowledge base (105), - a knowledge base (105) constructed and configured to communicate with the plurality of applicative nodes (101), said knowledge base (105) containing genomic data; - a plurality of private spaces (106) contained within the knowledge base (105), each private space (106) being constructed and configured: - to contain metadata and results from a computation performed by the computing system (102) of an authenticated entity comprising 1 to n user(s), characterized in that said results are encrypted using an encryption key of said entity, and - to be accessible in read-mode only to users of the authenticated entity; - an identification and authentication system for a user, the identification and authentication system being configured: - to identify said user in an identification database containing for each user the encryption key of the entity corresponding to said user, - to use a password known only to the user to release said encryption key of said entity corresponding to said user and return said encryption key encrypted with an asymmetric system to said user.
2. Architecture (100) for genomic data analysis according to claim 1, said architecture (100) being characterized in that it comprises a plurality of shared spaces (107), each shared space (107) being constructed and configured: to contain metadata and results from collaborative computations performed by the computing systems (102) of an authenticated entity comprising 1 to n user(s), and to be accessible in read-mode only to users of authenticated entities and benefiting from a right of consultation rights of said shared space (107).
3. Architecture (100) for genomic data analysis according to any one of the preceding claims, said architecture (100) being characterized in that it comprises a common space (108), said common space (108) being accessible to all users of the authenticated entities of the architecture (100).
4. Architecture (100) according to any one of the preceding claims, said architecture (100) being characterized in that the computing system (102) is constructed and configured to classify short numerical sequences of biological material.
5. Architecture (100) according to the preceding claim 4, said architecture (100) being characterized in that the computing system (102) is constructed and configured to aggregate said short numerical sequences of biological material of a same class.
6. Architecture (100) according to any one of the preceding claims, said architecture (100) being characterized in that the computing system (102) is constructed and configured to identify a biomarker sequence responding to a biological question.
7. Architecture (100) according to any one of the preceding claims, said architecture (100) being <b>characterized in that it comprises a plurality of resources (109), said resources (109) being formed by: - An identification system of an entity and elements it manipulates; - An identification system of a user, - An identification system of a computation associated to at least one user and one entity, - A system allowing to correlate performed computations, - A logical grouping system of computation identifiers associated to a user and an entity, - A system allowing to associate a meta-information to data of the knowledge base (105).
8. Architecture (100) according to any one of the preceding claims, said architecture (100) being characterized in that the knowledge base (105) is formed by a database ring.
9. Architecture (100) according to any one of the preceding claims, said architecture (100) being characterized in that the access interface (104) to the knowledge base (105) is a REST API.