Single namespace for high performance computing systems

By introducing a data processing controller into a high-performance computing system, intelligent arbitration and standardized format storage between local and remote computing clusters are achieved, solving the problems of low data transmission and management efficiency and improving the overall efficiency and security of the system.

CN122122572APending Publication Date: 2026-05-29GUARDANT HEALTH INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUARDANT HEALTH INC
Filing Date
2024-10-25
Publication Date
2026-05-29

Smart Images

  • Figure CN122122572A_ABST
    Figure CN122122572A_ABST
Patent Text Reader

Abstract

A data processing fabric controls data processing arbitration in a high performance computing system that includes one or more sites. The individual sites can include one or more server computers that execute instances of a local file system and that include one or more temporary data storage devices. The individual instances of the local file system can access files stored in objects of a primary data store. The individual objects of the primary data store can be accessed using a common identifier that indicates a storage location of the individual objects in the primary data store.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority Statement This application claims priority to U.S. Provisional Patent Application No. 63 / 593,354, filed October 26, 2023, entitled “Single Namespace for Bioinformatics System”; U.S. Provisional Patent Application No. 63 / 627,636, filed January 31, 2024, entitled “Data Processing Abstraction for Bioinformatics System”; and U.S. Provisional Patent Application No. 63 / 656,184, filed June 5, 2024, entitled “Data Processing Abstraction for Bioinformatics System”, which are incorporated herein by reference in their entirety. Technical Field

[0002] The implementation of this disclosure generally relates to the field of computer architecture, and more specifically to the implementation of computer architectures for controlling data processing operations in various systems, such as media streaming systems, scientific research systems, bioinformatics systems, content generation systems, generative machine learning systems, etc.

[0003] background The transmission and computation of massive amounts of data can be performed by high-performance computing systems. For example, bioinformatics can involve the analysis of massive amounts of data to analyze the causes of various biological conditions and identify treatments for many of them. In many cases, bioinformatics can involve the computational analysis of genomic data. Genomic data can include the nucleotide sequences of genetic material obtained from individual samples. Genomic data from a single individual can correspond to many megabytes of data storage, while genomic data from different groups of individuals can correspond to hundreds of gigabytes, or even trillions or quadrillions of data storage. In other instances, high-performance computing systems can be used to predict and model scenarios related to meteorological and geological data, as well as in the execution of machine learning algorithms and in fraud detection. Because of the massive amounts of data accessed and analyzed by bioinformatics systems and other systems utilizing high-performance computing, the storage and transmission of this data can be inefficient in terms of the network resources used, leading to performance lag. Brief description of the attached diagram Figure 1 An example architecture for processing patient data is shown, based on some examples.

[0005] Figure 2An example framework for allocating additional network resources during the transfer of patient data from local and remote data repositories is shown, based on several examples.

[0006] Figure 3 This is a flowchart illustrating example methods for efficiently processing patient data using a combination of local and remote data processing resources, based on several examples.

[0007] Figure 4 The illustration shows example computing architectures for storing, retrieving, and modifying data in a high-performance computing environment, based on several examples.

[0008] Figure 5 Example flowcharts are shown, based on some examples, of the processes for storing, retrieving, and modifying data in a high-performance computing environment.

[0009] Figure 6 This is a block diagram illustrating the components of a machine in the form of a computer system, according to some examples, which can read and execute instructions from one or more machine-readable media to perform any one or more methods described herein.

[0010] Figure 7 This is a block diagram illustrating representative software architectures that can be used in conjunction with one or more hardware architectures described herein, based on some examples.

[0011] Detailed description The following description and figures fully illustrate the particular implementations to enable those skilled in the art to practice them. Other implementations may be combined with structural, logical, electrical, procedural, and other variations. Some portions and features of some implementations may be included in or replaced by those portions and features of other implementations. The implementations set forth in the claims include all available equivalents of these claims.

[0012] Due to the massive amounts of data accessed and analyzed by high-performance computing systems, the storage, transmission, and processing of this data can be inefficient in terms of network resource utilization and can lead to lag when analyzing large datasets such as scientific, technical, media content, and bioinformatics data. Some computing architectures allow such data to be processed by remote computing engines or cloud computing systems. However, managing data arbitration is a daunting task. Specifically, users need to manually select which data operations are processed locally and which are processed remotely, which requires significant time and effort and involves navigating multiple pages of information. Furthermore, considering whether remote resources are even available to perform the requested operations and / or whether the requested operations can be performed within cost constraints is difficult and time-consuming, especially since availability and cost can change over time. Finally, sometimes the data to be processed includes sensitive and private information. Managing how such data is processed adds another level of time and cost, reducing the overall security and efficiency of conventional systems.

[0013] Since the creation of the first computer, data movement, curation, and management have been a challenge. With the use of multiple high-performance computing clusters, cloud assets, remote sites, and various storage technologies, a variety of sophisticated tools have been created to track, coordinate, retrieve, and manage data generated by a site or project. These typically require additional resources to be managed by expert teams rather than standard systems teams, and require dedicated or collaborative teams for development and maintenance. Furthermore, this necessitates extensive training for users to locate, retrieve, and archive their data, as these systems are file system-specific and cannot be utilized by any other type of file system.

[0014] The disclosed technology addresses these drawbacks by providing a data processing controller that automatically and intelligently arbitrates data processing operations locally and / or remotely using various computing clusters. Specifically, the disclosed technology involves a data processing controller on a centralized server storing multiple files in a standardized format and receiving a first request from a first computing system to access a first file among the multiple files, the first computing system generating the request using a first type of file system. The disclosed technology also receives a second request from a second computing system to access a second file among the multiple files, the second computing system generating the request using a second type of file system different from the first type of file system of the first computing system. The disclosed technology controls access to the first and second files in response to receiving the first and second requests.

[0015] In this way, data processing operations can be run and executed more efficiently with minimal user effort and interaction. Furthermore, files stored on a centralized server can be accessed by any type of file system, creating a single, file system-agnostic namespace for the files. This allows users trained to operate various user interfaces and operating systems to access a centralized collection of files stored and managed using those different operating systems and user interfaces. This improves the overall efficiency of operating the device.

[0016] Figure 1 An example architecture 100 is illustrated, based on some examples, for managing and arbitrating data between local or remotely processed data using local and / or remote computing clusters. Architecture 100 may include a life science service provider 102 and may be used to provide high-performance computing services to the life science service provider 102. The high-performance computing system may include a cluster of processors capable of performing computations in massively parallel mode. In at least some examples, the high-performance computing system can perform computations and transfer data volumes hundreds, thousands, or even millions of times larger than those of a typical desktop, laptop, or server system. The high-performance computing system may use thousands, tens of thousands, or even millions of processors to perform computations and can perform up to 10^18 floating-point operations per second. The life science service provider 102 may include an entity that provides at least one of the following: a product or service to an individual. The life science service provider 102 may include at least one of the following: an educational organization, a non-profit organization, a privately owned enterprise, or a publicly owned enterprise. In one or more examples, the life science service provider 102 may include an entity that develops treatments for one or more biological conditions. For example, the life science service provider 102 may include a pharmaceutical company that develops and / or manufactures pharmaceutical substances to treat one or more biological conditions.

[0017] In some examples, life science service provider 102 may include a diagnostic organization that develops tests to detect the presence of one or more biological conditions in a subject. Life science service provider 102 may also include a medical device entity that develops and / or manufactures medical devices to treat or detect at least one of one or more biological conditions. Furthermore, life science service provider 102 may include an organization that develops or manufactures devices, equipment, supplies, and / or combinations thereof for detecting and / or treating one or more biological conditions. In some examples, life science service provider 102 may include a healthcare service provider that provides testing, medical services, and / or treatment for one or more biological conditions. In various examples, life science service provider 102 may include one or more healthcare providers.

[0018] As used herein, a healthcare provider may refer to an entity, individual, or group of individuals involved in providing care to an individual in connection with at least one of the treatments or preventions of one or more biological conditions. Furthermore, as used herein, a biological condition may refer to a functional and / or structural abnormality in an individual, to the extent that a detectable feature produces or threatens to produce an abnormality. A biological condition may be characterized by external and / or internal characteristics, signs, and / or symptoms that indicate a deviation from the biological norm in one or more populations. A biological condition may include at least one of one or more diseases, one or more disorders, one or more injuries, one or more syndromes, one or more disabilities, one or more infections, one or more isolated symptoms, or other atypical variations in an individual's biological structure and / or function.

[0019] As used herein, "treatment" can refer to substances, procedures, routines, devices, and / or other interventions that can be administered or performed to alleviate one or more effects of a biological condition in an individual. In some examples, treatment may include substances metabolized by the individual. The substance may include compositions of substances, such as pharmaceutical compositions. The substance may be delivered to the individual by various methods, such as ingestion, injection, absorption, or inhalation. Treatment may also include physical interventions, such as one or more surgical procedures.

[0020] In at least some examples, the life science service provider 102 may store, access, process, and / or analyze at least one of the following: data corresponding to multiple subjects 104. In one or more examples, a sample 106 may be extracted from the subject 104. The sample 106 may be derived from at least one of bodily fluids or tissues obtained from the subject 104. At operation 108, the sample 106 may undergo at least one of one or more diagnostic tests or one or more analytical tests. In various examples, one or more diagnostic tests and / or one or more analytical tests performed at operation 108 may be performed to detect one or more biological conditions that may be present in the subject 104. In some examples, at least one of the one or more diagnostic tests or one or more analytical tests performed at operation 108 (also referred to as data processing operations) may include one or more assays related to the detection of one or more forms of cancer.

[0021] One or more diagnostic tests and / or one or more analytical tests performed at operation 108 may generate patient data 110. Patient data 110 may include data derived from one or more diagnostic tests and / or analytical tests performed at operation 108. For example, patient data 110 may include genomic information, genetic information, metabolomics information, transcriptomics information, fragmentomics information, immune receptor information, methylation information, epigenomic information, proteomics information, immunohistochemistry (IHC) and immunofluorescence (IF) and / or personally identifiable information (PII). PII may include information that can identify an individual when used alone or in conjunction with other relevant data. PII may contain direct identifiers that can uniquely identify an individual (e.g., passport information), or quasi-identifiers (e.g., ethnicity) that can be combined with other quasi-identifiers (e.g., date of birth) to successfully identify an individual. PII may include sensitive personally identifiable information such as full name, social security number, driver's license, financial information, and / or medical records.

[0022] As used herein, “fragment information” can in particular include information related to the length analysis of DNA or RNA fragments to determine the presence or absence of a tumor and to characterize the tumor. In at least some examples, fragment information may correspond to nucleosome structure and transcription factor binding sites. In some examples, fragment information may include fragment endpoint density, plasma DNA size, endpoints, nucleosome footprint, DNA fragments aligned with base positions in the genome, the number of DNA fragments that begin or end at specific base positions in the genome, fragment origins and lengths associated with a specific condition, heterogeneity patterns of cfDNA localization in cancer, nucleosome occupancy, nucleosome dynamics, chromatin organization, structure and function, chromatin state, consequences of genomic aberrations, and / or DNA epigenetic changes associated with health and disease.

[0023] In some examples, “genomic information” may correspond to a nucleic acid sequence derived from sample 106. Genomic information may indicate one or more mutations corresponding to the genes of subject 104. Mutations in the genes of subject 104 may correspond to differences between the nucleic acid sequence of subject 104 and one or more reference genomes. Reference genomes may include known reference genomes, such as hg19. In various examples, mutations in the genes of subject 104 may correspond to differences in the germline genes of subject 104 relative to the reference genome. In some examples, the reference genome may include the germline genome of subject 104. In one or more additional examples, mutations in the genes of subject 104 may include somatic mutations. Mutations in the genes of subject 104 may be related to insertions, deletions, single nucleotide variants, loss of heterozygosity, duplications, amplifications, translocations, fusion genes, or one or more combinations thereof. In at least some examples, genomic information may correspond to non-coding regions of the genome. Non-coding regions may be associated with the regulation of one or more genes. In one or more examples, analysis of non-coding regions may detect one or more epigenetic characteristics in one or more patients.

[0024] In some examples, the genomic information included in patient data 110 may include genomic maps of tumor cells present in one or more subjects 104. In these cases, the genomic information may be derived from analysis of genetic material, such as deoxyribonucleic acid (DNA) and / or ribonucleic acid (RNA) found in blood samples of one or more subjects 104, present due to degradation by tumor cells present in one or more subjects 104. In some examples, the genomic information of tumor cells in one or more subjects 104 may correspond to one or more target regions. The presence of one or more mutations in one or more target regions may indicate the presence of tumor cells in one or more subjects 104.

[0025] In some examples, the genetic material analyzed to generate genomic information may be derived from one or more samples 106, including but not limited to tissue samples or tumor biopsies, circulating tumor cells (CTCs), exosomes or necrosomes, or from circulating nucleic acids. In various examples, circulating nucleic acids may be referred to herein as “cell-free DNA.” “Cell-free DNA,” “cfDNA molecule,” or simply “cfDNA” includes DNA molecules present in the subject 104 in an extracellular form (e.g., in blood, serum, plasma, or other bodily fluids such as lymph, cerebrospinal fluid, urine, or sputum), and includes DNA that is not contained within cells or otherwise bound to cells upon separation from the subject 104. Although DNA is originally present in one or more cells of a large, complex biological organism (e.g., a mammal) or in other cells of that organism (e.g., bacteria), the DNA has been released from the cell into fluids found in that organism. cfDNA includes, but is not limited to, cell-free genomic DNA of subject 104 (e.g., genomic DNA of a human subject) and cell-free DNA of microorganisms (e.g., bacteria) residing in subject 104 (whether pathogenic or bacteria typically found in common colonization sites such as the gut or skin of a healthy control group), but excludes cell-free DNA of microorganisms that have only contaminated a body fluid sample. Typically, cfDNA can be obtained by obtaining a volume of fluid without requiring an in vitro cell lysis step and includes the removal of cells present in the fluid (e.g., blood centrifugation to remove cells).

[0026] In some examples, patient data 110 may include information about subject 104 (e.g., PII). Specifically, patient data 110 may include an identifier for subject 104, physical characteristics of subject 104 (e.g., weight, height), age of subject 104, personal information of subject 104, ethnic background of subject 104, one or more combinations thereof, etc. Furthermore, patient data 110 may include a medical record corresponding to patient data 110. For illustration, subject 104's medical record may be generated alongside and / or in conjunction with patient data 110. The medical record may include imaging information, laboratory test results, diagnostic test information, clinical observation results, dental health information, healthcare professional records, medical history forms, diagnostic request forms, medical procedure order forms, medical infographics, one or more combinations thereof, etc. The medical record may also indicate lifestyle information, such as smoking status, alcohol consumption, sleep habits, one or more combinations thereof, etc. Any of this information may be characterized as a PII, or more generally as personal information.

[0027] Life science service provider 102 may include a bioinformatics system 114 that performs one or more data processing operations to analyze at least one of patient data 110. The bioinformatics system 114 may implement one or more statistical techniques to analyze at least one of the patient data 110. In some examples, the bioinformatics system 114 may implement one or more machine learning techniques (or machine learning models (ML)) to perform data processing operations to analyze at least one of the patient data 110. In various examples, the bioinformatics system 114 may analyze at least one of the patient data 110 to determine characteristics of a subject 104 in which a biological condition exists. For example, the bioinformatics system 114 may analyze at least one of the patient data 110 to determine one or more genomic characteristics of at least a portion of a subject 104 in which at least one form of cancer exists. The disclosed examples provide a data processing controller (which may be implemented locally by one or more local computing clusters and / or remotely by one or more cloud computing clusters located away from the local computing clusters).

[0028] The data processing controller can store patient data 110 and / or various other data used by the bioinformatics system 114 in a centralized management manner. Specifically, the data processing controller can be implemented by one or more cloud computing clusters or servers, and can store patient data 110 and / or various other data in one or more files. Files can be stored in a standardized format or a format compatible with the operating system or file system of the data processing controller. The data processing controller can maintain various metadata or data objects associated with each file system. Data objects indicate the specific storage location of the file system on the cloud computing cluster, the locking status of each file (e.g., indicating whether the file is currently being used by a client computing system), and changes associated with the file. Files stored by the cloud computing cluster can serve as sources of the true version of the file.

[0029] Each computing system can implement different types of file systems or operating systems and can present different types of user interfaces. The computing system can communicate with a data processing controller on a cloud computing cluster or server to gain access to certain files among multiple files stored by the cloud computing cluster. The data processing controller can receive requests for individual files and can access the data objects associated with the file in response to receiving a request from the individual computing system. The data processing controller can retrieve individual files based on storage locations indicated in one or more data objects. The data processing controller can also determine whether a file is currently in use based on one or more data objects. Then, the data processing controller can control the computing system's access to the file based on retrieving the individual file and whether the file is currently being accessed by another computing system.

[0030] If an individual file is currently being used by another computing system, the data processing controller can block access to the individual file and notify the requesting computing system that the individual file is currently in use. Alternatively, the data processing controller can access the cache of the other computing system currently using the file to determine changes associated with the file. The data processing controller can then update the individual file and generate a copy of the updated individual file. The data processing controller can then provide the copy of the individual file to the requesting computing system. The data processing controller can track changes made by each computing system currently accessing the same file in one or more data objects, such as by periodically requesting updates from the corresponding caches stored on the computing systems. The data processing controller can periodically merge changes stored in one or more data objects into a source copy of the file's true version accessed by multiple computing systems.

[0031] In some examples, the data processing controller can stream individual files to the requesting compute system. The requesting compute system can store a copy of the individual file in a local cache and can convert the file from a standardized format to a format associated with the requesting compute system's file system. The requesting compute system can then present the contents of the individual file on its user interface. The requesting compute system can track changes made to the individual files in its local cache. The requesting compute system can upload the tracked changes from the local cache to one or more objects managed and stored by the data processing controller on the cloud computing cluster. After verifying the changes, the data processing controller can update the source of the file's true version based on the changes stored in the corresponding object associated with the individual file.

[0032] In some examples, the bioinformatics system 114 can analyze at least one of the patient data 110 to identify one or more cohorts corresponding to multiple groups of subjects 104. In some examples, the bioinformatics system 114 can analyze at least one of the patient data 110 to determine the effectiveness of one or more treatments provided to at least a portion of subjects 104 relative to one or more biological conditions present in a group of subjects 104. Additionally, the bioinformatics system 114 can analyze at least one of the patient data 110 to determine recommendations for treatment of at least a portion of subjects 104 related to one or more biological conditions present in a group of subjects 104. Furthermore, the bioinformatics system 114 can analyze at least one of the patient data 110 to determine the amount of progression of biological conditions present in at least a portion of subjects 104. In at least some examples, the bioinformatics system 114 can analyze at least one of the patient data 110 to determine biological conditions present in at least a portion of subjects 104. In some examples, the bioinformatics system 114 can analyze at least one of the patient data 110 to determine a diagnosis for at least a portion of subjects 104.

[0033] Life science service provider 102 may include one or more computing devices 116 that can access bioinformatics system 114. The one or more computing devices 116 may include at least one of one or more desktop computing devices, one or more laptop computing devices, one or more tablet computing devices, one or more mobile computing devices, one or more smartphones, one or more wearable computing devices, or combinations thereof. Life science service provider 102 may also include and / or be coupled to a local computing cluster 118. In one or more examples, the local computing cluster 118 may include one or more data repositories and servers or computer systems located at at least one site of life science service provider 102. In various examples, the local computing cluster 118 may be coupled to at least one of the one or more computing devices 116 or bioinformatics system 114 via one or more physical network connections, one or more of which are maintained, controlled, or managed by life science service provider 102. The local computing cluster 118 may store and process at least one of at least a portion of patient data 110 (referred to as a batch of data). In some cases, local computing clusters 118 may be physically located away from computing devices 116, but may be securely coupled via physical lines and internal network security specifically associated with the life science service provider 102. Each local computing cluster 118 may represent an individual computing system within a computing system that accesses a cloud computing system or server managed by a data processing controller.

[0034] In some examples, life science service provider 102 may communicate with remote computing cluster 120 (also referred to as a cloud computing cluster or cloud cluster). Remote computing cluster 120 may be located off-site relative to one or more locations of life science service provider 102 and may be controlled, maintained, or managed by at least one entity different from life science service provider 102 (e.g., a third-party entity relative to life science service provider 102). In one or more examples, remote computing cluster 120 may be controlled, maintained, or managed by at least one of one or more third-party cloud computing service providers. Specifically, cluster 120 may include a first set of computing clusters or systems provided by a first entity that is a third party relative to life science service provider 102, and may include a second set of computing clusters or systems provided by a second entity that is a third party relative to life science service provider 102. The first set of computing clusters may be associated with a different set of costs and resources available to life science service provider 102 compared to the second set of computing clusters.

[0035] In some examples, the first set of computing clusters or systems may implement a first type of file system running on a first type of operating system, while the second set of computing clusters or systems may implement a second type of file system running on a second type of operating system. Even though the first and second sets of computing clusters operate using different types of file systems, they can still share access to and receive files stored on the centralized computing system. To this end, the first and second sets of computing clusters can implement a conversion engine to convert files from the standardized format of the centralized computing system to the corresponding file format of the first and second sets of computing clusters.

[0036] In various examples, local computing cluster 118 and / or remote computing cluster 120 may store at least one portion of patient data 110. In some cases, a batch of data including patient data 110 may be stored at a centralized location, for example, via one or more clusters 120. Links to this batch of data can be generated and made available to computing device 116. Computing device 116 may use the links to instruct local computing cluster 118 and / or cluster 120 to perform one or more data processing operations. That is, a data processing controller may instruct any combination of local computing clusters 118 and cluster 120 to use the links to perform or run one or more operations. This can minimize the bandwidth and time spent moving data for processing. For example, instead of sending the batch of data from one device to another over a network, the data processing controller may send a link to the computing cluster selected to perform the data processing operation. The computing clusters can then use the links to retrieve the batch of data from a centralized storage device, thus accelerating the processing of such data. After the batch of data has been processed, the results or processed data are provided back to the centralized storage device (which can be implemented by cluster 120) so that it can be used by other computing clusters.

[0037] Life science service provider 102 can communicate with remote computing cluster 120 via physical communication network 122. Physical communication network 122 may include a communication network infrastructure controlled, maintained, or managed by an entity other than life science service provider 102. Cluster 120 is publicly accessible via the Internet and can perform operations simultaneously for multiple parties or entities, and is non-exclusively associated with life science service provider 102. For example, physical communication network 122 may include at least one physical networking device controlled, maintained, or managed by network management system 124 of a network service provider. Network management system 124 can control network resources used by multiple different entities that utilize physical communication network 122 for data transmission, processing, and / or access. For example, network management system 124 can allocate bandwidth and / or processing resources of physical communication network 122 to an entity using physical communication network 122 for at least one data transmission, processing, or access, where bandwidth corresponds to the amount of network resources allocated to one or more entities. Network management system 124 may also implement one or more technologies and / or protocols to facilitate efficient data transmission between endpoints of physical communication network 122.

[0038] In some examples, virtual communication network 126 may couple life science service provider 102 to remote computing cluster 120. Virtual communication network 126 may correspond to a portion of physical communication network 122 allocated to life science service provider 102 at a given time. For illustration, various portions of the network resources of physical communication network 122 may be allocated to multiple different entities at a given time. In at least some examples, the amount of network resources of physical communication network 122 dedicated to the virtual communication network between remote computing cluster 120 and life science service provider 102 may change over time. In one or more illustrative examples, the bandwidth of virtual communication network 126 may be modified based on the amount of data to be transmitted between remote computing cluster 120 and life science service provider 102.

[0039] Architecture 100 may also include a database management system 128. The database management system 128 may be coupled to a remote computing cluster 120 and a local computing cluster 118. In one or more examples, computing device 116 may use the database management system 128 to access data stored by the local computing cluster 118 and the remote computing cluster 120. In various examples, the database management system 128 may facilitate access to at least one of files or objects stored by the local computing machine 118 and the remote computing cluster 120 in response to a request generated by at least one of computing device 116 or bioinformatics system 114.

[0040] Life science service provider 102 may utilize the storage resources of one or more cloud storage providers to store at least one portion of patient data 110 in a remote computing cluster 120. In one or more examples, life science service provider 102 may acquire and / or generate data volumes that may exceed the capacity of local computing machine 118. In these scenarios, excess data may be stored by remote computing cluster 120. Additionally, at least one portion of patient data 110 may be stored by remote computing cluster 120 for other reasons, such as storing medical record information to comply with one or more regulatory frameworks. In one or more additional examples, remote computing cluster 120 may store at least one portion of patient data 110 to minimize the cost of storing and retrieving information for life science service provider 102 and / or improve the efficiency of storing and retrieving information for life science service provider 102.

[0041] In at least some examples, the memory resources used to store patient data 110 may be larger than the memory resources used to store other data. In various examples, the memory resources used to store patient data 110 may be two, five, ten, twenty, fifty, up to 100, 1000, 10,000, 100,000, or more times larger than the memory resources used to store other data. In one or more illustrative examples, patient data 110 may include DNA sequences and expression values ​​for multiple genomic regions (such as dozens, hundreds, or thousands of genomic regions) of an individual patient and may consume up to hundreds of gigabytes of memory resources. In one or more further illustrative examples, patient data 110 for an individual patient may include sample identifiers, batch information, and patient characteristics, which may be stored in a text file consuming approximately several hundred gigabytes of memory resources, although in at least some cases, the amount of memory resources used to store patient data 110 for an individual patient may be larger, such as from approximately tens of megabytes to hundreds of megabytes or more.

[0042] In one or more examples, local computing cluster 118 may store a first set of patient data 110, and remote computing cluster 120 may store a second set of patient data 110. In some examples, local computing cluster 118 may include a cache memory that stores at least a portion of the patient data 110 while analysis of at least one of the patient data 110 is performed by bioinformatics system 114.

[0043] In various examples, computing device 116 can be used to generate requests for at least one of the transfer, processing, and / or access to at least a portion of patient data 110 from a centralized storage device. Requests for the transfer, processing, and / or access to at least one portion of patient data 110 can be generated to analyze at least a portion of patient data 110 (e.g., a batch of data) using bioinformatics system 114. In some examples, requests for the transfer, processing, and / or access to at least one portion of patient data 110 can be generated based on one or more application programming interface (API) calls to database management system 128. In some cases, the request to process the batch of data is routed to... Figure 2The data processing controller 202 shown may be processed by or be part of the data processing controller 202. The data processing controller 202 may be implemented by computing devices of the life science service provider 102 and / or by computing devices of cluster 120. The data processing controller 202 may select a computing cluster (which may include any combination of local computing cluster 118 and / or cluster 120) to perform operations on the batch of data. In some examples, the data processing controller 202 may perform the selection of the computing cluster based on a determination of whether the batch of data includes sensitive or private information (e.g., PII). For example, in response to determining that the batch of data includes sensitive or private information, the data processing controller 202 may prevent the batch of data from being processed by cluster 120 and may ensure that the batch of data is exclusively processed by the local computing cluster 118.

[0044] Figure 2 An example framework 200 for arbitrating data processing operations between local and remote computing clusters is shown, based on several examples. Framework 200 may include information regarding... Figure 1 The described components include a life science service provider 102, computing devices 116, a local computing cluster 118, a remote computing cluster 120, a physical communication network 122, a network management system 124, and a database management system 128. Figure 2 In an illustrative example, life science service provider 102 includes a data processing controller 202 that monitors requests for processing data received from computing device 116. In some examples, the data processing controller 202 may analyze the request to determine whether such a request includes or is associated with a batch of data containing sensitive or private information (e.g., PII). To enhance security and ensure privacy remains intact, the data processing controller 202 selectively arbitrates data processing operations between local computing cluster 118 and cluster 120 based on whether such data processing requests include or are associated with multiple batches of data containing sensitive or private information (e.g., PII).

[0045] Data processing controller 202 may receive a request to perform one or more data processing operations on a batch of data in a first file stored in multiple files centrally stored and managed by a cloud computing system or server. In some examples, the batch of data includes patient data, which includes genomic information of multiple subjects. In some examples, the data processing operations include performing an analysis of at least a portion of the batch of data by a bioinformatics system 114 implemented by a selected computing cluster, and determining one or more characteristics of subjects corresponding to at least a portion of the batch of data based on the performed analysis. In some examples, one or more characteristics include one or more genomic mutations present in nucleic acids derived from samples obtained from one or more subjects, and the nucleic acids correspond to cell-free deoxyribonucleic acid (DNA) extracted from bodily fluid samples obtained from one or more subjects. In some examples, one or more characteristics include resistance to treatments provided to one or more subjects that address a biological condition present in one or more subjects. In some examples, the biological condition corresponds to a form of cancer. In some examples, the analysis includes determining recommendations for treatments provided to one or more subjects to treat the biological condition present in one or more subjects.

[0046] In some examples, data processing controller 202 stores multiple files in a standardized format on a centralized server. Data processing controller 202 receives a first request from a first computing system (e.g., local computing cluster 118) to access a first file among the multiple files, the first computing system generating the request using a first type of file system. Data processing controller 202 receives a second request from a second computing system (e.g., cluster 120) to access a second file among the multiple files, the second computing system generating the request using a second type of file system different from the first type of file system of the first computing system. Life science service provider 102 controls access to the first and second files in response to receiving the first and second requests. Any operations performed by data processing controller 202 as discussed herein can be performed by any other component, such as local computing cluster 118, cluster 120, database management system 128, etc.

[0047] In some aspects, the data processing controller 202 uses data objects or other information stored in association with the first file to identify the storage location of the first file on the centralized server. The data processing controller 202 retrieves the first file and stores an indication in the data object associated with the first file that the first file is currently accessed by the first computing system. The data processing controller 202 then streams the first file to the first computing system, which stores the first file in its local cache. The first computing system may convert the first file from a standardized format to a format compatible with a first type of file system. The first computing system tracks changes to the first file in its cache, such as based on user interactions with the presentation of the first file by applications (e.g., user interfaces) of the first computing system.

[0048] In some respects, the first computing system updates one or more objects associated with the first file, stored on a centralized server, based on changes to the first file in the first computing system's cache. The first file stored on the centralized server represents the source of the true version of the first file. In response to receiving an update from the first computing system, the data processing controller 202 can verify the change and can periodically update the source of the true version of the first file. The data processing controller 202 can also determine whether the first file is currently being accessed by another computing system, such as a second computing system. This can be done by the data processing controller 202 querying a locked file or other data objects associated with the first file that indicate where the first file is currently being accessed.

[0049] In some examples, if the second computing system is also currently accessing the first file, the data processing controller 202 can transfer changes stored in one or more objects (based on changes received from cached copies of the file from the first computing system) to the second computing system. The second computing system can then update its local cached copy of the first file and continue rendering and tracking modifications to the first file.

[0050] In some examples, the data processing controller 202 may prevent multiple computing systems from accessing the same file based on a locking state stored in one or more objects associated with the file. For example, after a first computing system requests access to a first file, the data processing controller 202 stores an indication in an object that the first file is currently locked and used by the first computing system. Later, the data processing controller 202 may receive a request to access the first file from a second computing system. The data processing controller 202 may query the data object associated with the first file to determine whether the first file is currently being used by the first computing system. In response, the data processing controller 202 may notify the second computing system that the first file is currently in use and prevent the second computing system from accessing the first file. Alternatively, the data processing controller 202 may query the first computing system to see if the first computing system has completed its access to the first file. If so, the data processing controller 202 may receive changes to the first file stored in a cache from the first computing system. In response to updating the data object representing the changes stored in the cache of the first computing system, the first computing system may delete the cache of the first file and / or delete any changes that have been transmitted to the data processing controller 202 from the cache. For example, each time one or more changes are sent from the cache of the first computing system to the data processing controller 202, the first computing system deletes those changes from the cache and begins storing the new changes into the cache. This reduces the overhead of the amount of data sent to the data processing controller 202 over the network.

[0051] The data processing controller 202 first updates the provenance of the first file, then streams the updated first file to the second computing system. The data processing controller 202 updates the lock state to now indicate that the first file is being accessed by the second computing system, not the first computing system. The lock file can store a history of access patterns indicating when and which computing systems accessed each associated file.

[0052] In some examples, the data processing controller 202 determines that a second computing system is associated with a higher priority than the first computing system. In this case, the data processing controller 202 can automatically retrieve changes to a first file stored in a cache from the first computing system and remove the lock state associated with the first computing system. The data processing controller 202 first updates the provenance of the first file and then automatically streams the updated first file to the second computing system. The data processing controller 202 updates the lock state to now indicate that the first file is being accessed by the second computing system instead of the first computing system.

[0053] Figure 3This is a flowchart of an example method 300 (or process) for arbitrating data processing between local and / or remote computing clusters, based on some examples. At operation 302, method 300 may include multiple files stored in a standardized format by a data processing controller 202 on a centralized server. At operation 304, data processing controller 202 receives a first request from a first computing system to access a first file among the multiple files, the first computing system generating the request using a first type of file system. At operation 306, data processing controller 202 receives a second request from a second computing system to access a second file among the multiple files, the second computing system generating the request using a second type of file system different from the first type of file system of the first computing system. At operation 308, data processing controller 202 controls access to the first and second files in response to receiving the first and second requests.

[0054] Figure 4 Example computing architecture 400 for storing, retrieving, and modifying data in a high-performance computing environment is illustrated, based on several examples. Computing architecture 400 can be implemented with respect to service providers, academic institutions, non-profit entities, research entities, commercial enterprises, or one or more combinations thereof. In various examples, at least a portion of the components, devices, systems, and / or data repositories of computing architecture 400 can be controlled, maintained, or managed by at least one single entity. Additionally, the various components, devices, systems, and / or data repositories of computing architecture 400 can be controlled, maintained, or managed by at least one different entity.

[0055] In one or more illustrative examples, the data and metadata stored, analyzed, and transmitted by the computing architecture 400 may be generated by multiple sensors and include meteorological data, molecular data, transportation-related data (such as data generated by autonomous vehicle sensors), geological data, genetic data, data generated by one or more diagnostic tests, data generated by one or more medical devices, etc. In one or more additional illustrative examples, the data stored, analyzed, and transmitted by the computing architecture 400 may also include media content, communication data, financial data, scientific data, etc. Metadata may include timing information associated with the data, source information associated with the data, identification information, priority information, destination information, parameters used to generate the data, etc. In one or more additional illustrative examples, the information stored, analyzed, and transmitted by the computing architecture 400 may be related to a patient undergoing at least one of one or more biological condition tests or treatments.

[0056] High-performance computing (HPC) systems can include clusters of processors capable of performing computations in massively parallel fashion. In at least some examples, HPCs can perform computations and transfer hundreds, thousands, or even millions of times more data than typical desktop, laptop, or server systems. HPCs can use thousands, tens of thousands, or even millions of processors to perform computations and can perform up to 10^18 floating-point operations per second. Additionally, the data to be read from or written to a HPC can be on the order of petaflops and exaflops. In HPCs, the movement of data to storage devices can be offloaded from the system performing the data-related computational operations or performed independently of the system performing the data-related computational operations. In this way, HPCs can efficiently read and write data to storage devices while dedicating additional computing resources to performing computations on the data.

[0057] The computing architecture 400 may include a master object data repository 402, which includes multiple data storage devices. The master object data repository 402 may store information in data containers 404. In various examples, the data container 404 may be referred to as a bucket. The data container 404 may store information associated with one or more data files 406. In one or more examples, the information stored by the data files 406 may be generated based on input obtained via one or more applications. Each data file 406 may include multiple objects. For example, each data file 406 may include a data object 408. The data object 408 may include information to be accessed and / or manipulated by one or more users or at least one of one or more applications. Each data file 406 may also include a metadata object 410. The metadata object 410 may include information associated with the data object 408. For example, the metadata object 410 may include an identifier for the data object 408. The identifier of data object 408 may include at least one of a pointer, a Uniform Resource Locator (URL), or an Internet Protocol (IP) address, indicating the storage location of data object 408 in the master object data repository 402 and being accessible to data object 408. In one or more additional examples, metadata object 410 may indicate a directory accessible to data object 408. In other examples, metadata object 410 may indicate attributes of data object 408. For illustration, metadata object 410 may indicate one or more owners of data object 408, access permissions for data object 408, timestamps indicating one or more events associated with data object 408, one or more combinations thereof, etc.

[0058] In at least some examples, each data file 406 may include a lock object 412. A lock object 412 may be generated in response to a write request for data object 408. Lock object 412 may indicate that the current user providing the write request can modify data object 408 and other users cannot modify data object 408. In one or more illustrative examples, lock object 412 may include an identifier of the user who submitted the request to write to data object 408, a file system identifier, or both. In various examples, lock object 412 may be deleted from or otherwise removed from data file 406. For example, lock object 412 may be removed from data file 406 after one or more write operations on data object 408 have been completed. In one or more additional examples, lock object 412 may expire after a period of time. In one or more illustrative examples, lock object 412 may have a duration of several microseconds, milliseconds, seconds, minutes, or hours. Lock object 412 may be removed from data file 406 in response to its expiration. After lock object 412 is removed, another lock object 412 may be generated in response to a write request related to data object 408 from the same user who provided the write request that caused lock object 412 to be generated, or in response to a write request to data object 408 from another user. In one or more additional examples, lock object 412 may be removed in response to streaming data object 408 to one or more file systems.

[0059] although Figure 4 The illustrative example shows the primary object data repository 402 as the only object data repository included in the computing architecture 400; however, in one or more additional implementations, the computing architecture 400 may include multiple object data repositories. In these scenarios, the primary object data repository 402 may be a source of baseline truth data for the computing architecture 400, and one or more additional object data repositories may operate as backup object data repositories. In one or more further examples, the primary object data repository 402 may be maintained, controlled, or managed by at least one of a first object data storage service provider, and at least a portion of one or more additional object data repositories 402 may be maintained, controlled, or managed by at least one of one or more additional object data storage service providers. In other implementations, the primary object data repository 402 and one or more additional object data repositories may be maintained, controlled, or managed by at least one of the same object data storage service provider.

[0060] The master object data repository 402 may be coupled to the object data repository management system 414. The object data repository management system 414 may control and / or manage access to the data file 406 stored by the master object data repository 402. In one or more examples, the object data repository management system 414 may receive a request to read data object 408 or a request to write to data object 408. The object data repository management system 414 may retrieve data object 408 from data container 404 and send a copy of data object 408 to the computing device requesting access to data object 408. Upon receiving a write request, the object data repository management system 414 may generate a lock object 412 associated with the data object 408, which is the subject of the write request.

[0061] The computing architecture 400 may also include a first site data repository management system 416 that electronically communicates with a first site data storage device 418. Information stored by the first site data storage device 418 can be accessed using a first site file system 420. Additionally, the computing architecture 400 may include a second site data repository management system 422 that electronically communicates with a second site data storage device 424. Information stored by the second site data storage device 424 can be accessed using a second site file system 426. In one or more examples, the first site data repository management system 416, the first site data storage device 418, and the first site file system 420 may correspond to a first location, and the second site data repository management system 422, the first site data storage device 424, and the second site file system 426 may correspond to a second location. In one or more illustrative examples, the first location and the second location may be separated by a relatively large distance, such as tens of miles, hundreds of miles, or up to thousands of miles. In one or more additional illustrative examples, the first location and the second location may be separated by a relatively short distance, such as less than a few miles. The first location and the second location may correspond to different sites of one or more entities, which maintain, control, manage or supervise at least one of the following: the first site data repository management system 416, the first site data storage device 418, the first site file system 420, the second site data repository management system 422, the second site data storage device 424, and the second site file system 426.

[0062] In various examples, computing architecture 400 may also include a remote data repository management system 428 that electronically communicates with a remote data repository 430. In one or more examples, the remote data repository management system 428 may be controlled, maintained, monitored, or managed by at least one of one or more cloud storage service providers. Information stored by the remote data repository 430 may be accessed via a remote file system 432. Furthermore, computing architecture 400 may include a computing operating system 434. The computing operating system 434 may provide processing resources that can be used to perform computational operations on data objects 408 stored by the master object data repository 402. In at least some examples, the processing resources provided by the computing operating system 434 may be provided by one or more cloud computing providers. In one or more additional examples, the processing resources may be available on one or more servers of one or more cloud computing providers.

[0063] Furthermore, the computing architecture 400 may include one or more user devices 636. One or more user devices 636 may include one or more computing devices, such as one or more laptop computing devices, one or more tablet computing devices, one or more desktop computing devices, one or more mobile computing devices, or one or more combinations thereof. One or more user devices 636 may access at least one of the first site file system 420 or the second site file system 426. One or more user devices 636 may receive input to access information stored by one or more data files 406 of the master object data repository 402. For example, one or more user devices 636 may generate read or write requests that can be used to access one or more data objects 408.

[0064] In one or more illustrative examples, user device 436 may correspond to a first site and have access to the first site file system 420. In these cases, user device 436 may be used to identify an identifier of a file to be accessed via user device 436. A request including the file identifier may then be sent to the first site file system 420. In various examples, the file identifier may also correspond to a directory of the first site file system 420. Based on the file identifier, the first site file system 420 may identify metadata stored in the first site data storage device 418. For example, the first site data storage device 418 may store first metadata 438. The first metadata 438 corresponds to a data file 406 stored by the master object data repository 402, which is accessible by one or more user devices 636. In one or more additional illustrative examples, the first metadata 438 may indicate the location of the data file 406 within the master object data repository 402. In at least some examples, in response to receiving a file identifier in a request to access the data file 406, the first site file system 420 may determine a portion of the first metadata 438 corresponding to the requested data file 406. The first site file system 420 can determine the location of the requested data file 406 based on the file identifier included in the access request, and send a request to the object data repository management system 414 to retrieve the data object 408 corresponding to the requested data file 406. The object data repository management system 414 can generate a lock object 412 for the data file 406 and send a copy of the data object 408 to the first site data repository management system 416. The copy of the data object 408 can be stored in a first temporary data object repository 440. In one or more additional illustrative examples, the first temporary data object repository 440 may include a cache memory.

[0065] User equipment 436 requesting data file 406 can capture input that can be used to modify a copy of data object 408 stored by a first temporary data object repository 440. In one or more examples, at least a portion of the information included in the copy of data object 408 may undergo one or more computational operations. One or more computational operations may include at least one of one or more text processing operations, one or more mathematical operations, one or more statistical operations, or one or more machine learning operations. In one or more examples, one or more computational operations may be performed by a first site data repository management system 416. In one or more additional examples, one or more computational operations may be performed by a computing operating system 434. In at least some examples, the computing operating system 434 identified as performing one or more computational operations may be determined based on at least one of the cost, capacity, or availability of performing one or more computational operations. For example, the cost, availability, and / or capacity of multiple cloud service providers performing one or more computational operations may be analyzed to determine a specific cloud service provider performing one or more computational operations.

[0066] After the selected cloud service provider has performed one or more computing operations, the modified data object can be sent to the first site data repository management system 416. The modified data object can be stored in the first temporary data object repository 440. In one or more examples, the first site file system 420 can send the modified data object to the object data repository management system 414. The object data repository management system 414 can store the modified data object in the main object data repository 402. In at least some examples, the modified data object can be stored in the data container 404 of the original data object 408. In other examples, the modified data object can be stored in the data file 406 of the original data object 408. In various examples, the lock object 412 corresponding to the original data object can be removed from the data file 406. Furthermore, the metadata 410 corresponding to the original data object 408 can be modified. For example, the storage location of the modified data object can be updated. In the case of updating the metadata object 410 of the modified data object, the object data repository management system 414 can send the modification of the metadata object 410 to the first site data repository management system 416. In these scenarios, the first site file system 420 can modify the first metadata 438 to correspond to the update of the metadata object 410.

[0067] In response to an update to the first metadata 438, the first site file system 420 can send the update to the second site file system 426. The second site file system 426 can then update the second metadata 442 stored by the second site data storage device 424. In this way, updates to metadata made by different file systems at different sites are propagated to other sites. Therefore, the various data objects stored by the master object data repository 402 can be accessed by both the first site file system 420 and the second site file system 426. When the user device 436 sends a request to access data from the master object data repository 402 using the second site file system 426, the second site data storage device 424 includes a second temporary data object repository 444 to store data objects 408 accessed and / or generated via the second site file system 426, and modified data objects. Additionally, the data objects and / or modified data objects stored in the first temporary data object repository 440 and the second temporary data object repository 444 may be removed after a specified period of time or in response to one or more commands generated by one or more user devices 636, the first site file system 420, or the second site file system 426. Additionally, where the remote data repository 430 is a backup storage device for the primary object data repository 402, the modified data objects and modified metadata may also be stored by the remote data repository 430 via the remote file system 432.

[0068] Figure 5 An example flowchart of a process 500 for storing, retrieving, and modifying data in a high-performance computing environment, implemented according to one or more examples, is shown. Process 500 may include obtaining a request at 502 to access a file stored in a master data repository. In one or more examples, the computing device may display a user interface corresponding to a local file system. For example, a user device at the entity's location may access the local file system via one or more data control systems of the entity. In one or more illustrative examples, one or more data control systems of the entity may include information about... Figure 2 and Figure 3The described data processing controller 202. The user interface may include one or more user interface elements. One or more user interface elements may correspond to one or more files accessible to the computing device via a local file system. The user interface elements may be selectable to provide at least one of one or more requests, one or more commands, or one or more application programming interface (API) calls to the local file system, enabling the computing device to access one or more files. In one or more illustrative examples, at least one of the computing device or one or more data control systems of an entity corresponding to the computing device may, in response to the selection of one or more user interface elements corresponding to one or more files, provide one or more requests, one or more commands, and / or one or more API calls to the local file system. In at least some examples, the local file system may be an instance of file system instructions executed by at least one of one or more servers of an entity or one or more third-party servers. In various examples, one or more third-party servers may be at least one of being supervised, maintained, or controlled by an entity associated with the file system.

[0069] In one or more examples, the files accessible to the computing device can be based on the credentials of the computing device's user. For illustration, a first user may have first credentials indicating that the first user can access one or more first files accessible via the local file system, and a second user may have second credentials indicating that the second user can access one or more second files accessible via the local file system. The user's credentials may be based on the user's location within an entity, one or more of the user's job responsibilities, one or more of the user's locations, or a combination thereof. In various examples, access to files may be restricted based on the type of data included in the files. For example, personal information, health-related information, clinical trial information, or a combination thereof may be subject to restricted access.

[0070] At 504, process 500 may include accessing metadata corresponding to the file using a local file system. The metadata may be stored in at least one accessible data repository on the computing device or local file system. In one or more examples, the metadata may be stored in memory located in the same location as the computing device. That is, the metadata may be stored in local memory. The metadata may also be stored in memory maintained, controlled, or supervised by at least one of the entities associated with the computing device. In other examples, the metadata may be stored in memory of the local file system. In one or more other examples, the metadata may be stored in one or more third-party data repositories. In one or more illustrative examples, the metadata may indicate one or more storage locations of the data included in the file in the master data repository. In various examples, one or more storage locations of the data may include one or more Uniform Resource Locators (URLs). In one or more additional examples, one or more storage locations of the data may include paths associated with the file. In at least some examples, the metadata may indicate one or more types of data included in the file, one or more access restrictions on the file, one or more users associated with the file, at least a portion of the file's version history, one or more combinations thereof, etc.

[0071] Process 500 may further include, at 506, causing one or more requests based on metadata and sent by the local file system to the master data repository to retrieve files. In one or more examples, the local file system may send one or more requests in response to at least one of one or more commands, one or more API calls, or other instructions provided to the local file system by at least one of a computing device or a computing system of an entity associated with the computing device. In at least some examples, the request may include at least one of an identifier of the file to be retrieved or the storage location of the file to be retrieved. In response to one or more requests, the local file system may provide one or more communications to the master data repository to retrieve the files. In various examples, the local file system may communicate with a data repository management system coupled to the master data repository. The local file system may send at least one of one or more commands, one or more API calls, or one or more additional instructions to the data repository management system to retrieve files from the master data repository and send the files to the local file system. In one or more illustrative examples, the master data repository may include a data repository that is remote relative to at least one of the computing device or the local file system. For example, the master data repository may include storage included in a cloud computing architecture that is controlled, maintained, or supervised by at least one third party. In one or more additional illustrative examples, the master data repository may include an object data repository.

[0072] Additionally, at 508, process 500 may include storing the file in a temporary data repository accessible to the local file system. In one or more examples, the file may be stored in a cache memory of the local file system. In one or more additional examples, the file may be stored in a cache memory of the computing system, which is accessible to the local file system and is maintained, controlled, or supervised by at least one of the following:

[0073] Furthermore, process 500 may include, at 510, causing one or more operations to be performed on the file to produce a modified version of the file stored by a temporary data repository. In one or more examples, one or more operations may include modifications to data included in the file. One or more operations may also include at least one of computational analysis, machine learning analysis, or statistical analysis of the data included in the file. In one or more additional examples, one or more operations may include one or more generative operations to generate additional content based on information included in the file. One or more operations may be performed by one or more instances of hardware and / or software executed by a computing device. Additionally, one or more operations may be performed by one or more instances of hardware and / or software executed locally relative to the computing device on one or more computing systems. Furthermore, one or more operations may be performed by one or more instances of at least one of hardware or software executed remotely relative to the computing device. For example, one or more operations may be performed by one or more computing systems of a cloud computing architecture configured to execute software and / or implement hardware to perform one or more operations.

[0074] At 512, process 500 may include storing a modified version of the file by a master data repository. In one or more examples, the storage of the modified version of the file by the master data repository may be performed in response to determining that the file will no longer be modified. For example, at least one of a computing device, a local file system, or a data control system of an entity associated with the computing system may determine that the modified version of the file has been closed. In various examples, at least one of the computing device, the local file system, or a data control system of an entity associated with the computing device may determine that at least one of the computing device, the local file system, or a data control system of an entity associated with the computing system has generated one or more commands and / or one or more requests to close the file or stop modifying file data.

[0075] In one or more illustrative examples, a session may be initiated in response to a request for a file from a computing device. The session may be initiated by at least one of the computing device or a local file system. Data generated and / or modified during the session may be stored as a modified version of the file. In various examples, the modified version of the file may correspond to one or more additional files. For example, analysis may be performed on data included in the file, and additional data may be generated as part of the analysis. The additional data may be stored in one or more additional files. In these cases, the modified version of the file may be at least one of being physically or logically associated with one or more additional files. For illustration, metadata of the modified version of the file may indicate the association between the modified version of the file and one or more additional files. In various examples, the modified version of the file and one or more additional files may be stored in relation to each other. In other examples, the modified version of the file may correspond to a modification of at least one of the data included in the file or metadata associated with the file.

[0076] In at least some examples, the differences between files and their modified versions can be stored in a master data repository. For example, the differences between a previous version of a file and its current modified version can be used to update a previous version of the file to produce a modified version of the file in the master data repository. At least one of the local file system or data control system of the entity associated with the computing device can track changes made to previous versions of the file to produce the modified version of the file. In this way, changes made to previous versions of the file are provided to the master data repository, and the remaining data of the file can be de-de-saturated. That is, in one or more illustrative examples, the local file system and / or data control system can electronically transmit changes to the initial version of the file to the master data repository while deleting the remaining data associated with the file, since the remaining data associated with the file is already stored in the storage location of the initial version of the file.

[0077] In at least some examples, modified versions of a file can be stored in the same storage location in the master data repository as previous versions of the file. For example, a modified version of a file can be stored in an object in the master data repository corresponding to the original file. In various examples, at least one of a computing device, data control system, or local file system can generate one or more commands and / or API calls to store the modified version of the file in the same object in the master data repository as the previous version of the file. In this way, when the initial version of the file is created, an object can be created in the master data repository, and subsequent versions of the file can also be stored in the same object. Therefore, the object stored in the file's master data repository can serve as a source of the file's ground truth and enable multiple file systems to access the file using a common storage identifier.

[0078] Process 500 may include removing a modified version of the file from the temporary data repository at point 514. In one or more examples, the modified version of the file may be deleted from the temporary data repository after it has been stored in the master data repository. In various examples, the modified version of the file may be deleted from the temporary data repository after receiving an indication from the database management system of the master data repository that the modified version of the file has been stored in the location of the master data repository. In at least some examples, at least one of the local file system, the computing device, or the data control system of an entity associated with the computing device may provide one or more command and / or API calls to remove the modified version of the file from the temporary data repository. In one or more illustrative examples, removing the modified version of the file from the temporary data repository may be part of a process of dehydrating the file-related data in the temporary data repository, storing updated metadata of the file stored by the local file system. As part of the dehydration process, a stub file may be generated indicating at least one of the file's identifier or the file's storage location in the master data repository.

[0079] Additionally, process 500 may include, at 516, storing updated metadata of the file on the local file system. In one or more examples, the updated metadata may include the storage location of a modified version of the file. In various examples, the metadata may include a stub file. The stub file may indicate the storage location of a modified version of the file, allowing the local file system to retrieve the modified version of the file from the master data repository, providing one or more computing devices with access to the modified version of the file.

[0080] In at least some examples, updated metadata may be provided to one or more additional file systems. In one or more examples, the local file system may be one of multiple local file systems of the entity. In various examples, each local file system may correspond to one or more locations of the entity. In one or more illustrative examples, each local file system may correspond to a different location of the entity. In this way, a single local file system may provide access to files for one or more computing devices located at a single location of the entity. Additionally, each local file system may be coupled to or have access to local data storage devices and / or computing resources, such as one or more server computers, at different locations of the entity.

[0081] If one of multiple local file systems determines that the metadata of a file accessed by that local file system has been modified, the local file system can make the updated metadata available to other local file systems of the entity. Therefore, computing devices corresponding to other local file systems can access the updated version of the file generated by the local file system. For illustration, a local file system associated with the entity and used to access the initial version of a file and then store modified versions of the file can transmit information to other local file systems of the entity that the initial version of the file has been modified, and provide this information to other local file systems so that computing devices corresponding to those other local file systems can access the modified versions of the file. In this way, a single identifier for the storage location of modified versions of a file (e.g., an object of a master data repository) can be provided across multiple local file systems, enabling computing devices in multiple locations and served by multiple local file systems to access modified versions of the file and control access to the file, ensuring that multiple users do not modify the file simultaneously. In one or more illustrative examples, the object of the master data repository storing the file can be the true source of data associated with multiple versions of the file. Therefore, the object of the master data repository created in response to the initial version of the file being created can be used as the true source of subsequent versions of the generated file.

[0082] In various examples, one or more local file systems may be dedicated to accessing data with one or more access restrictions. For example, one or more local file systems may be dedicated to accessing at least one of personal data, personal health information, medical data, genomic data, electronic medical records, insurance information, clinical trial information, financial information, etc. Additionally, data with one or more access restrictions may be stored in one or more master data repositories that comply with one or more regulations relating to the storage of personal data, personal health information, medical data, genomic data, electronic medical records, insurance information, clinical trial information, financial information, and one or more combinations thereof. In this way, access to and storage of data with one or more access restrictions can be controlled by one or more local file systems dedicated to accessing and storing data with one or more access restrictions.

[0083] Although process 500 has been described with respect to a modified version of generating and storing existing files, at least a portion of the operations of process 500 can also be implemented to create and store newly created files. For example, a user of a computing device can execute an application to create content. The content may include at least one of text content, image content, video content, or audio content. Content can be created in relation to a file. In one or more examples, while the content of a file is being created or after an initial version of the file content has been generated, the file may be stored in a temporary data repository accessible to the local file system. The local file system can be used to determine the identifier of the file. In various examples, the identifier of the file may be based at least in part on input provided by the user of the computing device.

[0084] Additionally, the local file system can communicate with the master data repository to create a new object of the master data repository in which the file is stored. For illustration, the local file system can transmit at least one of one or more commands or API calls to the master data repository to cause the master data repository to create a new object for the file. The local file system can also transmit at least one of one or more commands or API calls to the master data repository to cause data included in the file to be stored in the object corresponding to the file in the master data repository. Additionally, the local file system can provide metadata of the newly created file to other file systems. In this way, a computing device accessing an additional file system can access the file by communicating with the master data repository based on the file's metadata. After the initial version of the file is stored in the object of the master data repository, process 500 can be implemented to modify the file.

[0085] In one or more illustrative examples, the data to be accessed from the master data repository may include genetic data. In various examples, genetic data associated with one or more diagnostic tests may be generated to identify the presence or absence of biological conditions in a subject. In at least some examples, numerous computational operations may be performed on the genetic data. Computational operations may include analysis of the genetic data. In one or more examples, the analysis of the genetic data may modify the genetic data. In various examples, the modified genetic data may include additional data derived from the genetic data. Computational operations may be performed locally using processing resources associated with a file system. Computational operations may also be performed by one or more cloud computing providers. In response to the completion of computational operations and in response to one or more users no longer accessing the genetic data and / or modified genetic data, the genetic data and / or modified genetic data may be removed from temporary storage at the site and stored within the master object data repository. Updates to the metadata corresponding to the genetic data and / or modified genetic data may be propagated to every file system within the organization, allowing users at different locations within the organization to access the genetic data and / or modified genetic data in a uniform manner.

[0086] Example Example 1. A method comprising: storing a plurality of files in a standardized format by a data processing controller on a centralized server; receiving, by the data processing controller, a first request to access a first file in the plurality of files from a first computing system, the first computing system generating the request using a first type of file system; receiving, by the data processing controller, a second request to access a second file in the plurality of files from a second computing system, the second computing system generating the request using a second type of file system different from the first type of file system of the first computing system; and, in response to receiving the first request and the second request, controlling access to the first file and the second file by the data processing controller.

[0087] Example 2. The method according to Example 1 further includes: converting the first file from a standardized format to a format compatible with a first type of file system by the first computing system.

[0088] Example 3. The method according to Example 2 further includes: storing the first file on a cache of the first computing system; and tracking changes to the first file in the cache of the first computing system.

[0089] Example 4. The method according to Example 3 further includes: updating a data object stored on a centralized server associated with the first file based on changes to the first file in the cache of the first computing system, wherein the first file stored on the centralized server represents the source of the true version of the first file.

[0090] Example 5. According to the method described in Example 4, the data object is periodically updated to reflect changes to the first file.

[0091] Example 6. The method according to any one of Examples 4-5 further includes: merging changes to the first file stored in the data object with the source of the actual version of the first file.

[0092] Example 7. The method according to any one of Examples 4-6 further includes: in response to determining that the second computing system is currently accessing the first file, transmitting changes to the first file stored in the data object to the second computing system.

[0093] Example 8. The method according to any one of Examples 1-7 further includes: storing one or more data objects associated with each of the plurality of files, the one or more data objects representing the storage location of each file on a centralized server, the locking state of each file, and changes made to each file by a cached copy of each file on the respective computing system.

[0094] Example 9. The method according to Example 8 further includes: in response to receiving a first request, using one or more data objects associated with the first file to determine whether the first file is currently being used by another computing system.

[0095] Example 10. The method according to any one of Examples 1-9, wherein the first computing system includes a server associated with a cloud service provider that is different from the cloud service provider corresponding to the centralized server.

[0096] Example 11. The method according to any one of Examples 1-10, wherein the first computing system is associated with and managed by a life science service provider.

[0097] Example 12. The method according to Example 11, wherein the second computing system is associated with and managed by one or more third parties relative to the life science service provider.

[0098] Example 13. The method according to any one of Examples 1-12, wherein the first file includes patient data, which includes genomic information of multiple subjects.

[0099] Example 14. The method according to any one of Examples 1-13 further includes: performing analysis on at least a portion of a batch of data in a first file via a bioinformatics system implemented by a first type of file system; and determining one or more characteristics of a subject corresponding to at least a portion of the batch of data based on the performed analysis.

[0100] Example 15. The method according to Example 14, wherein: one or more features include one or more genomic mutations present in nucleic acids derived from samples obtained from one or more subjects; and the nucleic acids correspond to cell-free deoxyribonucleic acid (DNA) extracted from bodily fluid samples obtained from one or more subjects.

[0101] Example 16. The method according to any one of Examples 14-15, wherein one or more features include resistance to treatments provided to one or more subjects in connection with biological conditions present in one or more subjects.

[0102] Example 17. The method described in Example 16, wherein the biological condition corresponds to a form of cancer.

[0103] Example 18. The method according to any one of Examples 14-16, wherein the analysis includes determining recommendations for treatments provided to one or more subjects to treat biological conditions present in one or more subjects.

[0104] Example 19. A system comprising: one or more hardware processing units; and one or more computer-readable storage media storing computer-executable instructions that, when executed by the one or more hardware processing units, cause the system to perform operations including: storing a plurality of files in a standardized format by a data processing controller on a centralized server; receiving, by the data processing controller, a first request to access a first file among the plurality of files from a first computing system, the first computing system generating the request using a first type of file system; receiving, by the data processing controller, a second request to access a second file among the plurality of files from a second computing system, the second computing system generating the request using a second type of file system different from the first type of file system of the first computing system; and, in response to receiving the first and second requests, controlling access to the first and second files by the data processing controller.

[0105] Example 20. A non-transitory computer-readable storage medium storing one or more computer-executable instructions that, when executed by one or more hardware processing units, cause a system to perform operations including: storing a plurality of files in a standardized format by a data processing controller on a centralized server; receiving, by the data processing controller, a first request to access a first file among the plurality of files from a first computing system, the first computing system generating the request using a first type of file system; receiving, by the data processing controller, a second request to access a second file among the plurality of files from a second computing system, the second computing system generating the request using a second type of file system different from the first type of file system of the first computing system; and, in response to receiving the first and second requests, controlling access to the first and second files by the data processing controller.

[0106] Example 21 is a method comprising: obtaining a request from a computing system including memory and one or more processors to access a file stored in a master data repository, wherein the computing system is maintained, controlled, or supervised by at least one of a first entity, and the master data repository is maintained, controlled, or supervised by at least one of a second entity including a cloud storage service provider; accessing metadata corresponding to the file by the computing system and using a local file system, wherein the metadata indicates an object in the master data repository, the object including data corresponding to the file; causing one or more requests based on the metadata and sent from the local file system to the master data repository to retrieve the file; causing the file to be stored in a temporary data repository of the first entity, the temporary data repository being accessible from the local file system; causing one or more computational operations to be performed on the file by the computing system to produce a modified version of the file stored in the temporary data repository; causing the modified version of the file to be stored in an object in the master data repository via the local file system by the computing system; causing the modified version of the file to be removed from the temporary data repository by the computing system; and causing updated metadata of the modified version of the file to be stored in additional storage of the first entity accessible from the local file system by the computing system.

[0107] In Example 22, the subject matter of Example 21 may optionally include: wherein the computing system receives a request to access a file from a user’s computing device, wherein the computing device corresponds to a first entity and has access to a local file system; and the computing system includes a data processing controller executed on one or more servers of the first entity.

[0108] In Example 23, the subject of Example 22 may optionally include: wherein a local file system and a computing device are associated with a location of a first entity, the location of which includes at least one server computer executing an instance of the local file system and includes a temporary data repository accessible by the local file system.

[0109] In Example 24, the subject matter of any one or more of Examples 22-23 may optionally include: wherein, in response to receiving a request to access a file, the computing system determines at least one of one or more commands or one or more application programming interface calls of the local file system to cause the local file system to retrieve metadata from the additional storage of the first entity and retrieve the file from the master data repository based on the metadata; and the computing system sends at least one of one or more commands or one or more application programming interface calls to the local file system.

[0110] In Example 25, the subject of Example 24 may optionally include: wherein a computing device displays one or more user interfaces comprising multiple user interface elements, the multiple user interface elements including user interface elements corresponding to files, and the user interface elements are optionally instances of a local file system that send at least one of one or more commands or one or more application programming interface calls to be executed by at least one server computer.

[0111] In Example 26, any one or more of the topics in Examples 22-25 may optionally include: wherein, in response to determining that no further modifications are made to the data corresponding to the file, a modified version of the file is stored by the master data repository.

[0112] In Example 27, the subject of Example 26 may optionally include: initiating a session of an instance of the application in response to one or more additional requests received from a computing device; wherein a file is modified according to one or more operations based on input provided during the session to produce a modified version of the file; and one or more operations correspond to at least one of modifying existing data of the file or adding additional data to the file.

[0113] In Example 28, the subject of Example 27 may optionally include: wherein an instance of the application is executed by at least one of a computing device, one or more servers, or one or more cloud computing services.

[0114] In Example 29, any one or more of the topics in Examples 26-28 may optionally include: wherein determining that no further modifications will be made to the data corresponding to the file includes: determining that the session of the application instance has been terminated.

[0115] In Example 30, the subject of any one or more of Examples 21-29 may optionally include: wherein removing a modified version of a file from a temporary data repository is part of a process of dehydrating data associated with the file in the temporary data, the process including storing updated metadata of the file stored by a local file system, and including generating a stub file during the dehydration process, wherein the stub file indicates at least one of the file's identifier or the file's storage location in the main data repository.

[0116] In Example 31, the subject matter of any one or more of Examples 21-30 may optionally include: wherein the data stored in the file includes scientific data, which includes at least one of genomic information, genetic information, metabolomics information, transcriptomics information, fragmentome information, immune receptor information, methylation information, epigenome information, or proteome information; and at least a portion of one or more computational operations for generating a modified version of the file is performed by the bioinformatics pipeline of the first entity.

[0117] Example 32 is a computing system comprising: one or more hardware processors; and a memory storing computer-readable instructions that, when executed by the one or more hardware processors, cause the one or more hardware processors to perform operations including: obtaining a request to access a file stored in a master data repository, wherein the computing system is maintained, controlled, or supervised by at least one of a first entity, and the master data repository is maintained, controlled, or supervised by at least one of a second entity, including a cloud storage service provider; accessing metadata corresponding to the file using a local file system, wherein the metadata indicates an object in the master data repository, the object including data corresponding to the file; causing one or more requests based on the metadata and sent from the local file system to the master data repository to retrieve the file; causing the file to be stored in a temporary data repository of the first entity, and the temporary data repository being accessible from the local file system; causing one or more computational operations to be performed on the file to produce a modified version of the file stored in the temporary data repository; causing the modified version of the file to be stored in an object in the master data repository via the local file system; causing the modified version of the file to be removed from the temporary data repository; and causing updated metadata of the modified version of the file to be stored in additional memory of the first entity accessible from the local file system.

[0118] In Example 33, the subject matter of Example 32 may optionally include: wherein the request to access a file is received from a user’s computing device, wherein the computing device corresponds to a first entity and has access to a local file system; the computing system includes a data processing controller executed on one or more servers of the first entity; the local file system and the computing device are associated with a location of the first entity, the location of the first entity including at least one server computer executing an instance of the local file system, and including a temporary data repository accessible by the local file system.

[0119] In Example 34, the subject of Example 33 may optionally include: wherein the venue is one of a plurality of venues of the first entity, each of the plurality of venues corresponding to a corresponding local file system, including at least one corresponding server computer executing the corresponding local file system, and including a corresponding temporary data repository accessible by the corresponding local file system.

[0120] In Example 35, the subject matter of Example 34 may optionally include: wherein the memory stores additional computer-readable instructions that, when executed by one or more hardware processors, cause one or more hardware processors to perform additional operations, the additional operations including: causing updated metadata of a file to be stored in additional memory of an additional location in multiple locations; obtaining an additional request from a computing device of the additional location to access a modified version of the file; and causing one or more additional requests to be sent, based on the updated metadata, by an additional file system of the additional location to a master data repository to retrieve the modified version of the file.

[0121] In Example 36, the subject of Example 35 may optionally include: wherein a modified version of the file is stored in the same object as the initial version of the file, and the updated metadata includes the storage location of the master data repository corresponding to that object.

[0122] In Example 37, the subject matter of any one or more of Examples 35-36 may optionally include: wherein the memory stores additional computer-readable instructions that, when executed by one or more hardware processors, cause one or more hardware processors to perform additional operations, the additional operations including: causing a modified version of the file to be stored in an additional temporary data repository in an additional location accessible by the additional file system; causing the additional modified version of the file to be stored by the additional local file system in an object identical to the initial version of the file in the main data repository; and causing additional updated metadata of the file to be stored in corresponding storage devices in multiple locations.

[0123] In Example 38, the subject matter described in Example 37 may optionally include: wherein the memory stores additional computer-readable instructions that, when executed by one or more hardware processors, cause one or more hardware processors to perform additional operations, the additional operations including: removing an additional modified version of a file from an additional temporary data repository; and generating additional updated metadata of the additional modified version of the file while removing the additional modified version of the file from the additional temporary data repository.

[0124] Example 39 is one or more computer-readable storage media storing computer-readable instructions that, when executed by one or more hardware processors, cause one or more hardware processors to perform operations including: obtaining a request to access a file stored in a master data repository, wherein a computing system is maintained, controlled, or supervised by at least one of a first entity, and the master data repository is maintained, controlled, or supervised by at least one of a second entity, including a cloud storage service provider; accessing metadata corresponding to the file using a local file system, wherein the metadata indicates an object in the master data repository, the object including data corresponding to the file; causing one or more requests based on the metadata and sent from the local file system to the master data repository to retrieve the file; causing the file to be stored in a temporary data repository of the first entity, and the temporary data repository being accessible from the local file system; causing one or more computational operations to be performed on the file to produce a modified version of the file stored in the temporary data repository; causing the modified version of the file to be stored in an object in the master data repository via the local file system; causing the modified version of the file to be removed from the temporary data repository; and causing updated metadata of the modified version of the file to be stored in additional storage of the first entity accessible from the local file system.

[0125] In Example 40, the subject of Example 39 may optionally include: wherein files and modified versions of files are stored in an object of a master data repository, the object having a storage location identifier that multiple file systems can access data of one or more versions of the file stored in the object through the storage location identifier, and the object being the true source of data of one or more versions of the file.

[0126] Figure 6 This is a block diagram illustrating components of a machine 600 in the form of a computer system implemented according to one or more examples, which can read and execute instructions from one or more machine-readable media to perform any or more methods described herein. Specifically, Figure 6A graphical representation of machine 600 in an example form of a computer system is shown, in which instructions 602 (e.g., software, programs, applications, applets, or other executable code) are used to cause machine 600 to perform any or more of the methods discussed herein. Thus, instructions 602 can be used to implement the modules or components described herein. Instructions 602 transform a general, unprogrammed machine 600 into a specific machine 600 programmed to perform the described and illustrated functions in the described manner. In alternative embodiments, machine 600 operates as a standalone device or may be coupled (e.g., networked) to other machines. In a networked deployment, machine 600 may operate as a server machine or client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. Machine 600 may include, but is not limited to, server computers, client computers, personal computers (PCs), tablet computers, laptop computers, netbooks, set-top boxes (STBs), personal digital assistants (PDAs), entertainment media systems, cellular phones, smartphones, mobile devices, wearable devices (e.g., smartwatches), smart home devices (e.g., smart home appliances), other smart devices, networked home appliances, network routers, network switches, bridges, or any machine capable of sequentially or otherwise executing instructions 602, which specify the actions to be taken by machine 600. Furthermore, although only a single machine 600 is shown, the term "machine" should also be understood to include a collection of machines that individually or jointly execute instructions 602 to perform any or more of the methods discussed herein.

[0127] Machine 600 may include processor 604, memory / storage device 606, and I / O components 608, which may be configured to communicate with each other, for example, via bus 610. In example embodiments, processor 604 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) may include, for example, processors 612 and 614 capable of executing instruction 602. The term "processor" is intended to include multi-core processor 604, which may include two or more independent processors (sometimes referred to as "cores") capable of executing instruction 602 simultaneously. Although Figure 6More than one processor 604 is shown, but machine 600 may include a single processor 612 with a single core, a single processor 612 with multiple cores (e.g., a multi-core processor), more than one processor 612, 614 with a single core, more than one processor 612, 614 with multiple cores, or any combination thereof.

[0128] Memory / storage device 606 may include memory, such as main memory 616 or other memory / storage devices, and storage cell 618, both of which may be accessed by processor 604, such as via bus 610. Storage cell 618 and main memory 616 store instructions 602 containing any one or more of the methods or functions described herein. During execution of instructions 602 by machine 600, instructions 602 may also reside wholly or partially within main memory 616, storage cell 618, at least one of processor 604 (e.g., within the processor's cache memory), or any suitable combination thereof. Thus, the memory of main memory 616, storage cell 618, and processor 604 are examples of machine-readable media.

[0129] I / O component 608 may include a wide variety of components to receive input, provide output, generate output, transmit information, exchange information, capture measurement values, and so on. The specific I / O component 608 included in a particular machine 600 will depend on the type of machine. For example, a portable machine such as a mobile phone may include a touch input device or other such input mechanism, while a headless server machine may not include such a touch input device. It should be understood that component 608 may include... Figure 6 Many other components are not shown. The grouping of I / O components 608 by function is merely for the purpose of simplifying the discussion below, and the grouping is by no means limiting. In various example embodiments, I / O components 608 may include user output components 620 and user input components 622. User output components 620 may include visual components (e.g., displays such as plasma display panels (PDPs), light-emitting diode (LED) displays, liquid crystal displays (LCDs), projectors, or cathode ray tube (CRT) displays), acoustic components (e.g., speakers), haptic components (e.g., vibration motors, resistive mechanisms), other signal generators, and so on. User input components 622 may include alphanumeric input components (e.g., keyboards, touchscreens configured to receive alphanumeric input, photoelectric keyboards, or other alphanumeric input components), point-based input components (e.g., mice, touchpads, trackballs, joysticks, motion sensors, or other pointing tools), haptic input components (e.g., physical buttons, touchscreens or other haptic input components that provide the position or force of a touch or touch gesture), audio input components (e.g., microphones), and so on.

[0130] In another example implementation, I / O component 608 may include biometric component 624, motion component 626, environmental component 628, or position component 630, as well as a wide range of other components. For example, biometric component 624 may include components for detecting facial expressions (e.g., hand expressions, facial expressions, vocal expressions, body posture, or eye tracking), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), and recognizing a person (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or EEG-based recognition). Motion component 626 may include accelerometer components (e.g., accelerometers), gravity sensor components, rotation sensor components (e.g., gyroscopes), etc. Environmental component 628 may include, for example, a lighting sensor component (e.g., a photometer), a temperature sensor component (e.g., one or more thermometers that detect ambient temperature), a humidity sensor component, a pressure sensor component (e.g., a barometer), an acoustic sensor component (e.g., one or more microphones that detect background noise), a proximity sensor component (e.g., an infrared sensor that detects nearby objects), a gas sensor (e.g., a gas detection sensor for the safe detection of hazardous gas concentrations or for measuring pollutants in the atmosphere), or other components that can provide indications, measurements, or signals corresponding to the surrounding physical environment. Position component 630 may include a position sensor component (e.g., a GPS receiver component), an altitude sensor component (e.g., an altimeter or barometer from which altitude can be derived), a direction sensor component (e.g., a magnetometer), etc.

[0131] Communication can be implemented using a variety of technologies. I / O component 608 may include communication component 632 operable to couple machine 600 to network 634 or device 636. For example, communication component 632 may include a network interface component or other suitable device to interface with network 634. In further examples, communication component 632 may include wired communication components, wireless communication components, cellular communication components, near field communication (NFC) components, Bluetooth® components (e.g., Bluetooth® Low Energy), Wi-Fi® components, and other communication components that provide communication via other modalities. Device 636 may be another machine 600 or any of a variety of peripheral devices (e.g., peripheral devices coupled via USB).

[0132] Furthermore, communication component 632 can detect identifiers or include components operable for detecting identifiers. For example, communication component 632 may include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor that detects one-dimensional barcodes such as Universal Product Code (UPC) barcodes, multi-dimensional barcodes such as Quick Response (QR) codes, Aztec codes, Data Matrix, Dataglyph, MaxiCode, PDF417, Ultra codes, UCC RSS-2D barcodes, and other optical codes), or an acoustic detection component (e.g., a microphone for identifying audio signals of the tag). Additionally, various information can be derived via communication component 632, such as location via Internet Protocol (IP) geolocation, location via Wi-Fi® signal triangulation, location by detecting NFC beacon signals that can indicate a specific location, and so on.

[0133] As used herein, a “component” refers to a device, physical entity, or logic having boundaries defined by function or subroutine calls, branch points, APIs, or other technologies that provide specific processing or control functionality, such as partitions or modularization. Components can be combined with other components through their interfaces to perform machine processes. A component can be a packaged functional hardware unit designed for use with other components, or it can be part of a program that typically performs related functions. Components can constitute software components (e.g., code implemented on a machine-readable medium) or hardware components. A “hardware component” is a tangible unit capable of performing certain operations and can be configured or arranged in some physical manner. In various example implementations, one or more computer systems (e.g., standalone computer systems, client computer systems, or server computer systems) or one or more hardware components (e.g., processors or a group of processors) of a computer system can be configured as hardware components by software (e.g., an application or a portion of an application) that operates to perform certain operations described herein.

[0134] Hardware components can also be implemented mechanically, electronically, or in any suitable combination thereof. For example, a hardware component may include dedicated circuitry or logic permanently configured to perform certain operations. A hardware component may be a dedicated processor, such as a field-programmable gate array (FPGA) or an ASIC. A hardware component may also include programmable logic or circuitry temporarily configured by software to perform certain operations. For example, a hardware component may include software executed by a general-purpose processor 604 or another programmable processor. Once configured by such software, the hardware component becomes a specific machine (or a specific component of machine 600), uniquely tailored to perform the configured functions, and is no longer the general-purpose processor 604. It should be understood that the decision to implement a hardware component mechanically, in dedicated and permanently configured circuitry, or in temporarily configured circuitry (e.g., configured by software) may be driven by cost and time considerations. Therefore, the phrase “hardware component” (or “hardware-implemented component”) should be understood to include tangible entities that are physically constructed, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmable) that operate or perform certain operations described herein. Given the implementation where hardware components are temporarily configured (e.g., programmed), each hardware component does not need to be configured or instantiated at any given time. For example, in the case where the hardware components include a general-purpose processor 604 configured by software as a dedicated processor, the general-purpose processor 604 can be configured as different dedicated processors (e.g., including different hardware components) at different times. The software accordingly configures a specific processor 612, processor 414, or processor 604, for example, constituting a specific hardware component in one time instance and different hardware components in different time instances.

[0135] Hardware components can provide information to other hardware components and receive information from other hardware components. Therefore, the described hardware components can be considered communicatively coupled. When more than one hardware component is present simultaneously, communication can be achieved through signal transmission between two or more hardware components (e.g., via appropriate circuitry and buses). In embodiments where more than one hardware component is configured at different times or exemplifies different hardware components, communication between these hardware components can be achieved, for example, by storing and retrieving information in a memory structure accessible to more than one hardware component. For example, one hardware component can perform an operation and store the output of that operation in a memory device communicatively coupled to it. Another hardware component can then access that memory device at a later time to retrieve and process the stored output.

[0136] Hardware components can also initiate communication with input or output devices and operate on resources (e.g., collections of information). Various operations of the example methods described herein can be performed at least in part by one or more processors 604, which are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors 604 can constitute processor-implemented components that operate to perform one or more operations or functions described herein. As used herein, "processor-implemented component" refers to a hardware component implemented using one or more processors 604. Similarly, the methods described herein can be implemented at least in part by processors, with a particular processor 612, processor 414, or processor 604 being an instance of hardware. For example, at least some operations of the methods can be performed by one or more processors 604 or processor-implemented components. Furthermore, one or more processors 604 can also operate to support the performance of relevant operations in a "cloud computing" environment or as "Software as a Service" (SaaS). For example, at least some operations can be performed by a group of computers (as instances of machine 400 including processor 604), and these operations can be accessed via network 634 (e.g., the Internet) and via one or more appropriate interfaces (e.g., APIs). The performance of some operations can be distributed among processors, residing not only within a single machine 600 but also deployed across many machines. In some example implementations, processor 604 or processor-implemented components can reside in a single geographic location (e.g., within a home environment, office environment, or server cluster). In other example implementations, processor 604 or processor-implemented components can be distributed across many geographic locations.

[0137] Figure 7 This is a block diagram illustrating a system 700 including an example software architecture 702, which can be used in conjunction with various hardware architectures described herein. Figure 7 This is a non-limiting example of a software architecture, and it will be understood that many other architectures can be implemented to facilitate the functionality described herein. Software architecture 702 can be used in, for example... Figure 6 The execution occurs on the hardware of machine 600, which includes processor 604, memory / storage device 606, and input / output (I / O) components 608, etc. A representative hardware layer 704 is shown, and can represent, for example... Figure 6 The machine 600. A representative hardware layer 704 includes a processing unit 706 having associated executable instructions 708. The executable instructions 708 represent executable instructions of the software architecture 702, including implementations of the methods, components, etc., described herein. Hardware layer 704 also includes at least one of a memory or storage module memory / storage device 710, which also has the executable instructions 708. Hardware layer 704 may also include other hardware 712.

[0138] exist Figure 7 In the example architecture, software architecture 702 can be conceptualized as a stack of layers, where each layer provides specific functionality. For example, software architecture 702 may include layers such as operating system 714, libraries 716, framework / middleware 718, applications 720, and presentation layer 722. Operationally, applications 720 or other components within a layer can invoke API calls 724 via the software stack and receive messages 726 in response to API calls 724. The illustrated layers are representative in nature; not all software architectures have all layers. For example, some mobile or special-purpose operating systems may not provide a framework / middleware 718, while others may provide such a layer. Other software architectures may include additional or different layers.

[0139] Operating system 714 can manage hardware resources and provide public services. Operating system 714 may include, for example, a kernel 728, services 730, and drivers 732. Kernel 728 can act as an abstraction layer between the hardware layer and other software layers. For example, kernel 728 may be responsible for memory management, processor management (e.g., scheduling), component management, networking, security settings, and so on. Services 730 can provide other public services to other software layers. Drivers 732 are responsible for controlling or interfacing with the underlying hardware. For example, depending on the hardware configuration, drivers 732 may include display drivers, camera drivers, Bluetooth® drivers, flash memory drivers, serial communication drivers (e.g., Universal Serial Bus (USB) drivers), Wi-Fi® drivers, audio drivers, power management drivers, and so on.

[0140] Library 716 provides common infrastructure used by application 720, other components, or layers. Library 716 provides functionality that allows other software components to perform tasks more easily than directly interfaceing with the underlying operating system 714 functions (e.g., kernel 728, services 730, drivers 732). Library 716 may include system libraries 734 (e.g., the C standard library), which provide functions such as memory allocation functions, string manipulation functions, mathematical functions, etc. Furthermore, library 716 may include API libraries 736, such as media libraries (e.g., libraries supporting the rendering and manipulation of various media formats such as MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.), graphics libraries (e.g., the OpenGL framework for rendering 2D and 3D graphics content on a display), database libraries (e.g., SQLite providing various relational database functions), web libraries (e.g., WebKit providing web browsing functionality), etc. Library 716 may also include a wide variety of other libraries 738 to provide many other APIs to application 720 and other software components / modules.

[0141] The framework / middleware 718 (sometimes also called middleware) provides a higher level of common infrastructure that can be used by applications 720 or other software components / modules. For example, the framework / middleware 718 can provide various graphical user interface functions, high-level resource management, high-level location services, and so on. The framework / middleware 718 can provide a broad spectrum of other APIs that applications 720 or other software components / modules can utilize, some of which may be specific to a particular operating system 714 or platform.

[0142] Application 720 includes built-in applications 740 and third-party applications 742. Examples of representative built-in applications 740 may include, but are not limited to, contact applications, browser applications, book reader applications, location applications, media applications, messaging applications, or game applications. Third-party applications 742 may include applications developed by entities other than the vendor of a specific platform using the Android™ or iOS™ Software Development Kit (SDK), and may be mobile software running on a mobile operating system such as iOS™, Android™, Windows® Phone, or other mobile operating systems. Third-party applications 742 may invoke API calls 724 provided by the mobile operating system (such as operating system 714) to facilitate the functionality described herein.

[0143] Application 720 can use built-in operating system functions (e.g., kernel 728, services 730, drivers 732), libraries 716, and frameworks / middleware 718 to create a UI for user interaction with the system. Optionally or additionally, in some systems, user interaction can occur through a presentation layer (such as presentation layer 722). In these systems, application / component "logic" can be decoupled from the various aspects of the application / component interacting with the user.

[0144] At least some of the processes described herein can be embodied in computer-readable instructions executable by one or more processors, such that operations of the processes can be performed, in part or in whole, by functional components of one or more computer systems. Therefore, in some cases, the computer-implemented processes described herein are examples for reference only. However, in other embodiments, at least some operations of the computer-implemented processes described herein can be deployed on a variety of other hardware configurations. Therefore, the computer-implemented processes described herein are not intended to be limited to those relating to… Figure 6 and Figure 7 The system and configuration described may be implemented wholly or partially by one or more other systems and / or components.

[0145] Although the flowcharts described herein may present operations as a sequential process, many operations can be executed in parallel or concurrently. Furthermore, the order of operations can be rearranged. When an operation of a process completes, it is terminated. A process can correspond to a method, program, algorithm, etc. The operations of a method can be executed in whole or in part, can be combined with some or all of the operations from other methods, and can be executed by any number of different systems (such as the system described herein) or any part thereof (such as a processor included in any system).

[0146] As used herein, a component can refer to a device, physical entity, or logic having boundaries defined by function or subroutine calls, branch points, APIs, or other technologies that provide specific processing or control functionality, such as partitions or modularization. Components can be combined with other components through their interfaces to perform machine processes. A component can be a packaged functional hardware unit designed for use with other components, or it can be part of a program that typically performs related functions. Components can constitute software components (e.g., code implemented on a machine-readable medium) or hardware components. A “hardware component” is a tangible unit capable of performing certain operations and can be configured or arranged in some physical manner. In various example implementations, one or more computer systems (e.g., standalone computer systems, client computer systems, or server computer systems) or one or more hardware components (e.g., processors or a group of processors) of a computer system can be configured as hardware components by software (e.g., an application or a portion of an application) that operates to perform certain operations described herein.

[0147] It should be understood that, as long as the teaching remains operational, the individual steps used in the method of this teaching can be performed in any order and / or simultaneously. Furthermore, it should be understood that, as long as the teaching remains operational, the apparatus and method of this teaching can include any number or all of the described implementations.

[0148] The steps of the methods disclosed herein or the steps performed by the systems disclosed herein may be performed at the same time or at different times, and / or in the same geographical location or in different geographical locations (e.g., countries). The steps of the methods disclosed herein may be performed by the same person or different people.

[0149] Various implementations of the systems, devices, and methods have been described herein. These implementations are given by way of example only and are not intended to limit the scope of the claimed invention. Furthermore, it should be understood that the various features of the described implementations can be combined in various ways to produce many additional implementations. Moreover, although various materials, sizes, shapes, configurations, and positions have been described for use with the disclosed implementations, other materials, sizes, shapes, configurations, and positions may be utilized without departing from the scope of the claimed invention.

[0150] Those skilled in the art will recognize that implementations may include fewer features than those shown in any of the individual implementations described above. The implementations described herein are not intended to be an exhaustive representation of how various features can be combined. Therefore, implementations are not mutually exclusive combinations of features; rather, as will be understood by those skilled in the art, implementations may include combinations of different individual features selected from different individual implementations. Furthermore, unless otherwise stated, elements described with respect to one implementation may be implemented in other implementations, even if not described in those implementations. While a dependent claim may refer in a claim to a specific combination of one or more other claims, other implementations may also include combinations of dependent claims with the subject matter of each other dependent claim, or combinations of one or more features with other dependent or independent claims. Such combinations are presented herein unless it is stated that a particular combination is not intended to be used. Furthermore, it is intended that features of a claim be included in any other independent claim, even if that claim is not directly dependent on that independent claim.

[0151] Furthermore, references to "one implementation," "implementation," or "some implementations" in the specification imply that a particular feature, structure, or characteristic described in connection with the implementation is included in at least one of the teachings. The phrase "in one implementation" appearing in various places in the specification does not necessarily refer to the same implementation.

[0152] Any inclusion by reference of the foregoing documents is limited such that no subject matter contrary to the express disclosure herein is incorporated. Any inclusion by reference of the foregoing documents is further limited such that any claims included in the documents are not incorporated herein by reference. Any inclusion by reference of the foregoing documents is also further limited such that any definitions provided in the documents are not incorporated herein by reference unless expressly included herein.

[0153] Although implementations have been described with reference to specific examples, it will be apparent that various modifications and changes can be made to these implementations without departing from the broader spirit and scope of this disclosure. Therefore, the specification and drawings are to be considered illustrative rather than restrictive. The drawings, which form part of this application, illustrate, by way of illustration and not limitation, specific implementations in which the subject matter can be practiced. The illustrated implementations are described in sufficient detail to enable those skilled in the art to implement the teachings disclosed herein. Other implementations may be used and derived therefrom, such that structural and logical substitutions and changes can be made without departing from the scope of this disclosure. Therefore, this detailed description should not be construed as restrictive, and the scope of the various implementations is defined only by the appended claims together with their equivalents, which enjoy the full scope of those claims.

[0154] While specific implementations have been described and illustrated herein, it should be understood that the illustrated specific implementations can be replaced by any arrangement calculated to achieve the same purpose. This disclosure is intended to cover any and all modifications to various implementations. After reading the above description, those skilled in the art will understand combinations of the above implementations and other implementations not specifically described herein.

[0155] In this document, the terms “a” or “an”, as is common in patent documents, are used to include one or more, and are not related to any other example or use of “at least one” or “one or more”. In this document, the term “or” is used to refer to a non-exclusive “or”, so unless otherwise stated, “A or B” includes “A but not B”, “B but not A”, and “A and B”. In this document, the terms “including” and “in which” are used as their plain English equivalents to the corresponding terms “comprising” and “wherein”. Furthermore, in the appended claims, the terms “including” and “comprising” are open-ended, meaning that a system, user equipment (UE), article, composition, formulation, or process that includes elements other than those listed after such terms in the claims is still considered to fall within the scope of the claims. Additionally, in the appended claims, the terms “first,” “second,” and “third,” etc., are used merely as labels and are not intended to impose numerical requirements on their objects.

Claims

1. A method comprising: A request is received from a computing system including memory and one or more processors to access files stored in a master data repository, wherein the computing system is maintained, controlled or supervised by at least one of a first entity, and the master data repository is maintained, controlled or supervised by at least one of a second entity including a cloud storage service provider. The computing system accesses metadata corresponding to the file using a local file system, wherein the metadata indicates objects in the master data repository, the objects including data corresponding to the file; The computing system enables one or more requests, based on the metadata and sent by the local file system, to the master data repository to retrieve the file; The computing system enables the file to be stored in a temporary data repository of the first entity, and the temporary data repository is accessible by the local file system; The computing system causes one or more computational operations to be performed on the file to produce a modified version of the file stored in the temporary data repository; The computing system causes the modified version of the file to be stored in the object of the master data repository via the local file system; The computing system causes the modified version of the file to be removed from the temporary data repository; and The computing system enables the updated metadata of the modified version of the file to be stored in the additional storage of the first entity, accessible by the local file system.

2. The method according to claim 1, wherein: The computing system receives a request to access the file from a user's computing device, wherein the computing device corresponds to the first entity and has access to the local file system; and The computing system includes a data processing controller that runs on one or more servers of the first entity.

3. The method according to claim 2, wherein, The local file system and the computing device are associated with a location of the first entity, the location of which includes at least one server computer executing an instance of the local file system and includes the temporary data repository accessible by the local file system.

4. The method according to claim 2, wherein: In response to receiving a request to access the file, the computing system determines at least one of one or more commands or one or more application programming interface calls of the local file system to cause the local file system to retrieve the metadata from the attached storage of the first entity and retrieve the file from the master data repository based on the metadata; and The computing system sends at least one of the one or more commands or the one or more application programming interface calls to the local file system.

5. The method according to claim 4, wherein, The computing device displays one or more user interfaces including multiple user interface elements, the multiple user interface elements including a user interface element corresponding to the file, and the user interface elements can be selected such that at least one of the one or more commands or the one or more application programming interface calls is sent to the instance of the local file system executed by the at least one server computer.

6. The method according to claim 2, wherein, In response to determining that no further modifications will be made to the data corresponding to the file, the modified version of the file is stored in the master data repository.

7. The method of claim 6, comprising: In response to one or more additional requests received from the computing device, a session for an instance of the application is initiated; Specifically, based on input provided during the session, the file is modified according to one or more operations to produce the modified version of the file; and The one or more operations correspond to at least one of modifying existing data in the file or adding additional data to the file.

8. The method according to claim 7, wherein, An instance of the application is executed by at least one of the computing device, the one or more servers, or the one or more cloud computing services.

9. The method according to claim 6, wherein, Determining that no further modifications will be made to the data corresponding to the file includes determining that the session of the application instance has been terminated.

10. The method according to claim 1, wherein, Removing the modified version of the file from the temporary data repository is part of a process of dehydrating the data associated with the file in the temporary data repository. The process includes storing updated metadata of the file stored by the local file system and includes generating a stub file during the dehydration process, wherein the stub file indicates at least one of the file's identifier or the file's storage location in the main data repository.

11. The method according to claim 1, wherein: The data included in the document includes scientific data, which includes at least one of the following: genomic information, genetic information, metabolomics information, transcriptomics information, fragmentome information, immune receptor information, methylation information, epigenome information, or proteome information. and At least a portion of the one or more computational operations used to generate a modified version of the file are performed by the bioinformatics pipeline of the first entity.

12. A computing system, comprising: One or more hardware processors; and A memory storing computer-readable instructions that, when executed by the one or more hardware processors, cause the one or more hardware processors to perform operations, the operations including: A request to access files stored in a master data repository is received, wherein the computing system is maintained, controlled, or supervised by at least one of a first entity, and the master data repository is maintained, controlled, or supervised by at least one of a second entity, including a cloud storage service provider. The local file system is used to access the metadata corresponding to the file, wherein the metadata indicates objects in the master data repository, the objects including data corresponding to the file; This enables one or more requests to be sent from the local file system to the master data repository based on the metadata in order to retrieve the file; This allows the file to be stored in a temporary data repository of the first entity, and the temporary data repository to be accessed by the local file system; This causes one or more computational operations to be performed on the file to produce a modified version of the file stored in the temporary data repository; This causes the modified version of the file to be stored in the object of the master data repository via the local file system; This causes the modified version of the file to be removed from the temporary data repository; and This causes the updated metadata of the modified version of the file to be stored in the additional storage of the first entity, which is accessible by the local file system.

13. The computing system according to claim 12, wherein: The request to access the file is received from the user's computing device, which corresponds to the first entity and has access to the local file system; The computing system includes a data processing controller that runs on one or more servers of the first entity; The local file system and the computing device are associated with a location of the first entity, the location of which includes at least one server computer executing an instance of the local file system and includes the temporary data repository accessible by the local file system.

14. The computing system according to claim 13, wherein, The location is one of a plurality of locations of the first entity, each of the plurality of locations corresponding to a corresponding local file system, including at least one corresponding server computer executing the corresponding local file system, and including a corresponding temporary data repository accessible by the corresponding local file system.

15. The computing system according to claim 14, wherein, The memory stores additional computer-readable instructions that, when executed by the one or more hardware processors, cause the one or more hardware processors to perform additional operations, the additional operations including: This causes the updated metadata of the file to be stored in an additional storage location within the plurality of locations; Obtain an additional request from the computing device at the additional location to access the modified version of the file; and This causes one or more additional requests to be sent from the additional file system of the additional location to the master data repository based on the updated metadata, in order to retrieve the modified version of the file.

16. The computing system according to claim 15, wherein, The modified version of the file is stored in the same object as the initial version of the file, and the updated metadata includes the storage location of the master data repository corresponding to the object.

17. The computing system according to claim 15, wherein, The memory stores additional computer-readable instructions that, when executed by the one or more hardware processors, cause the one or more hardware processors to perform additional operations, the additional operations including: This causes the modified version of the file to be stored in an additional temporary data repository in the additional location accessible by the additional file system; Such that additional modified versions of the file are stored by an additional local file system in the same object as the initial version of the file in the master data repository; and This allows additional updated metadata for the file to be stored in the corresponding storage devices at the plurality of locations.

18. The computing system according to claim 12, wherein, The memory stores additional computer-readable instructions that, when executed by the one or more hardware processors, cause the one or more hardware processors to perform additional operations, the additional operations including: This causes the additional modified version of the file to be removed from the additional temporary data repository; and Generate additional updated metadata for the additional modified version of the file, while simultaneously removing the additional modified version of the file from the additional temporary data repository.

19. One or more computer-readable storage media storing computer-readable instructions, said computer-readable instructions causing said one or more hardware processors to perform operations when executed by said hardware processors, said operations including: A request to access files stored in a master data repository is received, wherein the computing system is maintained, controlled, or supervised by at least one of a first entity, and the master data repository is maintained, controlled, or supervised by at least one of a second entity, including a cloud storage service provider. The local file system is used to access the metadata corresponding to the file, wherein the metadata indicates objects in the master data repository, the objects including data corresponding to the file; This enables one or more requests to be sent from the local file system to the master data repository based on the metadata in order to retrieve the file; This allows the file to be stored in a temporary data repository of the first entity, and the temporary data repository to be accessed by the local file system; This causes one or more computational operations to be performed on the file to produce a modified version of the file stored in the temporary data repository; This causes the modified version of the file to be stored in the object of the master data repository via the local file system; This causes the modified version of the file to be removed from the temporary data repository; and This causes the updated metadata of the modified version of the file to be stored in the additional storage of the first entity, which is accessible by the local file system.

20. The computer-readable storage medium of claim 19, wherein, The file and its modified versions are stored in an object in the master data repository. The object has a storage location identifier that allows multiple file systems to access data of one or more versions of the file stored in the object. The object is the true source of the data of the one or more versions of the file.