Data pipeline

A scalable cloud-based data processing platform with a data pipeline and distributed file system enhances data transfer and analysis, addressing inefficiencies in existing cloud computing services to expedite the generation of 3D models from cryo-electron microscopy data.

JP2026034448AInactive Publication Date: 2026-02-27REGENERON PHARMACEUTICALS INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025194069
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-08-27
Filing Date
2025-11-13
Publication Date
2026-02-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Current cloud computing services are inefficient for processing large volumes of data generated by research experiments like cryo-electron microscopy, leading to excessive upload times and slowing down the generation of 3D models, which is critical for scientific research.

Method used

A scalable cloud-based data processing and computing platform that includes a data pipeline with a data transfer filter, a distributed file system, and a data analysis application program to optimize data transfer and processing.

Benefits of technology

Significantly reduces data transfer times, allowing for faster generation of 3D models, thereby accelerating scientific research and drug discovery processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026034448000001_ABST
    Figure 2026034448000001_ABST
Patent Text Reader

Abstract

To provide a method for estimating a three dimensional molecular structure of a particle through reconstruction based on a two dimensional particle image of a cryo-electron microscope.SOLUTION: Determining a program template by iteratively processing the plurality of job parameter configurations using at least one optimization solver based on the one or more job parameters and the perturbatively generated plurality of job parameter configurations, wherein the program template defines a level of detail of extraction from a two-dimensional particle image for generating a three-dimensional particle representation at a specified resolution, and causing execution of a data analysis application program on the two-dimensional particle image by a plurality of compute nodes to extract at the level of detail from the two-dimensional particle image to generate the three-dimensional particle representation at the desired resolution.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED PATENT APPLICATIONS: This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 163,690, filed March 19, 2021, and U.S. Provisional Patent Application No. 63 / 237,904, filed August 27, 2021. The entire contents of the foregoing applications are incorporated herein by reference. [Background technology]

[0002] Currently, cloud computing services are provided globally to millions of users and customers residing in different locations (e.g., countries, continents, etc.). Various entities provide private or public cloud computing services globally to different customers across various sectors for critical and non-critical applications. These entities offer various cloud computing services, including, for example, software as a service (SaaS), infrastructure as a service (IaaS), and / or platform as a service (PaaS). To use such cloud computing services, users must transfer locally generated data to the cloud. However, research experiments generate surprisingly large amounts of data, and uploading such data to cloud computing services takes an unsatisfactory amount of time, slowing down research. For example, cryo-electron microscopy (Cryo-EM) reveals the structure of proteins by probing flash-frozen solutions with a beam of electrons and then combining two-dimensional (2D) images of individual molecules into three-dimensional (3D) images. Cryo-EM is a powerful scientific instrument that generates a vast amount of data in the form of high-resolution 2D photographs of proteins. Scientists must perform a series of computationally intensive calculations to convert 2D images into useful 3D models. These tasks can take weeks on a typical workstation or computer cluster with finite capacity. Excessive upload times associated with using cloud computing services for model generation further exacerbate the time required to generate 3D models, slowing down research. Summary of the Invention [Problem to be solved by the invention]

[0003] The disclosed methods and systems, individually or in combination, provide a scalable cloud-based data processing and computing platform that supports high-volume data pipelines. [Means for solving the problem]

[0004] In one embodiment, the present disclosure provides a method. The method includes receiving an indication of a synchronization request. The method further includes determining one or more files stored at a staging location based on the indication. The method further includes generating a data transfer filter based on the one or more files. The method further includes causing a transfer of the one or more files to a destination computing device based on the data transfer filter.

[0005] In one embodiment, the present disclosure provides a method. The method includes receiving, via a graphical user interface, a request to migrate a dataset from object storage to a distributed file system. The method further includes receiving, via the graphical user interface, an indication of a storage size of the distributed file system. The method further includes, based on the request and the indication, migrating the dataset from the object storage to the distributed file system associated with the storage size.

[0006] In one embodiment, the present disclosure provides a method. The method includes identifying a data analysis application program. The method further includes identifying a dataset associated with the data analysis application program. The method further includes determining, as a program template, one or more job parameters associated with the data analysis application program for processing the dataset. The method further includes causing execution of the data analysis application program on the dataset based on the program template.

[0007] In one embodiment, the present disclosure provides a method. The method includes receiving an indication of a synchronization request. The method further includes determining one or more files stored at a staging location based on the indication. The method further includes generating a data transfer filter based on the one or more files. The method further includes triggering a transfer of the one or more files to object storage on a destination computing device based on the data transfer filter. The method further includes receiving a request to migrate one or more files from the object storage to a distributed file system via a graphical user interface. The method further includes receiving an indication of a storage size of the distributed file system via the graphical user interface. The method further includes migrate the one or more files from the object storage to the distributed file system associated with the storage size based on the request and the indication. The method further includes identifying a data analysis application program associated with one or more files in the distributed file system. The method further includes determining one or more job parameters associated with the data analysis application program that processes the dataset as a program template. The method further includes triggering execution of the data analysis application program associated with the one or more files in the distributed file system based on the program template.

[0008] Additional advantages of the disclosed methods and compositions will be set forth in part in the description which follows, and in part will be understood from the description or may be learned by practice of the disclosed methods and compositions. The advantages of the disclosed methods and compositions will be realized and attained by means of the elements and combinations particularly pointed out in the appended claims. It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention as claimed.

[0009] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate several embodiments of the disclosed methods and compositions and, together with the description, serve to explain the principles of the disclosed methods and systems. [Brief explanation of the drawings]

[0010] [Figure 1] 1 illustrates an example of an exemplary operating environment. [Figure 2A] 1 illustrates an exemplary data pipeline. [Figure 2B] 1 illustrates an exemplary operating environment. [Figure 3] 1 illustrates an exemplary operating environment. [Figure 4A] 1 illustrates an exemplary operating environment. [Figure 4B] 1 illustrates an exemplary cloud-based storage system. [Figure 5] 1 illustrates an exemplary graphical user interface. [Figure 6A] 1 illustrates an exemplary graphical user interface. [Figure 6B] 1 illustrates an exemplary graphical user interface. [Figure 7] 1 illustrates an exemplary graphical user interface. [Figure 8A] 1 shows an exemplary program template. [Figure 8B] 1 illustrates an exemplary operating environment. [Figure 8C] 1 illustrates an exemplary operating environment. [Figure 9] 1 illustrates an exemplary operating environment. [Figure 10] 1 illustrates an exemplary operating environment. [Figure 11] 1 illustrates an exemplary operating environment. [Figure 12] An exemplary method is shown. [Figure 13] An exemplary method is shown. [Figure 14] An exemplary method is shown. [Figure 15] An exemplary method is shown. DETAILED DESCRIPTION OF THE INVENTION

[0011] The disclosed methods and systems may be more readily understood by reference to the following detailed description of specific embodiments and examples contained therein, as well as the figures and their preceding and following descriptions.

[0012] It is understood that the methods and systems of the present disclosure are not limited to the particular methodology, protocols, and reactants described, as these may vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to limit the scope of the present invention, which is limited only by the appended claims.

[0013] It should be noted that as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Thus, for example, reference to "an image" includes a plurality of images, etc.

[0014] "Optional" or "optionally" means that the subsequently described event, circumstance, or material may or may not occur or exist, and includes instances in which the event, circumstance, or material occurs or exists, as well as instances in which it does not occur or exist.

[0015] Throughout the description and claims, the words "include" and variations of words such as "comprises" and "comprising" are intended to mean, for example, but not limited to, other adjuncts, components, wholes, or steps. In particular, in methods described as including one or more steps or operations, each step is intended to include what is recited (unless that step includes a limiting term such as "consisting of"), and each step is not intended to exclude, for example, other adjuncts, components, wholes, or steps not recited in the step.

[0016] "Exemplary" means "an example of" and is not intended to convey an indication of a preferred or ideal configuration. "Such" is used for descriptive purposes, not in a limiting sense.

[0017] Ranges may be expressed herein as from "about" one particular value and / or to "about" another particular value. When such ranges are expressed, it is the range from one particular value and / or to the other particular value that is specifically intended and considered to be disclosed, unless the context specifically dictates otherwise. Similarly, when values ​​are expressed as approximations, the use of "about" will be understood to indicate that the particular value forms another specifically contemplated embodiment that should be considered to be disclosed, unless the context specifically dictates otherwise. It will be further understood that the endpoints of each range are significant both relative to the other endpoint, and independently of the other endpoint, unless the context specifically dictates otherwise. Finally, it will be understood that all individual values ​​and subranges of values ​​falling within an explicitly disclosed range are also specifically intended and should be considered to be disclosed, unless the context specifically dictates otherwise. The foregoing applies regardless of whether, in a particular instance, some or all of these embodiments are explicitly disclosed.

[0018] The disclosed technology can be used on a wide range of macromolecules, including but not limited to proteins, peptides, nucleic acids, and polymers. In some aspects, the protein may be an antibody or a fragment thereof.

[0019] Once the structure of a macromolecule, such as an antibody, is determined using the disclosed techniques, the macromolecule can be used in methods of treatment, detection, or diagnosis. For example, an antibody identified using the disclosed techniques can be administered to a subject to treat the subject's disease, disorder, and / or condition. In some aspects, the subject's disease, disorder, and / or condition can be cancer, a viral infection (e.g., coronavirus, influenza virus), or an inflammatory disorder (e.g., rheumatoid arthritis, lupus).

[0020] Various techniques can be used to attempt to determine the 3D structure of proteins, viruses, and other molecules from their images. Once an image of a target molecule is collected, determining the molecule's 3D structure requires the successful completion of a difficult reconstruction of the 3D structure from the image. The computational demands of processing the volume of 2D images and the subsequent generation of the 3D structure are extreme, easily resulting in the generation of terabytes of data per day.

[0021] To fully realize the potential of cryo-EM technology, researchers need both data and computing power available over a network in a timely manner. Researchers need an information technology (IT) infrastructure that can handle data growth, as well as the massive demand for computing in the cloud, using highly specialized packages for post-capture image processing. Processing data / images from cryo-EM microscopes requires scalable storage for large datasets (e.g., a medium size of 1.2 TB per sample), the fastest CPU / GPU-enabled hybrid environment to perform calculations related to large datasets, and a high-speed network to move data from the instrument to the cloud. To address storage, computing power, and operating costs, the present disclosure provides a high-performance computing (HPC) platform on the cloud. The disclosed method and system can provide final results (e.g., 3D models) in significantly less time than systems in the art.

[0022] Given the growth and pace of data, various considerations can be taken into account and components can be reconfigured to provide the best infrastructure to accommodate cryo-EM needs. Increasing data transfer rates provides elasticity to accommodate growing needs and minimize irresolvable bottlenecks, ensuring data is available for processing in the shortest possible time (e.g., less than one to two hours for each dataset). Storage optimization (e.g., implemented using FSx Lustre (Tier 1) and S3 (Tier 2)) allows for unlimited storage capacity while keeping operational costs low, providing long-term storage for raw data, and providing a highly parallel file system to sustain heavy I / O (or throughput). Computing optimization can be adjusted based on an evaluation of workload usage patterns. Software optimization allows job-based submission script optimization to be performed across resources, completing jobs more quickly with less computing power. Self-service storage management utilities allow users to manage the analysis of datasets. For example, instead of creating one larger file system, a distributed file system (DFS) and / or parallel file systems can be created for each dataset being analyzed. Such a distributed / parallel file system may reduce storage capacity and / or costs.

[0023] Efficient, high-speed big data transfer methods are disclosed that can support one or more data processing applications. For example, data processing applications that can be supported by the efficient, high-speed big data transfer techniques disclosed herein include 3D structure estimation from 2D electron cryomicroscopy images.

[0024] As shown in FIG. 1 , system 100 may include data source 102. Data source 102 may be any type of data-generating system, such as, for example, an imaging system, a gene sequencing system, or a combination thereof. Data source 102, in one embodiment, may include one or more components that provide data. The components may expose data in numerous ways according to one or more mechanisms. For example, a component may be embodied in or configured on a computing device that includes one or more types of data storage. Thus, data source 102 may include a network file system (NFS), a server message block (SMB), a Hadoop distributed file system (HDFS), and / or an on-premises object store.

[0025] In one embodiment, the data source 102 may comprise an imaging system consisting of one or more electron microscopes (e.g., cryo-electron microscopes (cryo-EM)). Cryo-EM is a computer vision-based approach to determining 3D macromolecular structure. Cryo-EM is applicable to medium- to large-sized molecules in their native state. This scope of application contrasts sharply with X-ray crystallography, which requires crystals of the target molecule, where crystals are often difficult (though certainly not infeasible) to grow. Such scope also contrasts sharply with nuclear magnetic resonance (NMR) spectroscopy, which is limited to relatively small molecules. Cryo-EM has the potential to shed light on the molecular and chemical nature of fundamental biology through the discovery of the atomic structures of previously unknown biological structures. Many of these atomic structures have proven difficult or impossible to study using traditional structural biology techniques.

[0026] In cryo-EM, a purified solution of a target molecule is first frozen into a thin (single-molecule thick) film on a carbon grid, and the resulting grid is then imaged in a transmission electron microscope. The grid is exposed to a low-dose electron beam in a microscope column, and 2D projections of the sample are collected at the base of the column using a camera (film, charge-coupled device (CCD) sensor, direct electron detector, or similar). A large number of such projections are obtained, each of which provides a micrograph containing hundreds of visible individual molecules. In a process known as particle harvesting, individual molecules are selected from the micrograph, resulting in a stack of cropped images of the molecule (called particle images). Each particle image provides a noisy view of the molecule in an unknown pose. Once a large set of 2D electron microscope particle images of a molecule is obtained, reconstructions can be performed to estimate the 3D density of the target molecule from the images.

[0027] In cryo-EM, millions of 2D particle images of a sample, often consisting of hundreds to thousands of copies of a single protein molecule or protein-drug complex (known as a "target"), are captured in an electron microscope. The particle images can then be computationally assembled to reconstruct and refine a 3D model of the target to a desired resolution. In particular, the number of images, level of detail, and level of noise in each image significantly exceed what a person can reasonably comprehend by inspecting the image, mentally, or by attempting to interpret the image. In other words, the richness and complexity of the imaging data contained in a particle image easily overwhelms the human mind to reconstruct and refine a 3D model of the target. In most cases, all or most (potentially on the order of millions) of the images, but not all of the corresponding viewing directions, are simultaneously aggregated in the abstract form of complex-valued coefficients of a Fourier series expansion arranged in a 3D grid, making the information contained in the image interpretable by the embodiments described herein. The embodiments described herein generally use approaches and symbols that are substantially impossible to implement without the use of a computing device, and therefore, the approaches, methods, and processes described herein can only be realistically and specifically implemented through the judicious use of a computing device as described herein.

[0028] Generally, the usefulness of a particular cryo-EM reconstruction for a given target depends on the resolution achievable for that target. High-resolution reconstructions, when particularly good, can resolve fine details, including atomic positions that can be interpreted from the reconstruction. In contrast, low-resolution reconstructions may only represent the large globular features of a protein molecule rather than the fine details, thus making it difficult to use the reconstruction in further chemical or biological research pipelines.

[0029] Especially for drug development pipelines, high-resolution reconstruction of a target can be substantially advantageous. For example, such high-resolution reconstruction can provide invaluable insight into whether a target is suitable for the application of a therapeutic agent (e.g., a drug). For another example, high-resolution reconstruction can be used to understand the types of drug candidates that may be suitable for the target. For another example, if the target is actually a protein and a specific drug candidate compound, high-resolution reconstruction can even reveal possible ways to optimize the drug candidate to improve its binding affinity and reduce off-target binding, thereby reducing the possibility of unwanted side effects. Therefore, in cryo-EM reconstruction, approaches that can improve the resolution of computationally reconstructed 3D results have high scientific and commercial value.

[0030] Resolution in the context of cryo-EM is generally measured and described in terms of the shortest resolvable wavelength of the 3D structure signal in the final 3D structure output of a structure refinement technique. In some cases, the shortest resolvable wavelength has a resolution that is the shortest wavelength with an accurate and verifiable signal. Wavelength is typically expressed in units of angstroms (Å, one-tenth of a nanometer). Lower wavelength values ​​indicate higher resolution.

[0031] As an example, an ultra-high resolution cryo-EM structure may have a resolution of about 2 Å, a medium resolution may have a resolution of about 4 Å, and a low resolution may be in the range of about 8 Å or worse. In addition to the numerical resolution, interpretability, and usefulness of a cryo-EM reconstruction, it may depend on the quality of the reconstructed 3D density map and whether a qualified user can examine the 3D density map with the naked eye and identify important features of the protein molecule, such as the backbone, side chains, bound ligands, etc. The user's ability to accurately identify these features depends heavily on the resolution quality of the 3D density map.

[0032] Thus, data source 102 may be configured to generate data 104. Data 104 may include image data, such as image data defining 2D electron cryomicroscopy images, also referred to as particle images. Data 104 may, in some cases, include sequence data.

[0033] The computing device 106 may be in communication with the data source 102. The computing device 106 may be, for example, a smartphone, a tablet, a laptop computer, a desktop computer, a server computer, etc. The computing device 106 may include one or more server devices. The computing device 106 may be configured to generate, store, maintain, and / or update various data structures, including databases for storing the data 104. The computing device 106 may be configured to operate one or more application programs, such as a data staging module 108, a data synchronization manager 110, and / or a data synchronization module 112. The data staging module 108, the data synchronization manager 110, and / or the data synchronization module 112 may be stored and / or configured to operate separately on the same computing device 106 or on separate computing devices.

[0034] In one embodiment, the computing device 106 may be configured to collect, acquire, and / or receive data 104 from the data source 102 for storage on the computing device 106 (or in a storage system operatively coupled to the computing device 106) via a data staging module 108. The storage system may comprise one or more memory devices and may be referred to as a staging location. The data staging module 108 may manage data stored in the storage system until such data is transferred from the staging location. Once the data is transferred from the staging location, the data staging module 108 may delete such data. The data staging module 108 may be configured to receive the data 104 through various mechanisms. In one embodiment, the staging location may be treated as a remote directory for the data source 102, such that the data 104 generated by the data source 102 is saved directly to the staging location. Additionally, or in another embodiment, the data staging module 108 may be configured to monitor one or more network storage locations to detect new data 104 upon identifying new data 104 at the network storage locations, and the data staging module 108 may transfer the new data 104 to the staging locations. Additionally, or in yet another embodiment, the data staging module 108 may be configured to allow a user to manually upload data to the staging locations.

[0035] The computing device 106 may be configured to transfer the data 104 from the staging location to the cloud platform 114 via the data synchronization manager 110 and the data synchronization module 112. In one embodiment, the computing device 106 may be configured to transfer the data 104 as it is received from the data source 102 via the data synchronization manager 110 and the data synchronization module 112. As disclosed, the system 100 represents an automated end-to-end processing pipeline capable of transmitting and processing over 1 TB / hour of raw data. In one embodiment, the data 104 may be transferred in near real time as the data 104 is acquired.

[0036] In one embodiment, the data synchronization module 112 may be a data synchronization application program configured to transmit the data 104 to the cloud platform 114. The data synchronization application program may be any data synchronization program, including, for example, AWS DataSync. AWS DataSync is a native AWS service configured to transmit large amounts of data between on-premises storage and Amazon's native storage services. In one example, the on-premises storage may reside within the computing device 106 or may be a staging location functionally coupled to it. However, because data synchronization application programs are "synchronization" utilities, such application programs do not function as one-way copy utilities. In the case of AWS DataSync, AWS DataSync transfers data through four phases: startup, preparation, transfer, and verification. In particular, during the preparation phase, AWS DataSync examines the source (e.g., computing device 106) and destination (e.g., cloud platform 114) file systems to determine which files to synchronize. AWS DataSync recursively scans the content and metadata of files on the source and destination file systems for differences. The time AWS DataSync spends on the preparation stage depends on the number of files in both the source and destination file systems, and large data transfers can take several hours. As the size of the data 104 stored in the source and / or destination increases, AWS DataSync The time DataSync spends in the preparation phase also increases. Currently, for an example of a 500TB data size on the destination (e.g., cloud platform 114), the preparation phase takes over two hours. Only after the scan is complete and differences are determined does AWS DataSync move to the transfer phase, transferring files and metadata from the source file system to the destination by copying changes to files with different content or metadata between the source file system and the destination.

[0037] As described herein, data sources 102 generate very large amounts of data 104. This very large amount of data needs to be made available as quickly as possible on a high-performance computing platform, such as a cloud platform 114. Making the data 104 available more quickly provides lead time for scientists to process and achieve results more quickly, directly impacting the timing of drug discovery. The current state of existing data synchronization application programs significantly increases the time required to transfer such data to a high-performance computing platform due to the time spent scanning local and remote file systems before transferring the data.

[0038] The system 100 is configured to implement an improved data pipeline 201 that addresses technical deficiencies of data synchronization application programs, as shown in FIG. 2 . The data pipeline 201 may include a multi-stage data transfer process for pushing data 104 from a staging location (e.g., on-premise) on a computing device 106 to a cloud system 114. As described herein, the data 104 may be generated by a data source 102. As part of the data staging process 202, the data 104 may be stored at the staging location by a data staging module 108. The purpose of the data staging process 202 is to preserve the data 104 and keep it ready for transmission. The data 104 at the staging location may be deleted once the data 104 is moved to the data destination (e.g., cloud platform 114). The synchronization conditions 203 dictate when the data transfer process 204 can be initiated. Thus, satisfying the synchronization conditions 203 may trigger initiation of the data transfer process 204. In one exemplary scenario, the data transfer process 204 is initiated periodically at a rate defined by a time interval, which may be configurable. Thus, the synchronization condition 203 dictates that the elapsed time since the last data transfer must equal the time interval. However, prior to execution of the data transfer process 204, the data synchronization manager 110 may be configured to determine the data 104 (e.g., identifying files and / or directories) currently available at the staging location. To that end, for example, the data synchronization manager 110 may fetch a list 205 of the data 104 currently available at the staging location. In one embodiment, the data synchronization manager 110 may connect to the staging location and / or any respective mount points / disk volumes. The data synchronization manager 110 may then execute a list command to fetch a list of available files. The data synchronization manager 110 may be configured to utilize a naming convention when fetching the list of available files.For example, a scientific instrument may be configured to generate data with a defined naming convention. The data synchronization manager 110 may utilize regular expressions (RegEx) to include (or exclude) one or more files in a list. In one embodiment, the data synchronization manager 110 may also rely on RegEx to include directories and / or files in a list.

[0039] The data synchronization manager 110 may be configured to generate the filter 206 using the list. The filter may include one or more of a file name, a file location, a file extension, a file size, a checksum, a creation date, a modification date, a combination thereof, and the like. Generating the filter 206 may include generating a message that invokes a function call to a cloud service (e.g., AWS DataSync), where the message passes the list of available files as arguments to the function call. The function call may initiate a task (or job) in the cloud service. The function call may be retrieved according to an API implemented by the data storage service. The cloud service may be provided by one or more components of the cloud platform 114. The filter may be dynamically generated in that the filter may be generated at each iteration of the data transfer process 204. In one embodiment, the filter may include references to partial files (e.g., files that are not yet complete or are in the process of being transferred to a staging location). If the filter includes a partial file, the partial file is transferred, and in subsequent iterations, the filter includes the complete file and updates the transferred partial file.

[0040] The data synchronization manager 110 then triggers the data transfer process 204 according to the filter 206. The filter 206 causes the data transfer process 204 to transfer only the files and / or directories specified by the filter 206. Thus, the filter 206 represents data 104 that exists only in the staging location. The data pipeline 201 represents an improvement in computer technology because a standard data transfer process compares the data available in the staging location with the data available in the cloud platform 114, determines all new and changed / updated files to transfer, and pushes the data to the cloud platform 114, resulting in a significant increase in the time to complete the data transfer process. In contrast, the dynamically generated filter of the present invention causes the data transfer process 204 to scan only a limited set of data in the staging location and the cloud platform 114, which significantly reduces the time required to complete the data transfer process 204. In an AWS DataSync embodiment, the filter 206 causes the preparation phase of the AWS DataSync task to scan only the files specified in the filter, rather than all files, minimizing the preparation phase time.

[0041] In some embodiments, various synchronization policies can be generated and / or applied to determine which data will and will not be synchronized. A synchronization policy may specify which files will be synchronized based on selected criteria, including data type, metadata, and location information (e.g., the electron microscope instrument generating the data). As shown in FIG. 2B , synchronization policies can be maintained in one or more memory devices 250 (referred to as data stores 250) in one or more data structures 260 (referred to as policies 260). The data stores 250 may be integrated with or operatively coupled to the computing device 106. In some cases, the data stores 250 may be part of a staging location. A synchronization policy can dictate how a filter 206 is generated. In one exemplary scenario, a scientist may flag certain data to prevent it from being synchronized, even though the data resides in a staging location. A synchronization policy may dictate that data flagged in this manner not be synchronized. As a result, the data synchronization manager 110 can be configured to use one or more lists of files and such synchronization policies to generate instances of filters 206. Therefore, that instance of the filter can be updated to include one or more flags (which may be called exclude flags) associated with each file. Exclude flags cause such files to be excluded from synchronization. Another synchronization policy can dictate the time-to-live period of the exclude flag, which defines the time interval during which the exclude flag is active. The TTL period synchronizes the data at a point in time and avoids unnecessarily holding data in a staging location.

[0042] Other types of flags or metadata may be defined to control how filters 206 are instantiated and applied to data synchronization. For example, some flags may automatically expire after the entire dataset is loaded into a staging location, preventing partial synchronization.

[0043] Figure 3 shows an example AWS architecture for implementing the data pipeline 201 of Figure 2. Data is generated in a data center / lab at 301. The generated data may be staged in NetApp storage located in a local data center at 302. An AWS Cloud viewing rule is configured to trigger a Lambda function at standard intervals (e.g., periodically, at a configurable rate, or time interval) according to the agreed-upon SLA at 303. The invoked Lambda function connects to the on-premises NetApp storage via NFS to fetch a list of available files at 304. Once the file list is available, the Lambda function filters out valid datasets (based on naming conventions) and passes them as a filter to the triggered DataSYNC job at 305. A Lambda environment variable may hold the DataSYNC job ID that needs to be triggered. The completion of Lambda execution (success / failure) is passed to an SNS topic at 306. The Lambda environment variable may hold the SNS topic ARN. Any success or failure messages may be sent to a subscribed email address at 307. The SNS subscription has a message attribute filter setting that, if picked up and fails, additionally sends a text to the administrator at 308. If a failure occurs, the administrator is notified immediately by text, allowing for faster response. The example AWS architecture of FIG. 3 significantly reduces the preparation phase timing, as shown in Table 1. Making data available for computation as quickly as possible is a key factor for faster drug discovery and analysis. With the improved data pipeline provided by embodiments of the present disclosure, data is available for computation significantly faster. In some cases, a speedup factor of approximately 4 can be achieved. [Table 1]

[0044] The data 104 received by the cloud platform 114 may be stored in one or more types of storage (e.g., a file system). The cloud platform 114 may include a distributed parallel file system (e.g., Lustre) and / or an object-based file system. In one embodiment, the data 104 received by the cloud platform 114 may be stored in a distributed parallel file system or an object-based file system. In one embodiment, the data 104 received by the cloud platform 114 is initially stored in an object-based file system and is moved to a distributed parallel file system as the data 104 is processed (e.g., analyzed).

[0045] A file system is a subsystem that an operating system or program uses to organize and keep track of files. File systems may be organized in different ways. For example, a hierarchical file system is a file system that uses directories to organize files into a tree structure. File systems provide the ability to search for one or more files stored within the file system. This is often done using a "directory" scan or search. In some operating systems, searches may include file version, file name, and / or file extension.

[0046] Operating systems provide their own file management systems, but third-party file systems may also be developed. These systems may interact smoothly with the operating system but offer more features, such as encryption, compression, file versioning, improved backup procedures, and enhanced file protection. Some file systems are implemented over a network. Two common systems include the Network File System (NFS) and the Server Message Block (SMB, now CFIS) system. A file system implemented over a network takes requests from the operating system, converts the requests into network packets, transmits the packets to a remote server, and then processes the responses. Other file systems are implemented as downloadable file systems, where the file system is packaged and delivered to the user as a unit.

[0047] File systems share an abstracted interface through which users may perform operations. These operations include, but are not limited to: mount / unmount, directory scan, open (create) / close, read / write, status, etc. The step of associating a file system with an operating system (e.g., tying the file system's virtual layer to the operating system) is collectively called "mounting." In common usage, a newly mounted file system is associated with a specific location in the hierarchical file tree. All requests for that part of the file tree are passed to the mounted file system. Different operating systems impose limits on the number of file system mounts and the depth of nested file system mounts. Unmounting is the opposite of mounting; the file system is detached from the operating system.

[0048] As shown in FIG. 4 , in one embodiment, analysis of the data 104 may be performed on a distributed computing and storage architecture, such as a cloud platform 114. Because the data sources 102 typically generate a significant amount of data 104 (e.g., data per experiment), it is not feasible to retain such data for long periods in hot storage 401 (e.g., a distributed parallel file system, a solid-state drive (SSD), etc.). Therefore, the data 104 may be maintained in warm storage 402 (e.g., object storage) instead of hot storage 401. When data processing is required, the data 104 can be moved to hot storage 401 via a self-service model using a dataset management (DSM) utility 116 disclosed herein. The DSM utility 116 may allow or otherwise facilitate users' creation of a POSIX distributed file system and retrieval of appropriate datasets from warm storage 402 to hot storage 401. The POSIX file system may be attached within an HPC cluster (e.g., computing nodes 403) for processing. As an example, Lustre is a high-performance distributed file system that can act as a front end to S3 data and present the S3 data in a POSIX-based file system to the compute nodes 403. However, such file systems are financially expensive. To minimize storage costs, the disclosed DSM utility 116 provides, for example, on-demand provisioning of cloud-based file systems. The Lustre file system is one example of a cloud-based file system that can be provided. A user can create a Lustre file system that points to a dataset when a job runs. The Lustre file system serves as staging storage for processing, and once the job is complete, the results are synchronized to S3 object storage, and the Lustre file system can be deleted using the DSM utility 116.

[0049] The DSM utility 116 can create a new custom-sized distributed file system by targeting the dataset to be processed. The DSM utility 116 can mount the distributed file system on the HPC cluster (e.g., compute nodes 403) to stage the processed data. The DSM utility 116 can synchronize the changed dataset to an S3 object store. The DSM utility 116 can enable browsing of files available in the S3 object store. The DSM utility 116 can enable self-service data lifecycle management. Typically, such functionality requires the support of a technically trained user, but the DSM utility 116 allows non-technical users to perform these tasks.

[0050] FIG. 5 illustrates a graphical user interface 501 for the DSM utility 116. The graphical user interface 501 provides users with the ability to create and manage file systems for distributed workloads. As shown in FIG. 5, the graphical user interface 501 provides a menu of selectable options, including a first selectable option 502 (labeled "Create Lustre") and a second selectable option 503 (labeled "Manage Lustre"). The first selectable option 502 allows users to browse data stores on S3, view files and directories, and create a file system (e.g., a Lustre file system) from any location on S3. The second selectable option 503 allows users to mount a Lustre file system once it has been created, view the file system from the operating system (O / S) level, and access data within the file system. The second selectable option 503 also allows users to store data in the Lustre file system once it has been created, similar to S3, and to create new data or modify existing data while manipulating the file system. To make this data persistent even after deleting the file system, the user may export the data to a data store (S3). A second selectable option 503 allows the user to view the status of the export job once the Lustre file system is created. The user can toggle between the export job status and the file system view. A second selectable option 503 allows the user to delete the file system once the Lustre file system is created and the task is complete.

[0051] As shown in FIG. 6A, upon selecting (e.g., clicking) the first selectable option 502 ("Create Lustre"), the graphical user interface 501 presents the contents of the data store, the warm storage. The user can drill down into any directory to view subfolders by double-clicking a particular directory. A visually selectable element 602 (labeled "Previous Directory") allows the user to go back one step at a time. A visually selectable element 603 (labeled "Update Dataset") allows the user to navigate to the top-level screen. The visually selectable element 603 can also serve as an update indicator to retrieve the latest data from the data store. After the user selects a directory to load, selecting the visually selectable element 604 (labeled "Load Dataset") initiates the creation of the Lustre file system. As shown in FIG. 6B, the graphical user interface 501 provides the user with the ability to adjust the size of the Lustre file system. By default, a Lustre file system can be created with 7.2 TB of storage capacity, but to change the storage capacity, move slider mark 610 left (decrease) or right (increase). A menu of selectable options is also shown in Figure 6B, including a first selectable option 605 (labeled "Proceed") and a second selectable option 606 (labeled "Cancel"). Selecting first selectable option 605 creates a Lustre file system.

[0052] As shown in FIG. 7 , upon selecting the second selectable option 503 (“Manage Lustre”), the graphical user interface 501 provides an upper window 710a that displays all file systems owned by the user and a lower window 710b that displays other file systems not owned by the user. The graphical user interface 501 shown in FIG. 7 also includes a menu of selectable options, including a first selectable option 701 (labeled “Mount File System”), a second selectable option 702 (labeled “Store Data in S3”), a third selectable option 703 (labeled “Show Repository Tasks”), and a fourth selectable option 704 (labeled “Delete Lustre FSx”). After the file system creation is complete, the user must mount the file system at the O / S level to access the files. To mount the file system, the user can select the file system to mount and then select (e.g., click) the first selectable option 701 (“Mount File System”). Running a data analysis job may create new files or modify existing files. To make new or changed data persistent, the data is saved to a data store (e.g., S3). For this operation, the user can select the file system to save to and then select (e.g., click on) the second selectable option 702 ("Save Dataset to S3"). To check the status of a repository task (e.g., save dataset to S3), the user can select the file system on which the "Save Dataset S3" operation will be performed and then select (e.g., click on) the third selectable option 703 ("Show Repository Tasks") to display a screen listing the status of the repository job. Once data analysis is complete on the file system, the user can delete the file system to save costs.To delete a file system, a user can select the file system to delete and then select (e.g., click) a fourth selectable option ("Delete Lustre FSx"). A later one of these selections can prompt the user to "Save dataset to S3" 702 before deleting the selected file system. After confirmation, deletion of the file system can begin.

[0053] For further explanation, FIG. 4B illustrates an example of a cloud-based storage system 418 of a cloud platform 114, according to some embodiments of the present disclosure. In one embodiment, the DSM Utility 104 may communicate with the cloud storage system 418, which in one embodiment is embodied in one or more of the components illustrated in FIG. 4B (e.g., a storage controller application, a software daemon, etc.). In the example illustrated in FIG. 4B, the cloud-based storage system 418 is created entirely within a cloud platform 114, such as, for example, Amazon Web Services (AWS')™, Microsoft Azure™, Google Cloud Platform™, IBM Cloud™, Oracle Cloud™, etc. The cloud-based storage system 418 illustrated in FIG. 4B includes two cloud computing instances 420, 422, each of which is used to support the execution of a storage controller application 424, 426. The cloud computing instances 420, 422 may be embodied as instances of cloud computing resources (e.g., virtual machines) provided by the cloud platform 114, for example, that may support the execution of software applications such as the storage controller applications 424, 426. For example, each of the cloud computing instances 420, 422 may run on an Azure VM, and each Azure VM may include high-speed temporary storage that may be utilized as a cache (e.g., a read cache). In one embodiment, the cloud computing instances 420, 422 may be embodied as Amazon Elastic Compute Cloud (“EC2”) instances. In such an example, an Amazon Machine Image (“AMI”) that includes the storage controller applications 424, 426 may be launched to create and configure a virtual machine that may run the storage controller applications 424, 426.

[0054] 4B , storage controller applications 424, 426 may be embodied as modules of computer program instructions that, when executed, perform various storage tasks. For example, storage controller applications 424, 426 may be embodied as modules of computer program instructions that, when executed, perform the same tasks related to writing data to cloud-based storage system 418, erasing data from cloud-based storage system 418, retrieving data from cloud-based storage system 418, monitoring and reporting disk usage and performance, performing redundancy operations, such as RAID or RAID-like data redundancy operations, compressing data, encrypting data, and deduplication of data. Because there are two cloud computing instances 420, 422, each including storage controller applications 424, 426, in some embodiments, one cloud computing instance 420 may act as a primary controller as described above, and the other cloud computing instance 422 may act as a secondary controller as described above. The storage controller applications 424, 426 depicted in FIG. 4B may include the same source code running within different cloud computing instances 420, 422, such as separate EC2 instances.

[0055] Other embodiments do not include primary and secondary controllers and are within the scope of this disclosure. For example, each cloud computing instance 420, 422 may act as a primary controller for some portion of the address space supported by cloud-based storage system 418, or each cloud computing instance 420, 422 may act as a primary controller where servicing of I / O operations directed to cloud-based storage system 418 is divided in some other manner. Indeed, in other embodiments where cost savings can take precedence over performance demands, there may be only a single cloud computing instance that includes the storage controller application.

[0056] The cloud-based storage system 418 depicted in Figure 4B includes cloud computing instances 440A, 440B, and 440n having local storage 430, 434, and 438. The cloud computing instances 440A, 440B, and 440n may be embodied as instances of cloud computing resources that may be provided by the cloud platform 114, for example, to support the execution of software applications. The cloud computing instances 440A, 440B, and 440n of Figure 4B differ from the cloud computing instances 420, 422 described above because the cloud computing instances 440A, 440B, and 440n of Figure 4B have local storage 430, 434, and 438 resources, while the cloud computing instances 420, 422 that support the execution of the storage controller applications 424, 426 need not have local storage resources. Cloud computing instances 440A, 440B, and 440n having local storage 430, 434, and 438 may be embodied, for example, as EC2 M5 instances including one or more SSDs, as EC2 R5 instances including one or more SSDs, or as EC2 I3 instances including one or more SSDs. In some embodiments, local storage 430, 434, and 438 may be embodied as solid-state storage (e.g., SSDs) rather than storage using hard disk drives. Hot storage 401 may include one or more of local storage 430, 434, and 438.

[0057] 4B , each of the cloud computing instances 440A, 440B, and 440n having local storage 430, 434, and 438 may include software daemons 428, 432, and 436 that, when executed by the cloud computing instances 440A, 440B, and 440n, can present themselves to the storage controller applications 424, 426 as if the cloud computing instances 440A, 440B, and 440n were physical storage devices (e.g., one or more SSDs). In such an example, the software daemons 428, 432, and 436 may include computer program instructions similar to those typically included on storage devices, such that the storage controller applications 424, 426 can send and receive the same commands that a storage controller sends to a storage device. In this way, the storage controller applications 424, 426 may include code that is the same (or substantially the same) as the code executed by the controllers in the storage systems described above. In these and similar embodiments, communication between storage controller applications 424, 426 and cloud computing instances 440A, 440B, and 440n with local storage 430, 434, and 438 may utilize iSCSI, NVMe over TCP, messaging, a custom protocol, or some other mechanism.

[0058] In the example depicted in FIG. 4B , each of the cloud computing instances 440A, 440B, and 440n with local storage 430, 434, and 438 may also be coupled to block storage 442, 444, and 446 provided by the cloud platform 114, such as, for example, Amazon Elastic Block Store (“EBS”) volumes. Hot storage 401 may include one or more of the block storage 442, 444, and 446. In such an example, the block storage 442, 444, and 446 provided by the cloud platform 114 may be utilized in a manner similar to how NVRAM devices are utilized, as described above, and software daemons 428, 432, and 436 (or some other module) executing within a particular cloud computing instance 440A, 440B, and 440n may initiate writing data to the attached EBS volume, as well as to its local storage 430, 434, and 438 resources, in response to receiving a data write request. In some alternative embodiments, data may be written only to local storage 430, 434, 438 resources within a particular cloud computing instance 440A, 440B, 440n. In alternative embodiments, instead of using block storage 442, 444, 446 provided by the cloud platform 114 as NVRAM, the actual RAM of each cloud computing instance 440A, 440B, 440n having local storage 430, 434, 438 may be used as NVRAM, thereby reducing network utilization costs that would be associated with using EBS volumes as NVRAM. In yet another embodiment, a high-performance block storage resource, such as one or more Azure Ultra Disks, may be utilized as NVRAM.

[0059] Storage controller applications 424, 426 may be used to perform various tasks, such as deduplicating the data included in the request, compressing the data included in the request, determining where to write the data included in the request, and ultimately sending the deduplicated, encrypted, and potentially updated version of the data to one or more of cloud computing instances 440A, 440B, 440n using local storage 430, 434, 438. In some embodiments, either of cloud computing instances 420, 422 may receive a request to read data from cloud-based storage system 418 and may ultimately send the request to read the data to one or more of cloud computing instances 440A, 440B, 440n using local storage 430, 434, 438.

[0060] When a request to write data is received by a particular cloud computing instance 440A, 440B, 440n using local storage 430, 434, 438, the software daemons 428, 432, 436 not only write the data to their own local storage 430, 434, 438 resources and appropriate block storage 442, 444, 446 resources, the software daemons 428, 432, 436 may also be configured to write the data to cloud object storage 448 attached to the particular cloud computing instance 440A, 440B, 440n. The cloud object storage 448 attached to the particular cloud computing instance 440A, 440B, 440n may be embodied as, for example, Amazon Simple Storage Service (“S3”). In other embodiments, cloud computing instances 420, 422, each including a storage controller application 424, 426, may begin saving data to local storage 430, 434, 438 of cloud computing instances 440A, 440B, 440n and cloud object storage 448, respectively. In other embodiments, rather than using both cloud computing instances 440A, 440B, 440n and cloud object storage 448 using local storage 430, 434, 438 (also referred to herein as "virtual drives") to store data, the persistent storage tier may be implemented in other ways. For example, one or more Azure Ultra disks may be used to persistently store data (e.g., after the data is written to the NVRAM tier). Warm storage 402 may include cloud object storage 448. Thus, in one embodiment, the DSM utility 116 may communicate with cloud object storage 448, local storage (430, 434, 438), and / or block storage (442, 444, and 446).As described herein, the DSM utility 116 may be configured to enable or otherwise facilitate a user's creation of a distributed file system and retrieval of data sets from warm storage 402 to hot storage 401. In this manner, the DSM utility 116 enables the creation of file systems on cloud object storage 448, local storage (430, 434, 438), and / or block storage (442, 444, and 446). The DSM utility 116 supports the transfer of data sets from cloud object storage 448 to local storage (430, 434, 438) and / or block storage (442, 444, and 446).

[0061] While local storage 430, 434, 438 resources and block storage 442, 444, 446 resources utilized by cloud computing instances 440A, 440B, 440n may support block-level access, the cloud object storage 448 attached to a particular cloud computing instance 440A, 440B, 440n only supports object-based access. Thus, software daemons 428, 432, 436 may be configured to take blocks of data, package those blocks into objects, and write the objects to the cloud object storage 448 attached to a particular cloud computing instance 440A, 440B, 440n.

[0062] Consider an example in which data is written in 1 MB blocks to the local storage 430, 434, 438 resources and the block storage 442, 444, 446 resources utilized by cloud computing instances 440A, 440B, 440n. In such an example, assume that a user of cloud-based storage system 418 issues a request to write data that, after being compressed and deduplicated by storage controller applications 424, 426, results in the need to write 5 MB of data. In such an example, writing data to the local storage 430, 434, 438 resources and the block storage 442, 444, 446 resources utilized by cloud computing instances 440A, 440B, 440n is relatively simple, as five 1 MB blocks are written to the local storage 430, 434, 438 resources and the block storage 442, 444, 446 resources utilized by cloud computing instances 440A, 440B, 440n. In such an example, software daemons 428, 432, 436 may also be configured to create five objects containing separate 1 MB chunks of data. Thus, in some embodiments, each object written to cloud object storage 448 may be identical (or nearly identical) in size. In such an example, metadata associated with the data itself may be included in each object (e.g., the first 1 MB of the object is the data, and the remaining portion is metadata associated with the data). Cloud object storage 448 may be incorporated into cloud-based storage system 418 to increase the durability of cloud-based storage system 418.

[0063] In some embodiments, all data stored by cloud-based storage system 418 may be stored in both 1) cloud object storage 448 and 2) at least one of local storage 430, 434, 438 or block storage 442, 444, 446 resources utilized by cloud computing instances 440A, 440B, 440n. In such embodiments, the local storage 430, 434, 438 and block storage 442, 444, 446 resources utilized by cloud computing instances 440A, 440B, 440n may effectively act as a cache generally containing all data that is also stored in S3, such that all data reads can be serviced by cloud computing instances 440A, 440B, 440n without requiring cloud computing instances 440A, 440B, 440n to access cloud object storage 448. However, in other embodiments, all data stored by cloud-based storage system 418 may be stored in cloud object storage 448, but all data stored by cloud-based storage system 418 may be stored in at least one of the local storage 430, 434, 438 resources or block storage 442, 444, 446 resources utilized by cloud computing instances 440A, 440B, 440n. In such embodiments, various policies may be utilized to determine whether a subset of data stored by cloud-based storage system 418 may reside in both 1) cloud object storage 448 and 2) at least one of the local storage 430, 434, 438 resources or block storage 442, 444, 446 resources utilized by cloud computing instances 440A, 440B, 440n.

[0064] One or more modules of computer program instructions executing within cloud-based storage system 418 (e.g., a monitoring module running on its own EC2 instance) may be designed to handle a failure of one or more of cloud computing instances 440A, 440B, 440n having local storage 430, 434, 438. In such an embodiment, the monitoring module may handle a failure of one or more of cloud computing instances 440A, 440B, 440n having local storage 430, 434, 438 by creating one or more new cloud computing instances using the local storage, retrieving data stored on the failed cloud computing instance 440A, 440B, 440n from cloud object storage 448, and storing the data retrieved from cloud object storage 448 in local storage on the newly created cloud computing instance.

[0065] Various performance aspects of the cloud-based storage system 418 may be monitored (e.g., by a monitoring module running on the EC2 instances) so that the cloud-based storage system 418 can scale up or out as needed. For example, if the cloud computing instances 420, 422 used to support the execution of the storage controller applications 424, 426 are undersized and are not adequately servicing I / O requests issued by users of the cloud-based storage system 418, the monitoring module may create a new, more powerful computing instance that includes the storage controller application (e.g., a type of cloud computing instance that includes more processing power, lots of memory, etc.) so that the new, more powerful cloud computing instance can begin operating as the primary controller. Similarly, if the monitoring module determines that the cloud computing instance 420, 422 used to support the execution of the storage controller application 424, 426 is oversized and cost savings may be obtained by switching to a smaller, less powerful cloud computing instance, the monitoring module may create a new, less powerful (and cheaper) cloud computing instance that includes the storage controller application so that the new, less powerful cloud computing instance can begin operating as the primary controller.

[0066] Returning to FIG. 1 , the cloud platform 114 may include multiple computing nodes (not shown in FIG. 1 for simplicity). The multiple computing nodes communicate with a storage system of the cloud platform 114. The multiple computing nodes may include respective processing devices of one or more processing platforms. For example, the multiple computing nodes may include respective virtual machines (VMs) each having a processor and memory, although numerous other configurations are possible. The multiple computing nodes may additionally or alternatively be part of a cloud infrastructure such as the Amazon Web Services (AWS) system. Other examples of cloud-based systems that can be used to provide computing nodes include Google Cloud Platform (GCP) and Microsoft Azure. In some embodiments, the multiple computing nodes illustratively provide computing services, such as the execution of one or more application programs, for one or more respective users associated with a respective computing node of the multiple computing nodes. The multiple computing nodes may be configured for parallel computing.

[0067] In one embodiment, the cloud platform 114 may be part of a data analysis system. For example, the cloud platform 114 may provide a 3D structure estimation service, a genetic data analysis service (e.g., GEWAS, PHEWAS, etc.), etc. The cloud platform 114 may be configured to perform such data analysis via one or more data analysis modules 118. The data analysis modules 118 may be configured to utilize the computing modules 120. The computing modules 120 may be configured to generate program templates that can be used by at least one of the data analysis modules 118 to govern the execution of one or more processes / tasks, such as using GPU-based computing. The data analysis modules 118 may be configured to output data analysis results, such as an estimated 3D structure of a target in a resulting 3D map (e.g., a 3D model). The cloud platform 114 may also include a remote display module 122. The remote display module 122 may include a high-performance remote display protocol configured to securely stream a remote desktop and application to another computing device 124. For example, the remote display module 122 may be configured as NICE DCV.

[0068] In one embodiment, the data analysis module 118 may be an application program configured to perform image reconstruction (e.g., a reconstruction module). Such an application program (e.g., a reconstruction module) may be configured to perform reconstruction techniques to determine likely molecular structures. Any known technique for determining likely molecular structures may be used. In one embodiment, the application program may include RELION, an open-source program configured to apply an empirical Bayesian approach, where optimal Fourier filters for alignment and reconstruction are derived from the data in a fully automated manner.

[0069] The computing module 120 may be configured to determine one or more job parameters for the data analysis module 118. The one or more job parameters may be referred to as a program template. The program template may enable an application program to manage programs and / or jobs. The program template may enable the application program to utilize computing resources, including, for example, CPU processing time and / or GPU processing time. As an example, a program template may enable an application program (e.g., a reconstruction module) to determine the level of detail to be extracted from the raw data 104 (e.g., raw image data files and / or raw video data files). In one embodiment, the job parameters may include one or more of the number of message passing interfaces (MPIs), the number of threads, the number of compute nodes, the desired wallclock time, combinations thereof, and the like. A particular configuration of job parameters constitutes a particular program template. In one example, a program template is defined by the number of MPIs, the number of threads, and the number of compute nodes. The computing module 120 may be configured to determine such job parameters for one or more portions of a given application program, to include each of one or more tasks or processes of the given application program. 8A shows an example of a program template. The program templates are identified by their respective template names. In some cases, the template names identify the files that contain the program templates, i.e., the files that contain one or more job parameters that define the program templates.

[0070] As described herein, the computing module 120 may assume that the greater the number of MPIs and threads for a job, the greater the performance (e.g., the less time it takes to complete the job). The computing module 120 may assume that disabling hyperthreaded cores may be beneficial to performance. The computing module 120 may implement one or more parameters that specify a multi-GPU and multi-core infrastructure setup with hyperthreaded cores disabled. The computing module 120 may be configured to run one or more simulations to determine one or more job parameters that define a satisfactory (e.g., optimal or near-optimal) program template for a program application or its tasks. For example, the computing module 120 may set the number of MPIs equal to the number of MPIs on available GPU cards and the number of threads equal to the number of MPIs on available CPU cores on a node. After this observation, combinations of multi-node jobs (e.g., 2-, 4-, 6-, 12-node jobs, etc.) may be performed, and performance benchmarks may be compiled. Based on the performance benchmarks, combinations of MPI thread numbers and compute node numbers for jobs that are performance-saturating may be determined, showing that no performance improvement is observed beyond this parallelism.

[0071] In one embodiment, a multi-queue model for executing GPU- and CPU-based computing jobs is disclosed. In one embodiment, the disclosed Cryo-EM system may use the RELION and CryoSPARC applications to process images. A workflow may include a series of jobs (e.g., eight jobs) that are executed to complete image processing. A workflow may include a volume of computationally light steps and a volume of steps that require significant resources (CPU vs. GPU). Configuring a computing node for GPU-based processing for all workflow processing can be costly when processing jobs that require only CPU-based processing.

[0072] In one embodiment, the multi-queue system may be implemented on a (HPC) cluster. An HPC cluster may include hundreds or thousands of computing servers networked together. Each server is called a node. The nodes in each cluster operate in parallel with each other to increase processing speed and achieve high performance computing. A queue may be configured to run on CPU-based computing instances, while another queue may be configured to run on GPU-based computing instances. Users may have the option to select the queue they need to run a particular job and / or workflow.

[0073] In one embodiment, a method for using the best available resources is disclosed. As previously mentioned, RELION is an open-source software package configured to process cryo-EM data and generate protein structural images. The execution of the software depends on various job parameters that determine how the software uses the underlying computing resources. Misconfiguration of these job parameters can lead to underutilization of resources, significantly increasing operational costs and job execution times.

[0074] Cryo-EM job resource usage for all job types (CPU-based jobs and GPU-based jobs) within a cluster can be determined over time. The disclosed method can effectively manage the resources available in the cluster to reduce job runtimes and costs associated with computing and distributed storage. The disclosed method can be applied to multiple stages of job execution. The disclosed method can observe Cryo-EM job resource usage data over time and determine optimized patterns in a template file for future use. The optimized patterns define program templates, i.e., defined sets of job parameters. These optimized patterns can complete jobs many times (e.g., 6-8 times) faster by using fewer computational resources.

[0075] In one embodiment, as shown in FIG. 8B , a computing environment 800 may generate a program template according to aspects described herein. The computing environment 800 may include a job creation module 810 that can receive data 802. In some cases, the data may be received from a data source 102. In other cases, the data 802 may be synthetic in that it may be generated by a computing device for the purpose of performing a simulated reconstruction. The job creation module 810 may generate a job, or tasks associated with a job, to reconstruct one or more targets. In some cases, rather than solving a realistic reconstruction, the job creation module 810 may select a subset of the data 802 and generate or otherwise schedule a job directed to performing a summary simulation (or reconstruction).

[0076] Jobs generated in such a manner may be sent to the template generator module 820, which may generate various configurations of the job parameters. Such configurations may be referred to as job configurations. Each job configuration includes specific values ​​for each job parameter. Thus, such job configurations correspond to each candidate program template. The template generator module 820 may apply multiple strategies to generate job configurations. In some cases, the template generator module 820 may randomly generate job configurations. In other cases, the template generator module 820 may rely on a perturbation approach, in which the template generator module 820 generates variations of existing configurations used in the target production (or actual) reconfiguration. The template generator module 820 may send the job configurations to the compute module 120 for execution on the cloud platform 114 according to the job parameters defined in the job configuration. The template generator module 820 may collect or otherwise receive metrics indicative of the performance of the job's execution using a particular job. Numerous metrics may be collected. Example metrics include wallclock time, GPU time, CPU time, number of I / O operations, execution cost, etc. The collected metric values ​​serve as feedback regarding the suitability of the job configuration for the job. The template generator module 820 can iteratively generate job configurations for the job until satisfactory performance is achieved. To that end, the template generator module 820 may search the space of job parameters using one of a variety of optimization solvers, such as deepest gradient descent, Monte Carlo simulation, genetic algorithms, or the like. The job configuration that results in satisfactory performance (e.g., optimal performance) can determine satisfactory values ​​for the job parameters. Such values ​​define the program template.

[0077] Similar optimizations can be performed for various types of reconfigurations or tasks that are part of a reconfiguration, each of which produces a program template.

[0078] The data analysis module 118 may execute one or more jobs according to the program template to analyze the data. Thus, in some cases, the computing module 120 may select computing nodes within the cloud platform 114 to execute a computing job or a task that is part of a computing job. The selected computing nodes may be part of the computing nodes 403 (FIG. 4). In one embodiment, as shown in FIG. 8C , the computing module 120 includes an interface module 850 that may receive a program template 844 and data 846 defining a job. The program template 844 specifies a set of job parameters and serves as criteria for the selection of computing nodes within the cloud platform 114. For example, the program template may specify n MPIs, m threads, and q computing nodes for a task to be executed (e.g., a reconstruction task). The cloud platform 114 may include multiple sets of q computing nodes that can be selected to execute the task. Furthermore, at least some of the computing nodes may have respective processors, each having multiple cores that may support m threads. Similarly, other computing nodes may support, for example, n MPIs. Thus, the cloud platform 114 may support multiple placements or allocations consistent with a program template.

[0079] In one embodiment, as shown in FIG. 8C , the computing module 120 includes a selection module 860 that can evaluate candidate placements that match the program template. To evaluate the candidate placements, the evaluation component 864 may determine respective performance metrics of each workload on each computing node forming the candidate placement. Each workload may include a computing job defined by data 846. The computing device 106 (FIG. 1) may request the computing job. The evaluation component 864 may determine each performance metric based on measured performance data of each of the computing nodes in the candidate placement. The computing module 120 may obtain measured performance data from one or more components within the cloud platform 114. The measured performance data may include, for example, the current use or supply of one or more resources or other data. The measured performance data may also include or be based on processed data, e.g., values ​​derived from the measured data, such as statistics of the measured data. For example, the average CPU usage and / or average GPU usage on a computing node may be included in the measured performance data for the nodes of the candidate placement.

[0080] The selection module 860 may include a composition component 868 that can traverse a set of multiple candidate placements and evaluate each (or in some cases, at least some) of the candidate placements. This traversal can result in multiple fitness scores for each candidate placement. The composition component 868 can rank the multiple candidate placements according to fitness score and then select the highest or highest ranked one of the candidate placements as the node placement 850 to be utilized to perform the computational job defined by the data 846.

[0081] The data analysis module 118 may store the results of any data analysis in the file system of the cloud platform 114 and / or return the results to the computing device 106. The DSM utility 116 may be used to save the results of the data analysis to a data store and delete the file system from the file system.

[0082] 9 and 10 illustrate an exemplary system and method in which data may be generated via electron microscopes and cached on respective support computing devices. Multiple electron microscopes may generate imaging data as part of their respective electron microscopy experiments. Support computing devices operatively coupled to each of the electron microscopes may acquire and cache the imaging data. The imaging data from the support computing devices may be pushed to a local staging area. On a schedule (e.g., hourly, daily, at a defined time, etc.), the imaging data from the staging area may be pushed to a storage system such as cloud-based storage (e.g., AWS S3). A separate scheduled data synchronization task may continue to push the data to respective data store buckets (e.g., S3 buckets). The imaging data is visible from a storage gateway. Scheduled automatic cache updates may be used. Datasets required for processing may be mounted on a master / computation notebook via the DSM utility, and the storage used may be distributed and / or parallel (e.g., FSx-Lustre).

[0083] 11 is a block diagram illustrating an environment 1100 including a non-limiting example of a computing device 106 and a cloud platform 114 connected through a network 1104. In one aspect, some or all steps of the described methods may be performed on the computing device and / or cloud platform described herein. The computing device 106 may include one or more computers configured to store one or more of the data 104, the data synchronization manager 110, and / or the data synchronization module 112. The cloud platform 114 may include a high-throughput storage system 1106 configured to store the data 104, the DSM utility 116, the data analysis module 118, the calculation module 120, the remote display module 122, and / or one or more computational nodes 1108 configured to process the data 104. The cloud platform 114 can communicate with the computing device 106 via the network 1104.

[0084] In terms of hardware architecture, the computing device 106 and the cloud platform 114 may generally be one or more digital computers including a processor 1110, a memory system 1112, an input / output (I / O) interface 1114, and a network interface 1116. These components (1110, 1112, 1114, and 1116) are communicatively coupled via a local interface 1118. The local interface 1118 may be, for example, but not limited to, one or more buses or other wired or wireless connections as known in the art. The local interface 1118 may have additional elements that enable communication, such as controllers, buffers (caches), drivers, repeaters, and receivers, which are omitted for simplicity. Furthermore, the local interface may include address, control, and / or data connections to enable appropriate communication between the aforementioned components.

[0085] Processor 1110 may be one or more hardware devices for executing software, particularly stored in memory system 1112. Processor 1110 may be any custom-made or commercially available processor, a central processing unit (CPU), a coprocessor among several processors associated with computing device 106 and cloud platform 114, a semiconductor-based microprocessor (in the form of a microchip or chipset), or generally any device for executing software instructions. When computing device 106 and / or cloud platform 114 are operating, processor 1110 may be configured to execute software stored in memory system 1112, communicate data to and from memory system 1112, and generally control the operation of computing device 106 and cloud platform 114 in accordance with the software.

[0086] The I / O interface 1114 can be used to receive user input from one or more devices or components and / or provide system output. User input can be provided, for example, via a keyboard and / or a mouse. System output can be provided via a display device and a printer (not shown). The I / O interface 1114 can include, for example, a serial port, a parallel port, a small computer system interface (SCSI), an infrared (IR) interface, a radio frequency (RF) interface, and / or a universal serial bus (USB) interface.

[0087] The network interface 1116 may be used to transmit and receive data from the computing device 106 and / or the cloud platform 114 over the network 1104. The network interface 1116 may be, for example, a 10BaseT Ethernet (registered trademark) Adaptor, 100BaseT Ethernet Adaptor, LAN PHY Ethernet Adaptor, Token The network interface 1116 may include a Ring Adaptor, a wireless network adapter (e.g., WiFi, cellular, satellite), or any other suitable network interface device. The network interface 1116 may include address, control, and / or data connections to enable appropriate communication over the network 1104.

[0088] The memory system 1112 may include any one or combination of volatile memory elements (e.g., random access memory (e.g., RAM, such as DRAM, SRAM, SDRAM, etc.)) and non-volatile memory elements (e.g., ROM, hard drives, tape, CD-ROM, DVD-ROM, etc.). Additionally, the memory system 1112 may incorporate electronic, magnetic, optical, and / or other types of storage media. The memory system 1112 may have a distributed architecture, where various components are located remotely from one another but can be accessed by the processor 1110.

[0089] 11, the software in the memory system 1112 of the computing device 106 may include data 104, a data staging module 108, a data synchronization manager 110, a data synchronization module 112, a policy 260, a suitable operating system (O / S) 1120, and / or any other modules (e.g., the example modules disclosed in FIG. 1). In the example of FIG. 11, the software in the high-throughput storage system 1106 of the cloud platform 114 may include data 104, a DSM utility 116, a data analysis module 118, a calculation module 120, a remote display module 122, a suitable operating system (O / S) 1120, and / or any other modules (e.g., the example modules disclosed in FIG. 1). The operating system 1120 essentially controls the execution of other computer programs and provides scheduling, input-output control, file and data management, memory management, and communication control and related services.

[0090] For purposes of illustration, application programs and other executable program components, such as operating system 1120, are illustrated herein as separate blocks, with the understanding that such programs and components may reside at various times in different storage components of computing device 106 and / or cloud platform 114. An implementation of data synchronization manager 110, data synchronization module 112, DSM utility 116, data analysis module 118, calculation module 120, and / or remote display module 122 may be stored on or transmitted across some form of computer-readable media. Any of the disclosed methods may be implemented by computer-readable instructions embodied on a computer-readable medium. A computer-readable medium may be any available medium that can be accessed by a computer. By way of example, and not intended to be limiting, computer-readable media may include “computer storage media” and “communications media.” "Computer storage media" may include volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules, or other data. Exemplary computer storage media may include RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage, or any other medium that can be used to store the desired information and that can be accessed by a computer.

[0091] In one embodiment, the data synchronization manager 110 and / or the data synchronization module 112 are configured to perform an example method 1200 shown in FIG. 12 . The example method 1200 may be performed in whole or in part by a single computing device, multiple electronic devices, and the like. The example method 1200 may include, at block 1210, receiving an indication of a synchronization request. Receiving the indication of the synchronization request may be based on a synchronization condition. In some cases, the synchronization condition is a time interval. The indication includes payload data that conveys that data synchronization is to be implemented. In some cases, the indication may be embodied in a message that invokes, for example, a function call to a data storage service.

[0092] The example method 1200 may include, at block 1220, determining one or more files stored at the staging location based on the instructions. Various types of files may be determined. For example, the one or more files may include sequence data, particle images, or a combination of sequence data and particle images.

[0093] The example method 1200 may include, at block 1230, generating a data transfer filter based on the one or more files. Generating the data transfer filter may include generating a message that invokes a function call to a cloud service (e.g., AWS DataSync), where the message passes the list of available files as arguments to the function call. The function call may initiate a task (or job) in the cloud service. The function call may be retrieved according to an API implemented by the data storage service. In some instances, the data transfer filter includes a list of one or more files stored in a staging location.

[0094] The example method 1200 may include causing a transfer of one or more files to a destination computing device based on the data transfer filter at block 1240. Triggering such a transfer based on the data transfer filter may include having the data synchronization application program scan the staging location and the destination computing device only for the one or more files.

[0095] The example method 1200 may include receiving one or more files from a data source device, at block 1250. The data source device may include one or more of a sequencer or an electron microscope.

[0096] The example method 1200 may include, at block 1260, deleting the one or more files from the staging location based on the transfer of the one or more files to the destination computing device.

[0097] In one embodiment, the DSM utility 116 may be configured to perform an example method 1300 shown in Figure 13. Method 1300 may be performed in whole or in part by a single computing device, multiple electronic devices, and the like. The example method 1300 may include, at block 1310, receiving a request via a graphical user interface to migrate a dataset from object storage to a distributed file system.

[0098] The example method 1300 may include, at block 1320, receiving an indication of a storage size of the distributed file system via a graphical user interface.

[0099] The example method 1300 may include, at block 1330, migrating the dataset from the object storage to a distributed file system associated with the storage size based on the request and the indication.

[0100] The example method 1300 may include, at block 1340, receiving a request to perform an operation involving the distributed file system. The example method 1300 may include, at block 1350, performing an operation. The operation may be one or more operations involving the distributed file system. In one scenario, at block 1340, the example method 1300 includes receiving a request to mount the distributed file system via a graphical user interface. Additionally, at block 1350, the example method 1300 includes mounting the distributed file system. In another scenario, at block 1340, the example method 1300 includes receiving a request to save data in the distributed file system to object storage via a graphical user interface. Additionally, at block 1350, the example method 1300 includes saving data in the distributed file system to object storage. In yet another scenario, at block 1340, the example method 1300 includes receiving a request to delete the distributed file system via a graphical user interface. Additionally, at block 1350, the example method 1300 includes deleting the distributed file system.

[0101] In one embodiment, the data analysis module 118 and / or the computing module 120 may be configured to perform a method 1400, shown in Figure 14. The method 1400 may be performed in whole or in part by a single computing device, multiple electronic devices, and the like. The method 1400 may include, at block 1410, identifying a data analysis application program.

[0102] The example method 1400 may include, at block 1420, identifying a data set associated with the data analysis application program.

[0103] The example method 1400 may include determining one or more job parameters associated with the data analysis application program that processes the dataset as a program template at block 1430. Determining the one or more job parameters associated with the data analysis application program that processes the dataset may include determining one or more job parameters for each task of the multiple tasks. The one or more job parameters may include one or more of a message passing interface (MPI), a number of threads, or a number of compute nodes.

[0104] The example method 1400 may include, at block 1440, causing execution of a data analysis application program on the dataset based on the program template.

[0105] The example method 1400 may include, at block 1450, determining a plurality of tasks executable by the data analysis application program.

[0106] In one embodiment, the data synchronization manager 110, the data synchronization module 112, the DSM utility 116, the data analysis module 118, and / or the computation module 120 may be configured to perform an example method 1500 shown in FIG. 15 . The example method 1500 may be performed in whole or in part by a single computing device, multiple electronic devices, and the like. The example method 1500 may include, at block 1510, receiving an indication of a synchronization request. The indication includes payload data that conveys that data synchronization is to be implemented. In some cases, the indication may be embodied in a message that invokes, for example, a function call to a data storage service. Receiving the indication of a synchronization request may be based on a synchronization condition. In some cases, the synchronization condition is a time interval.

[0107] The example method 1500 may include, at block 1520, determining one or more files stored at the staging location based on the instruction.

[0108] The example method 1500 may include, at block 1530, generating a data transfer filter based on the one or more files.

[0109] The example method 1500 may include, at block 1540, causing a transfer of the one or more files to object storage of the destination computing device based on the data transfer filter.

[0110] The example method 1500 may include, at block 1550, receiving a request via a graphical user interface to migrate one or more files from the object storage to the distributed file system.

[0111] The example method 1500 may include, at block 1560, receiving an indication of a storage size of the distributed file system via a graphical user interface.

[0112] The example method 1500 may include, at block 1570, migrating one or more files from the object storage to a distributed file system associated with the storage size based on the request and the instruction.

[0113] The example method 1500 may include, at block 1580, identifying a data analysis application program associated with one or more files in the distributed file system.

[0114] The example method 1500 may include, at block 1590, determining, as a program template, one or more job parameters associated with a data analysis application program that processes the data set.

[0115] Exemplary method 1500 may include, at block 1595, causing execution of a data analysis application program associated with one or more files in the distributed file system based on the program template. Numerous other embodiments emerge from the foregoing detailed description and accompanying drawings. For example, example 1 of these embodiments includes a method that includes receiving an indication of a synchronization request, determining one or more files stored at a staging location based on the indication, generating a data transfer filter based on the one or more files, and causing transfer of the one or more files to a destination computing device based on the data transfer filter.

[0116] Example 2 of many embodiments includes the method of example 1, wherein receiving the indication of the synchronization request is based on a synchronization condition.

[0117] Example 3 of many embodiments includes the method of Example 2, in which the synchronization condition is a time interval.

[0118] Example 4 of many embodiments includes the method of example 1, wherein the data transfer filter includes a list of one or more files stored in the staging location.

[0119] Example 5 of many embodiments includes the method of Example 1, wherein generating the data transfer filter based on the one or more files includes generating a message that invokes a function call to a cloud service, the message passing one or more parameters that identify the one or more files as arguments to the function call.

[0120] Example 6 of many embodiments includes the method of Example 1, wherein causing the transfer of one or more files to the destination computing device based on the data transfer filter includes causing the data synchronization application program to scan the staging location and the destination computing device only for the one or more files.

[0121] Example 7 of the numerous embodiments includes the method of Example 1, further including receiving one or more files from the data source device.

[0122] Example 8 of the numerous embodiments includes the method of Example 7, wherein the data source device includes one or more of a sequencer or an electron microscope.

[0123] Example 9 of many embodiments includes the method of Example 8, wherein the one or more files include sequence data, particle images, or both.

[0124] Example 10 of the many embodiments includes the method of Example 1, and further includes deleting one or more files from the staging location based on the transfer of the one or more files to the destination computing device.

[0125] Example 11 of these many other embodiments includes a method including receiving, via a graphical user interface, a request to migrate a dataset from object storage to a distributed file system; receiving, via the graphical user interface, an indication of a storage size of the distributed file system; and, based on the request and the indication, migrating the dataset from the object storage to the distributed file system associated with the storage size.

[0126] Example 12 of many embodiments includes the method of Example 11, further including receiving a request to mount the distributed file system via a graphical user interface, and mounting the distributed file system.

[0127] Example 13 of many embodiments includes the method of Example 11, further including receiving a request via a graphical user interface to store data in the distributed file system in object storage, and storing the data in the distributed file system in the object storage.

[0128] Example 14 of many embodiments includes the method of Example 11, further including receiving a request to delete the distributed file system via a graphical user interface and deleting the distributed file system.

[0129] Example 15 of many embodiments includes a method including identifying a data analysis application program; identifying a dataset associated with the data analysis application program; determining one or more job parameters associated with the data analysis application program for processing the dataset as a program template; and causing execution of the data analysis application program on the dataset based on the program template.

[0130] Example 16 of many embodiments includes the method of Example 15, wherein the one or more job parameters include one or more of a number of message passing interfaces (MPIs), a number of threads, or a number of compute nodes.

[0131] Example 17 of many embodiments includes the method of Example 15, further including determining a plurality of tasks executable by the data analysis application program.

[0132] Example 18 of many embodiments includes the method of Example 17, wherein determining one or more job parameters associated with the data analysis application program that processes the dataset includes determining one or more job parameters for each task of the plurality of tasks.

[0133] Example 19 of numerous embodiments includes a method including receiving an indication of a synchronization request; determining one or more files stored at a staging location based on the indication; generating a data transfer filter based on the one or more files; causing a transfer of the one or more files to object storage on a destination computing device based on the data transfer filter; receiving a request to migrate one or more files from the object storage to a distributed file system via a graphical user interface; receiving an indication of a storage size of the distributed file system via the graphical user interface; moving the one or more files from the object storage to the distributed file system associated with the storage size based on the request and the indication; identifying a data analysis application program associated with the one or more files in the distributed file system; determining one or more job parameters associated with the data analysis application program that processes the dataset as a program template; and causing execution of the data analysis application associated with the one or more files in the distributed file system based on the program template.

[0134] Example 20 of many embodiments includes a computing system having at least one processor and at least one memory device having stored thereon processor-executable instructions that, in response to execution by the at least one processor, cause the computing system to receive an indication of a synchronization request, determine one or more files stored at a staging location based on the instructions, generate a data transfer filter based on the one or more files, and cause a transfer of the one or more files to a destination computing device based on the data transfer filter.

[0135] Example 21 of many embodiments includes the method of Example 20, wherein receiving the indication of the synchronization request is based on a synchronization condition.

[0136] Example 22 of many embodiments includes the method of Example 21, wherein the synchronization condition is a time interval.

[0137] Example 23 of many embodiments includes the method of Example 20, wherein the data transfer filter includes a list of one or more files stored in the staging location.

[0138] Example 24 of many embodiments includes the method of Example 20, wherein generating a data transfer filter based on the one or more files includes generating a message that invokes a function call to a cloud service, the message passing one or more parameters that identify the one or more files as arguments to the function call.

[0139] Example 25 of many embodiments includes the method of Example 20, wherein causing the transfer of one or more files to the destination computing device based on the data transfer filter includes having the data synchronization application program scan the staging location and the destination computing device only for the one or more files.

[0140] Example 26 of many embodiments includes the method of Example 20, wherein the at least one memory device further stores processor-executable instructions that, in response to execution by the at least one processor, further cause the computing system to receive one or more files from the data source device.

[0141] Example 27 of the numerous embodiments includes the method of Example 26, wherein the data source device includes one or more of a sequencer or an electron microscope.

[0142] Example 28 of many embodiments includes the method of Example 27, wherein the one or more files include sequence data, particle images, or both.

[0143] Example 29 of many embodiments includes the method of Example 20, wherein the at least one memory device further stores processor-executable instructions that, in response to execution by the at least one processor, further cause the computing system to delete one or more files from the staging location based on the transfer of the one or more files to the destination computing device.

[0144] Example 30 of many embodiments includes a computing system having at least one processor and at least one memory device having stored thereon processor-executable instructions that, in response to execution by the at least one processor, cause the computing system to receive a request to migrate a dataset from object storage to a distributed file system via a graphical user interface, receive an indication of a storage size of the distributed file system via the graphical user interface, and, based on the request and the indication, migrate the dataset from object storage to a distributed file system associated with the storage size.

[0145] Example 31 of many embodiments includes the computing system of Example 30, wherein the at least one memory device further stores processor-executable instructions that, in response to execution by the at least one processor, further cause the computing system to receive a request to mount the distributed file system via a graphical user interface and mount the distributed file system.

[0146] Example 32 of many embodiments includes the computing system of Example 30, wherein the at least one memory device further stores processor-executable instructions that, in response to execution by the at least one processor, further cause the computing system to receive a request via a graphical user interface to store data in the distributed file system in object storage, and receive a request to store data in the distributed file system in object storage.

[0147] Example 33 of many embodiments includes the computing system of Example 30, wherein the at least one memory device further stores processor-executable instructions that, in response to execution by the at least one processor, further cause the computing system to receive a request to delete the distributed file system via a graphical user interface and delete the distributed file system.

[0148] Example 34 of many embodiments includes a computing system comprising at least one processor and at least one memory device having stored thereon processor-executable instructions that, in response to execution by the at least one processor, cause the computing system to identify a data analysis application program, identify a dataset associated with the data analysis application program, determine one or more job parameters associated with the data analysis application program for processing the dataset as a program template, and cause execution of the data analysis application program on the dataset based on the program template.

[0149] Example 35 of many embodiments includes the computing system of Example 34, wherein the one or more job parameters include one or more of the number of message passing interfaces (MPIs), the number of threads, or the number of compute nodes.

[0150] Example 36 of many embodiments includes the computing system of Example 34, wherein the at least one memory device further stores processor-executable instructions that, in response to execution by the at least one processor, further cause the computing system to determine a plurality of tasks executable by the data analysis application program.

[0151] Example 37 of many embodiments includes the computing system of Example 36, wherein determining one or more job parameters associated with the data analysis application program that processes the dataset includes determining one or more job parameters for each task of the plurality of tasks.

[0152] Example 38 of many embodiments includes an apparatus having at least one processor and at least one memory device having stored thereon processor-executable instructions that, in response to execution by the at least one processor, cause a computing system to receive an indication of a synchronization request, determine one or more files stored at a staging location based on the instructions, generate a data transfer filter based on the one or more files, and cause a transfer of the one or more files to a destination computing device based on the data transfer filter.

[0153] Example 39 of numerous embodiments includes the apparatus of example 38, wherein receiving the indication of the synchronization request is based on a synchronization condition.

[0154] Example 40 of many embodiments includes the apparatus of example 39, wherein the synchronization condition is a time interval.

[0155] Example 41 of many embodiments includes the apparatus of example 38, wherein the data transfer filter includes a list of one or more files stored in the staging location.

[0156] Example 42 of many embodiments includes the apparatus of Example 38, wherein generating a data transfer filter based on the one or more files includes generating a message that invokes a function call to a cloud service, the message passing one or more parameters that identify the one or more files as arguments to the function call.

[0157] Example 43 of many embodiments includes the apparatus of Example 38, wherein causing the transfer of one or more files to the destination computing device based on the data transfer filter includes causing the data synchronization application program to scan the staging location and the destination computing device only for the one or more files.

[0158] Example 44 of many embodiments includes the apparatus of Example 38, wherein the at least one memory device further stores processor-executable instructions that, in response to execution by the at least one processor, further cause the computing system to receive one or more files from the data source device.

[0159] Example 45 of the numerous embodiments includes the apparatus of Example 44, wherein the data source device includes one or more of a sequencer or an electron microscope.

[0160] Example 46 of many embodiments includes the apparatus of Example 45, wherein the one or more files include sequence data, particle images, or both.

[0161] Example 47 of many embodiments includes the apparatus of Example 38, and further includes deleting one or more files from the staging location based on the transfer of the one or more files to the destination computing device.

[0162] Example 48 of many embodiments includes an apparatus having at least one processor and at least one memory device having stored thereon processor-executable instructions that, in response to execution by the at least one processor, further cause a computing system to receive, via a graphical user interface, a request to migrate a dataset from object storage to a distributed file system, receive, via the graphical user interface, an indication of a storage size of the distributed file system, and, based on the request and the indication, migrate the dataset from object storage to a distributed file system associated with the storage size.

[0163] Example 49 of many embodiments includes the apparatus of Example 48, wherein the at least one memory device further stores processor-executable instructions that, in response to execution by the at least one processor, further cause the computing system to receive a request to mount the distributed file system via a graphical user interface and mount the distributed file system.

[0164] Example 50 of many embodiments includes the apparatus of Example 48, wherein the at least one memory device further stores processor-executable instructions that, in response to being executed by the at least one processor, further cause the computing system to receive, via a graphical user interface, a request to store data in the distributed file system in object storage, and store the data in the distributed file system in the object storage.

[0165] Example 51 of many embodiments includes the apparatus of Example 48, wherein the at least one memory device further stores processor-executable instructions that, in response to execution by the at least one processor, further cause the computing system to receive, via a graphical user interface, a request to delete the distributed file system and delete the distributed file system.

[0166] Example 52 of many embodiments includes an apparatus comprising at least one processor and at least one memory device having stored thereon processor-executable instructions that, in response to execution by the at least one processor, cause a computing system to identify a data analysis application program, identify a dataset associated with the data analysis application program, determine one or more job parameters associated with the data analysis application program for processing the dataset as a program template, and cause execution of the data analysis application program on the dataset based on the program template.

[0167] Example 53 of many embodiments includes the apparatus of Example 52, wherein the one or more job parameters include one or more of the number of message passing interfaces (MPIs), the number of threads, or the number of compute nodes.

[0168] Example 54 of many embodiments includes the apparatus of Example 52, wherein the at least one memory device further stores processor-executable instructions that, in response to execution by the at least one processor, further cause the apparatus to determine a plurality of tasks executable by the data analysis application program.

[0169] Example 55 of many embodiments includes the apparatus of Example 54, wherein determining one or more job parameters associated with the data analysis application program that processes the dataset includes determining one or more job parameters for each task of the plurality of tasks.

[0170] Example 56 of many embodiments includes at least one computer-readable non-transitory storage medium having processor-executable instructions stored thereon that, in response to execution, cause a computing system to receive an indication of a synchronization request, determine one or more files stored at a staging location based on the instructions, generate a data transfer filter based on the one or more files, and transfer the one or more files to a destination computing device based on the data transfer filter.

[0171] Example 57 of numerous embodiments includes the at least one computer-readable non-transitory storage medium of Example 56, wherein receiving an indication of a synchronization request is based on a synchronization condition.

[0172] Example 58 of numerous embodiments includes the at least one computer-readable non-transitory storage medium of Example 57, wherein the synchronization condition is a time interval.

[0173] Example 59 of many embodiments includes at least one computer-readable non-transitory storage medium of Example 56, wherein the data transfer filter includes a list of one or more files stored in the staging location.

[0174] Example 60 of many embodiments includes at least one computer-readable non-transitory storage medium of Example 56, wherein generating a data transfer filter based on the one or more files includes generating a message that invokes a function call to a cloud service, the message passing one or more parameters identifying the one or more files as arguments to the function call.

[0175] Example 61 of many embodiments includes at least one computer-readable non-transitory storage medium of Example 56, and causing the transfer of one or more files to a destination computing device based on a data transfer filter includes causing a data synchronization application program to scan the staging location and the destination computing device only for the one or more files.

[0176] Example 62 of many embodiments includes at least one computer-readable non-transitory storage medium of Example 56, wherein the processor-executable instructions, in response to further execution, further cause the computing system to receive one or more files from the data source device.

[0177] Example 63 of many embodiments includes at least one computer-readable non-transitory storage medium of Example 62, wherein the data source device includes one or more of a sequencer or an electron microscope.

[0178] Example 64 of many embodiments includes at least one computer-readable non-transitory storage medium of Example 63, wherein the one or more files include sequence data, particle images, or both.

[0179] Example 65 of many embodiments includes at least one computer-readable non-transitory storage medium of Example 56, wherein the processor-executable instructions, in response to further execution, further cause the computing device to delete one or more files from the staging location based on the transfer of the one or more files to the destination computing device.

[0180] Example 66 of many embodiments includes at least one computer-readable non-transitory storage medium having processor-executable instructions stored thereon that, when executed, cause a computing system to receive, via a graphical user interface, a request to migrate a dataset from object storage to a distributed file system, receive, via the graphical user interface, an indication of a storage size of the distributed file system, and, based on the request and the indication, migrate the dataset from the object storage to the distributed file system associated with the storage size.

[0181] Example 67 of many embodiments includes at least one computer-readable non-transitory storage medium of Example 66, wherein the processor-executable instructions, in response to further execution, further cause the computing system to receive a request to mount the distributed file system via a graphical user interface and mount the distributed file system.

[0182] Example 68 of many embodiments includes at least one computer-readable non-transitory storage medium of Example 66, wherein the processor-executable instructions, in response to further execution, further cause the computing system to receive, via a graphical user interface, a request to store data in the distributed file system in object storage, and store the data in the distributed file system in the object storage.

[0183] Example 69 of many embodiments includes at least one computer-readable non-transitory storage medium of Example 66, wherein the processor-executable instructions, in response to further execution, further cause the computing system to receive, via a graphical user interface, a request to delete the distributed file system and delete the distributed file system.

[0184] Example 70 of numerous embodiments includes at least one computer-readable non-transitory storage medium having processor-executable instructions stored thereon that, in response to execution, cause a computing system to identify a data analysis application program, identify a dataset associated with the data analysis application program, determine one or more job parameters associated with the data analysis application program for processing the dataset as a program template, and cause execution of the data analysis application program on the dataset based on the program template.

[0185] Example 71 of many embodiments includes at least one computer-readable non-transitory storage medium of Example 70, wherein the one or more job parameters include one or more of the number of message passing interfaces (MPIs), the number of threads, or the number of compute nodes.

[0186] Example 72 of many embodiments includes at least one computer-readable non-transitory storage medium of Example 70, wherein the processor-executable instructions, in response to further execution, further cause the computing system to determine a plurality of tasks executable by the data analysis application program.

[0187] Example 73 of many embodiments includes at least one computer-readable non-transitory storage medium of Example 70, wherein determining one or more job parameters associated with a data analysis application program that processes the dataset includes determining one or more job parameters for each task of the plurality of tasks.

[0188] Example 74 of many embodiments includes a computing device having at least one processor and at least one memory device further having processor-executable instructions stored thereon that, in response to execution by the at least one processor, further cause the computing system to receive an indication of a synchronization request, determine one or more files stored in a staging location based on the instructions, generate a data transfer filter based on the one or more files, cause the one or more files to be transferred to object storage on a destination computing device based on the data transfer filter, receive a request to migrate one or more files from the object storage to a distributed file system via a graphical user interface, receive an indication of a storage size of the distributed file system via the graphical user interface, and migrate one or more files from the object storage to the distributed file system associated with the storage size based on the request and the instructions, identify a data analysis application program associated with one or more files in the distributed file system, determine one or more job parameters associated with the data analysis application program that processes a dataset as a program template, and cause execution of the data analysis application program associated with the one or more files in the distributed file system based on the program template.

[0189] Example 75 of many embodiments includes at least one computer-readable non-transitory storage medium having processor-executable instructions stored thereon that, in response to execution, cause a computing system to receive an indication of a synchronization request, determine one or more files stored at a staging location based on the indication, determine one or more files stored at a staging location based on the indication, generate a data transfer filter based on the one or more files, cause the transfer of the one or more files to object storage on a destination computing device based on the data transfer filter, receive a request to migrate one or more files from the object storage to a distributed file system via a graphical user interface, receive an indication of a storage size of the distributed file system via the graphical user interface, and based on the request and the indication, migrate one or more files from the object storage to the distributed file system associated with the storage size, identify a data analysis application program associated with one or more files in the distributed file system, determine one or more job parameters associated with the data analysis application program that processes a dataset as a program template, and cause execution of the data analysis application program associated with the one or more files in the distributed file system based on the program template.

[0190] The disclosed methods and systems can be configured for big data collection and real-time analysis. The disclosed methods and systems are configured for ultra-fast end-to-end processing of raw Cryo-EM data and reconstruction of electron density maps that are easy to incorporate into model-building software.

[0191] The disclosed methods and systems optimize reconstruction algorithms and GPU acceleration at one or more stages, from pre-processing, to particle collection, 2D particle classification, 3D ab-initio structure determination, high-resolution refinement, and heterogeneity analysis.

[0192] The disclosed methods and systems enable real-time Cryo-EM data quality assessment and decision-making during live data collection, as well as a rapid, streamlined workflow for processing already available data.

[0193] The disclosed methods and systems work with specialized and unique tools (eg, RELION) for therapeutically relevant targets, membrane proteins, and continuously flexible structures.

[0194] The disclosed method and system includes a computing platform that has good bandwidth on processing, storage for faster processing, and computation thereby reducing computing execution time, which are costly resources.

[0195] The disclosed methods and systems can be configured as a self-service, cloud-based computing platform, allowing scientists to run multiple analytical processes on demand without IT dependencies or computational design decisions. The disclosed methods and systems have broad and flexible applications regardless of data type, size, or experiment type.

[0196] The disclosed methods and systems are capable of determining the detailed structure of the binding complex between a potential therapeutic antibody and a target protein.

[0197] The disclosed methods and systems can be configured as a platform that enables scientists to scale and process vast amounts of images in a timely manner, with high levels of quality and agility, while containing costs.

[0198] The disclosed methods and systems can be configured as automated end-to-end processing pipelines by utilizing AWS Datasync, Apache Airflow (for orchestration), Luster Filesystem (for high-throughput storage) NextFlow, and AWS Parallel Cluster Framework for model development, enabling the transmission and processing of large amounts of data (e.g., 1 TB / hour of raw data) over long periods of time.

[0199] The disclosed methods and systems may integrate RELION for real-time Cryo-EM data quality assessment and decision making during data collection.

[0200] The disclosed methods and systems may extend the AWS parallel computing framework to accommodate GPU-based computing.

[0201] The disclosed methods and systems may include data management and tiering tools that allow user management of the lifecycle of data.

[0202] The disclosed method and system may implement high performance remote display protocols such as NICE DCV to provide graphics-intensive applications to remote users, streaming the user interface to any client machine and eliminating the need for a dedicated workstation.

[0203] The disclosed methods and systems can utilize blue-green high-performance computing, a concept generally limited to software development, to address cryo-EM data quality assessment and decision-making during collection. As a result, job processing is sped up and scaled up.

[0204] Unlike previous data pipelines with similar workload characteristics, which take 3-5 days to preprocess the data, the disclosed methods and systems can speed up the Cryo-EM pipeline to approximately 60 minutes / 1TB of data, for example, by capturing raw data, preprocessing, classifying, reconstructing, and refining the 3D map while the sample is still in the microscope.

[0205] Using the methods and systems disclosed herein, scientists can test and refine collection strategies and fine-tune sample preparation while data collection is ongoing. As a result, time wasted on poor samples can be minimized, allowing scientists to make on-the-fly decisions while performing microscopy experiments. Leveraging RELION, scientists can use 2D and 3D information to assess desired orientation by adjusting imaging and processing parameters in real time. The disclosed methods and systems may also enable PLUGIN interoperability with the NIFTY processing framework to address limited noisy signal and resolution quality and improve signal quality (e.g., signal-to-noise ratio).

[0206] The disclosed method and system can be configured as a managed service that provides users with instant access to RELION and its related applications from anywhere.

[0207] The disclosed method and system represent a scalable, cloud-based data processing and computing platform that supports large volume Cryo-EM type data pipelines. The disclosed method and system offer significant benefits in a cloud-based solution that is scalable, agile, and responsive to ever-changing research needs.

[0208] The disclosed methods and systems can be applied to other areas of research, such as large-scale sequencing, imaging, and other high-throughput biological research efforts.

[0209] Although particular configurations are described, the configurations herein are intended in all respects to be possible configurations rather than limiting, and therefore the scope is not intended to be limited to the particular configurations stated.

[0210] Unless expressly stated otherwise, it is in no way intended that any method described herein be construed as requiring its steps to be performed in a particular order. Thus, where a method claim does not actually recite the order in which its steps are to be followed, or where the claim or description does not otherwise specifically state that the steps are limited to a particular order, no order is intended to be inferred in any respect. This preserves any possible implicit basis for interpretation, including matters of logic regarding the placement or operational flow of steps, apparent meaning derived from grammatical construction or punctuation, and the number or type of constructions described herein.

[0211] It will be apparent to those skilled in the art that various modifications and variations can be made without departing from the scope or spirit of the present invention. Other configurations will be apparent to those skilled in the art from consideration of the specification and practice described herein. It is intended that the specification and described configurations be considered as exemplary only, with a true scope and spirit being indicated by the following claims.

Claims

1. 1. A method for estimating the three-dimensional molecular structure of a particle through reconstruction based on two-dimensional particle images from a cryo-electron microscope, comprising: receiving the two-dimensional particle image for reconstruction using a data analysis application program; generating synthetic data and selecting one or more subsets of the synthetic data for simulating summary reconstruction; determining one or more job parameters associated with the data analysis application program based on a plurality of computing nodes and the summary reconstruction, the one or more job parameters characterizing a possible utilization of the plurality of computing nodes by the data analysis application program; perturbatively generating a plurality of job parameter configurations based on a plurality of existing configurations used in at least one of a previous production reconfiguration and the summary reconfiguration; determining a program template based on the one or more job parameters and the perturbatively generated plurality of job parameter configurations by iteratively processing the plurality of job parameter configurations using at least one optimization solver, the program template defining a level of detail to extract from the two-dimensional particle image to generate a three-dimensional particle representation at a specified resolution; causing the plurality of computing nodes to execute the data analysis application program on the two-dimensional particle image to extract the level of detail from the two-dimensional particle image based on the program template; causing generation of said three-dimensional particle representation at a desired resolution.

2. The method of claim 1 , wherein the one or more job parameters include one or more of a number of message passing interfaces (MPIs), a number of threads, or a number of the plurality of compute nodes.

3. determining a plurality of tasks executable by the data analysis application program; The method of claim 1 , wherein determining the one or more job parameters associated with the data analysis application program comprises determining one or more job parameters for each task of the plurality of tasks.

4. The method of claim 1 , wherein determining the one or more job parameters comprises determining, for each of the plurality of compute nodes, at least one of a number of graphics processing units and a number of central processing units.

5. The method of claim 1 , wherein the at least one optimization solver comprises a steepest descent method, a Monte Carlo simulation, or a genetic algorithm.

6. The method of claim 1 , wherein satisfactory performance of causing the generation of the three-dimensional particle representation is based on the at least one optimization solver.

Citation Information

Patent Citations

  • Deferred copy-on-write of a snapshot

    US20030159007A1

  • Cloud gateway for ZFS snapshot generation and storage

    US20180196817A1