Data storage for scalable processing of large files generated by scientific instruments
The cache management system optimizes data access for large scientific instrument files by converting and naming data files for quick retrieval, addressing latency and cost issues in conventional systems, enabling efficient and scalable data access.
Patent Information
- Application Number
- JP2025532482
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-29
- Filing Date
- 2023-12-29
- Publication Date
- 2026-01-27
AI Technical Summary
Conventional data storage systems face challenges in efficiently accessing large data files generated by scientific instruments like TOF mass spectrometers, leading to significant data access delays and high storage costs, particularly when dealing with large binary data files that are accessed infrequently and require selective retrieval.
A cache management system is designed to optimize data access by converting large raw instrument data files into smaller files, using metadata and naming conventions for quick retrieval, and employing multi-tier caching with different eviction policies to reduce latency and improve scalability.
The solution provides faster and more selective access to scientific instrument data, reducing data access latency and storage costs while supporting horizontal scaling and parallel access from multiple client devices, enhancing data processing efficiency.
Smart Images

Figure 2026502818000001_ABST
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims priority to U.S. Provisional Patent Application No. 18 / 148,248, filed December 29, 2022, the entire contents of which are incorporated herein by reference. [Background technology]
[0002] This application relates generally to data storage devices, and more particularly, but not exclusively, to data processing and cache management.
[0003] For example, scientific instruments such as imaging instruments and spectrometers typically include complex arrangements of components, sensors, detectors, input and output ports, energy sources, and consumable elements. Some such scientific instruments generate relatively large amounts of data during operation, which can impact memory requirements and efficient data access. [Brief explanation of the drawings]
[0004]
[0013] The embodiments will be readily understood by the following detailed description taken in conjunction with the accompanying drawings, in which:
[0014] To facilitate this description, like reference numerals refer to like structural elements;
[0015] The embodiments are illustrated in the figures of the accompanying drawings, by way of example, and not by way of limitation. [Figure 1] FIG. 1 is a block diagram of an exemplary scientific instrument support module for performing support operations, according to various embodiments. [Figure 2] 2 is a block diagram illustrating exemplary data structures (objects) generated via the scientific instrument support module of FIG. 1 by applying automated processing to raw instrument data files, according to various embodiments. [Figure 3] FIG. 3 is a block diagram showing a specific example of an object shown in FIG. 2. [Figure 4] 2 is a flowchart of a data distribution method applied via the scientific instrument support module of FIG. 1 according to various embodiments. [Figure 5] 1 is a block diagram illustrating communication and data flow between various components of a distributed computing system, according to various embodiments. [Figure 6] FIG. 1 is a block diagram illustrating an example computing device that may implement some or all of the scientific instrumentation support methods and / or functionality disclosed herein, according to various embodiments. [Figure 7] FIG. 1 is a block diagram illustrating an exemplary scientific instrument support system capable of implementing some or all of the scientific instrument support methods and / or functionality disclosed herein, according to various embodiments. [Figure 8] FIG. 2 is a block diagram illustrating a cloud-hosted deployment used in the scientific instrument support module of FIG. 1 in accordance with various embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0005] Disclosed herein are scientific instrument support systems, as well as related methods, computing devices, storage devices, and computer-readable media. For example, in some embodiments, a support apparatus for a scientific instrument includes first logic, second logic, and third logic. The first logic is configured to acquire a first data file via one or more detectors of the scientific instrument, the first data file including unseparated data from multiple scans or channels of the one or more detectors. The second logic is configured to apply automated processing to the first data file to generate a corresponding plurality of second data files based on one or more configuration settings, each of the second data files being smaller than the first data file and named using a file naming convention that references the content of each of different ones of the second data files. The plurality of second data files are stored in object storage. The third logic is configured to process data requests received from a client device, the data request being for a data portion of the first data file. The third logic is further configured to provide the data portion back to the client device by accessing one or more of the first memory cache, the second memory cache, and the object storage to retrieve the corresponding portion of the corresponding one of the second data files identified based on the file naming convention, wherein the first memory cache and the second memory cache have different respective eviction policies for data loaded therein from the object storage.
[0006] The scientific instrument support embodiments disclosed herein may achieve improved performance compared to conventional approaches, for example, for time-of-flight (TOF) mass spectrometers, quadrupole mass spectrometers, ion trap mass spectrometers, instruments including multiple spectrometers and / or detectors, such as ultraviolet (UV) detectors, diode array detectors (DADs, e.g., photodiode array detectors, PDAs), and chromatography system detectors. For purposes of illustration, and without any implied limitation, some exemplary embodiments are described below with reference to TOF mass spectrometers. From the description provided, one skilled in the art will be able to fabricate and use additional support embodiments for scientific instruments using other types of spectrometers, detectors, devices, and various combinations thereof, without undue experimentation.
[0007] TOF mass spectrometry (TOFMS) is a mass spectrometry method in which the mass-to-charge ratio of ions is determined by time-of-flight measurements. Ions are accelerated through an electric field of known strength. This acceleration results in ions with the same kinetic energy as any other ion with the same charge. The velocity of the ions depends on their mass-to-charge ratio, such that heavier ions of the same charge achieve lower velocities than lighter ions. The time it takes the ions to reach a downstream detector is then measured. This time depends on the ion's velocity and therefore provides a measure of the ion's mass-to-charge ratio. From the measured TOF mass spectrum, the mass-to-charge ratios of its various components, and other known experimental parameters, the composition of the analyte can usually be determined.
[0008] A TOF mass analyzer typically includes a mass analyzer and a detector. An ion source (pulsed or continuous) is used to generate ions from the analyte. The TOF mass analyzer can be a linear flight tube or a reflectron. In various examples, the ion detector is a microchannel plate (MCP) detector or a secondary emission multiplier (SEM). The electrical signal from the detector is digitized by a time-to-digital converter (TDC) or an analog-to-digital converter (ADC). A TDC is a counting detector, and ion counting performed by it typically involves summing a large number (e.g., hundreds) of individual mass spectra, sometimes referred to as histogramming. Corresponding TOF mass analyzers typically operate at repetition rates of 5 kHz to 20 kHz to generate a sufficiently large number of mass spectra to be summed. The ADC typically operates at a rate of approximately 10 gigasamples / second, digitizing the pulsed ion current from the MCP detector at discrete time intervals. In various examples, the ADC has a dynamic range of 8 to 12 bits. The use of an ADC (as opposed to a TDC) is more beneficial for some specific types of TOF mass spectrometers, such as matrix-assisted laser desorption / ionization (MALDI)-TOF instruments, which have relatively high peak currents.
[0009] The raw data detected by a mass spectrometer is typically in the form of a signal distributed across the various m / z (mass-to-charge) values at which ions are detected. Centroid data includes raw data that has been processed via a suitable algorithm to retain only the local maxima in each mass range at which ions are detected. Such centroid data is often referred to as a "centroid scan."
[0010] In some instances, raw data files generated by TOFMS instruments range in size from 1 GB to 100 GB and contain data from multiple scans and / or channels in a non-segregated binary format, in a file format specifically optimized for recording, transferring, and packaging experimental data corresponding to a single analyte injection. In typical data storage systems, such binary data files take a relatively long time to transfer from object storage to a file system that enables a corresponding file reader associated with the instrument (e.g., provided by the instrument manufacturer) to extract selected scan and / or channel data for further processing, e.g., in response to a request from a client device. The corresponding data access delays typically worsen nonlinearly for larger data files and can present significant obstacles to users and operators of the corresponding scientific instrument and / or system.
[0011] In some use cases, an important consideration is the cost of data storage. For example, service fees associated with object storage are typically lower (e.g., in some specific cases, by a factor of about five) than service fees associated with general-purpose data storage. Given the fact that in typical real-world applications of various scientific instruments, raw data are accessed relatively infrequently, usually in a selective manner after their initial acquisition, object storage is often considered the preferred option.
[0012] Furthermore, the inventors have recognized that in some instances, significant performance improvements can be achieved by leveraging object storage and adapting the data structures used therefor to the way the corresponding data will be used after initial acquisition. In many use cases, such data structures are functionally distinct from data structures that facilitate high-speed recording during data acquisition. For example, the inventors have recognized that generating extracted ion chromatograms (XICs) can be approximately an order of magnitude faster using columnar storage than when the original raw data format is used for the same purpose.
[0013] In a network environment, various cache memory systems may be used to speed up data access in response to a data request from a client device. However, given the size and structure of typical raw data files, cache memory systems may not be able to provide sufficiently fast data access to client devices when dealing with large data files.
[0014] The above-mentioned problems in the state of the art, and possibly several other related problems, can be beneficially addressed using various examples, aspects, features, and embodiments of the systems and methods for data processing and cache management disclosed herein. In representative examples, a cache management system is designed and configured to consider various factors related to the type of measurement performed by an instrument and / or the data analysis criteria specified in a data access request received from a client device. In some examples, the cache management system operates in a cloud or enterprise environment to provide fast data access and data processing based on mass spectrometry criteria. In various examples, such mass spectrometry criteria include, but are not limited to, scan, scan type, instrument type, scan fragments corresponding to a specified mass range around a requested mass, and the like. At least some embodiments beneficially reduce data access latency relative to typical latency associated with previous data access solutions. At least some embodiments support horizontal scaling and beneficially mitigate latency variations associated with changes in the workload of a data delivery service.
[0015] Thus, the embodiments disclosed herein provide improvements to scientific instrument technology (e.g., improvements to the computer technology supporting such scientific instruments, among other things). For example, various embodiments disclosed herein may achieve improved (e.g., faster and more selective) access to scientific instrument data compared to conventional approaches. In various examples, the improvements are directed toward achieving one or more of the following goals: (a) the ability to store petabytes of raw data in inexpensive cloud storage with relatively fast selective access to any selected portion of the data; (b) the ability to scale caching and other data operations across multiple servers; (c) support for parallel (e.g., substantially simultaneous) access to data from many (e.g., up to 1000) client devices; (d) multi-stage caching to improve query performance; (e) fast traversal over time-series data using columnar file formats; (f) support for Representational State Transfer (REST) Application Programming Interfaces (APIs), unary general-purpose Remote Procedure Calls (gRPCs), and streaming calls for better intra-cluster performance; (g) compatibility with different types of object storage; (h) portability between on-premise and cloud deployments through relatively easy configuration changes; (i) flexibility in data access modes using, for example, FaaS capabilities, various programming languages, “big data” solutions, etc. As used herein, the acronym "FaaS" stands for Function as a Service, a category of cloud computing services that enable customers to develop, run, and manage application functionality without the complexities of building and maintaining the infrastructure typically associated with developing and launching an app.
[0016] In some embodiments, a data preservation method is provided that includes: (i) converting a large raw instrument data file into a corresponding plurality of smaller data files more suitable for storing time-series data; (ii) generating a common metadata file for a plurality of such smaller data files, the metadata detailing relevant parameters of the conversion process; (iii) assembling data portions from different regions of the raw instrument data file into records; (iv) separating different streams stored in the raw instrument data file into one or more streams of records; (v) further separating the streams into groups of, for example, a selected fixed number (e.g., 100) of centroid scans per stream; and (vi) using a suitable object storage naming convention (e.g., conceptually similar to a file path) to quickly find and then transfer required groups of streams from object storage to a cache memory to reduce the frequency of repeated transfers. In various examples, cache management policies governing cache memory operation provide for the transfer of only relatively small portions of data objects sufficient to resolve pending data requests. Exemplary benefits of such cache management policies include (i) a significant reduction in the amount of data stored in more expensive cache memory and (ii) faster data access due to smaller data portions transferred to cache memory in the event of a cache miss. In some examples, raw data REST and gRPC API services are designed to scale horizontally as demand increases, and cache tiers are shared across multiple compute instances. In such examples, synchronization of the transfer of data from object storage to cache memory is achieved using a distributed cache. The distributed cache allows one instance to initiate a copy to the cache and other instances to wait for the transfer to complete when they need to access the same cached data.
[0017] Various of the embodiments disclosed herein may improve upon conventional approaches to achieve the technical benefits of improved data behavior implemented with optimized cache management policies through file conversion and data segregation. Such technical benefits are routine and not achievable with conventional approaches, and all users of systems incorporating such embodiments can benefit from them (e.g., by helping users accelerate technical tasks such as processing and analyzing experimental data). Thus, the technical features of the embodiments disclosed herein, as well as various combinations of the features disclosed herein, are clearly unconventional in the field of instrument-associated data storage. As further described herein, various aspects of the embodiments disclosed herein can improve the functionality of the computer itself, for example, by operating instrument-associated data storage in an optimized manner, resulting in higher levels of productivity. The computational features disclosed herein not only involve the collection and comparison of information, but also apply new analytical and technical tools to modify the behavior of instrument-associated storage. Thus, the present disclosure introduces capabilities that neither conventional computing devices nor humans have been able to perform.
[0018] Thus, embodiments of the present disclosure may serve any of a number of technical purposes, such as controlling a particular technical system or process, determining how to control or configure a machine from measurements, or increasing the throughput of a data pipeline. Some examples disclosed herein provide solutions to technical problems, including, but not limited to, improvements to TOFMS instruments, such as improvements in the computer technology supporting TOFMS instruments, among other improvements.
[0019] In the following detailed description, reference is made to the accompanying drawings that form a part hereof, where like numerals refer to like parts throughout and which show, by way of illustration, embodiments that may be practiced. It is to be understood that other embodiments may be utilized and structural or logical changes may be made without departing from the scope of the present disclosure. Therefore, the following detailed description is not to be taken in a limiting sense.
[0020] Various operations may be described as multiple separate actions or operations, in the order most helpful in understanding the subject matter disclosed herein. However, the order of description should not be construed as implying that these operations are necessarily order dependent. In particular, these operations may not be performed in the order presented. The operations described may be performed in a different order than in the described embodiment. Various additional operations may be performed and / or described operations may be omitted in additional embodiments.
[0021] For purposes of this disclosure, the phrases “A and / or B” and “A or B” mean (A), (B), or (A and B). For purposes of this disclosure, the phrases “A, B, and / or C” and “A, B, or C” mean (A), (B), (C), (A and B), (A and C), (B and C), or (A, B, and C). Although some elements may be referred to in the singular (e.g., “processing device”), any suitable element may be represented by multiple instances of that element, and vice versa. For example, a set of operations described as being performed by a processing device may be implemented with different ones of the operations performed by different processing devices.
[0022] This description uses the phrases "one embodiment," "various embodiments," and "some embodiments," each of which may refer to one or more of the same or different embodiments. Furthermore, terms such as "comprising," "including," and "having," when used with respect to embodiments of the present disclosure, are synonymous. When used to describe a range of values, the phrase "between X and Y" represents a range that includes X and Y. As used herein, "apparatus" may refer to any individual device, a collection of devices, a portion of a device, or a collection of portions of devices. The drawings are not necessarily to scale.
[0023] FIG. 1 is a block diagram illustrating a scientific instrument support module 1000 for performing support operations, according to various embodiments. The scientific instrument support module 1000 may be implemented by circuitry (e.g., including electrical and / or optical components) such as a programmed computing device. The logic of the scientific instrument support module 1000 may be contained on a single common computing device or may be distributed across multiple computing devices that communicate with each other as needed. An example of a computing device that may implement the scientific instrument support module 1000, alone or in combination, is discussed herein with reference to the computing device 6000 of FIG. 6 , and an example of a system of interconnected computing devices in which the scientific instrument support module 1000 may be implemented across one or more of the computing devices is discussed herein with reference to the scientific instrument support system 7000 of FIG. 7 .
[0024] As shown in FIG. 1 , scientific instrument support module 1000 includes first logic 1002, second logic 1004, and third logic 1006 for implementing the support methods described herein for a scientific instrument, such as a TOFMS instrument. As used herein, the term “logic” may include an apparatus configured to perform a set of operations associated with the logic. For example, any of the logic elements included in scientific instrument support module 1000 may be implemented by one or more computing devices programmed with instructions that cause one or more processing devices of the computing devices to perform a set of associated operations. In particular embodiments, a logic element may include one or more non-transitory computer-readable media having instructions that, when executed by one or more processing devices of the one or more computing devices, cause the one or more computing devices to perform the associated set of operations. As used herein, the term “module” may refer to a collection of one or more logic elements that together perform a function associated with the module. Different logic elements within a module may take the same form or different forms. For example, some logic in a module may be implemented by a programmed general-purpose processing device, while other logic in the module may be implemented by an application-specific integrated circuit (ASIC). In another example, different ones of the logic elements in a module may be associated with different sets of instructions executed by one or more processing devices. A module may not include all of the logic elements depicted in an associated figure; for example, a module may include a subset of the logic elements depicted in an associated figure when the module performs a subset of the operations discussed herein with reference to that module.
[0025] The first logic 1002 may acquire one or more raw instrument data files corresponding to one or more analytes via one or more detectors of the TOFMS instrument. As described above, to perform such acquisition, ions generated via the ion source of the TOFMS instrument are directed through a mass analyzer and detected via one or more detectors of the TOFMS instrument. Thus, the first logic 1002 may acquire data via the TOFMS instrument at one or more settings of the mass analyzer and detector. The first logic 1002 may then include one or more parameters associated with the settings in the data files.
[0026] The second logic 1004 may apply automated processing to one or more raw instrument data files obtained via the first logic 1002. In various examples, the automated processing includes one or more of: (i) converting each large raw instrument data file into a corresponding plurality of smaller data files; (ii) generating a metadata file for the plurality of smaller data files based on one or more parameters related to TOFMS instrument settings included in the raw data by the first logic 1002 and further based on parameters related to the conversion of the large raw data file into the plurality of smaller data files; and (iii) naming the plurality of individual files using a naming convention suitable for referencing their contents. In some examples, the converting processing step includes some or all of the following substeps: (a) assembling different data portions of the raw data file into records; (b) separating different data streams of the raw data file into one or more record streams; and (c) further separating the streams into groups.
[0027] The third logic 1006 may process data requests received from various client devices. Therefore, based on the data request, the third logic 1006 may perform the following operations: First, the third logic 1006 may map the received data request to a corresponding data file among the multiple smaller data files generated via the second logic 1004, for example, based on a naming convention. Once the name of the smaller data file is identified through the mapping, the third logic 1006 may check a least recently used (LRU) cache in memory for the identified data file. If the LRU cache has the identified data file, the third logic 1006 retrieves the file from there, filters it to obtain the requested scan data, and returns it to the requesting client device. If the LRU cache does not have the identified data file, the third logic 1006 initiates a lookup for the file in a local file cache. If the file is found in the local file cache, the third logic 1006 causes the necessary data to be read from there and also causes a copy of the data file to be loaded into the LRU cache. If the file is not found in the local file cache, the third logic 1006 may query a corresponding database (e.g., Redis) to see if another service instance has the corresponding object. If another service instance has a copy of the object, the third logic 1006 may request the object copy from that instance, for example, using a fast binary serialization protocol for requests. If the service instance does not have the object, the third logic 1006 may request a copy of the object from object storage in the local file system and update Redis to indicate that a local copy of the object is current.The third logic 1006 can then cause the local file copy to be read, added to an in-memory LRU cache, and directed to the requesting client device after appropriate data filtering as described above.
[0028] As used herein, the term "Redis" refers to a Remote Dictionary Server. Redis is an advanced key-value store that can function as a Not only Structured Query Language (NoSQL) database or as a memory cache store to improve performance when serving data stored in system memory. Redis supports a variety of data structures such as strings, hashes, sets, lists, sorted sets, bitmaps, and geospatial indexes with radius queries, among others. In various additional examples, other distributed in-memory data stores (other than Redis) can also be used.
[0029] FIG. 2 is a block diagram illustrating a data structure (object) 2000 that may be generated via second logic 1004 by applying automated processing to a raw instrument data file, according to one embodiment. In the illustrated example, object 2000 is a column-oriented file divided into row groups. For illustrative purposes and without any implied limitation, object 2000 is shown in FIG. 2 as having two row groups, labeled row group 0 and row group 1, respectively. In other examples, the corresponding object 2000 may have a different number (from two) of row groups (see, e.g., FIG. 3). Each row group has columns, and each column is an array of data of the same data type. For illustrative purposes and without any implied limitation, row group 0 and row group 1 are each shown as having five respective columns, labeled “Column A” through “Column E,” respectively. In a representative example, the columns of a row group may be independently read from disk to other memory.
[0030] During operation, the scientific instrument support module 1000 may store multiple objects 2000 in object storage. When necessary, the scientific instrument support module 1000 may copy specific objects 2000 required to access the data stored therein to the service instance's local file cache. The scientific instrument support module 1000 may also cause individual columns of each row group to be individually read from the local file cache and cached in an in-memory LRU cache for faster access. In a representative example, the in-memory LRU cache is smaller than the local file cache, with exemplary sizes of approximately 1 Gb and 100 Gb, respectively. The LRU cache and the local file cache each have different eviction policies. In one example, the LRU cache is configured to evict the least recently accessed data when the LRU cache reaches its memory limit and additional data needs to be added to the cache. In contrast, the local file cache includes a local directory structure, which may be cleaned up asynchronously based on the total number of files cached therein and / or the total disk space used. For example, if individual files have metadata such as a "created" timestamp and / or a "last accessed" timestamp, files may be evicted from the local file cache based on age, e.g., the oldest files being deleted first, or based on time elapsed since last access, e.g., the least recently accessed files being deleted first.
[0031] FIG. 3 is a block diagram illustrating an example of an object 2000. In the illustrated example, the object 2000 has 100 row groups, labeled "RG1" through "RG100." Each row group RGn has six columns, labeled "3001" through "3006." The data types of columns 3001-3006 are indicated in a header row 3010. In this example, the data types are as follows: Centroid Scan Number, Mass, Peak Intensity, Baseline, Noise, and Resolution. The number of rows within a row group RGn can vary from one row group to another. In different examples, the object 2000 can have a different number of respective row groups RGn or the same fixed number of respective row groups RGn. In some examples, to optimize the scan traverse for a particular analyte and / or type of data analysis, the range of masses covered by the row groups can be reduced and the number of row groups per object 2000 can be increased.
[0032] In a typical example, a row group RGn in a columnar storage system is a group of arrays of manageable size. In some cases, a row group RGn may have columns such as scan number, mass, and intensity, each of which may have a large number of entries (rows) (e.g., 100,000). However, some smaller secondary files may have only one row group, with each row containing a centroid peak, such that the entire secondary file contains centroid peaks spanning 100 scans. Each column within each row group RGn is individually compressed to make the row group more suitable for functioning as a block of data that can be decompressed and loaded into memory. In some examples, there is no direct correspondence between centroid scans and row groups. In one such example, an optimized row columnar (ORC) file contains 100 centroid scans, but limits each row group RGn to 10 centroid scans, with each row within the row group representing a centroid peak.
[0033] In one embodiment, the LRU cache key is assembled from context information. An example of such context information is provided by the following string: {Injection Identifier} / {Columnar Group} / {Row Group} / {Column Identifier} (1)
[0034] The naming convention may be such that if the original raw data file is named "03_lumos_prg_sa_r1.raw", then the second logic 1004 may be configured to use this file name as the injection identifier in the above string. If the values are for centroid scan numbers 1 through 100 and the mass values for row group 0 are cached, then the corresponding LRU key may be in the form of the following string: 03_lumos_prg_sa_r1 / centroid_00001-00100 / 0 / mass (2)
[0035] The corresponding intensities of these masses can be accessed using another LRU key, represented for example by the following string: 03_lumos_prg_sa_r1 / centroid_00001-00100 / 0 / intensity (3)
[0036] When a value is retrieved from the LRU cache, it may be in the form of an array of double-precision floating-point numbers, for example, because this particular data type is designated for the corresponding column in the columnar storage format. In other embodiments, other suitable naming conventions and data types and formats may also be used.
[0037] 4 is a flowchart of a data distribution method 4000 according to one embodiment. In one example, the method 4000 is implemented using the scientific instrument support module 1000. The method 4000 is described below with continued reference to FIGS. 1-3.
[0038] The method 4000 includes the scientific instrument support module 1000 receiving a request for data from a client device (block 4002). In a typical example, the request identifies specific data that needs to be returned to the client device in response to the request. Such identification can be implemented using applicable naming conventions, illustrative examples of which are described above. For example, the request received at block 4002 may specify a centroid scan for injection "03_lumos_prg_sa_r1," scan number 42. Accordingly, the corresponding object 2000 may be the example object 2000 of FIG. 3.
[0039] The method 4000 further includes the scientific instrument support module 1000 checking an in-memory LRU cache for the requested data (block 4004). Such checking in block 4004 may include generating a corresponding LRU cache lookup key, which may be in the format indicated by strings (2) and (3) above. The method 4000 further includes determining whether the LRU cache contains the (smaller) data file identified by the key (decision block 4006). If the LRU cache contains the data file ("yes" at decision block 4006), processing of the method 4000 is directed to block 4020. Otherwise ("no" at decision block 4006), processing of the method 4000 is directed to block 4008.
[0040] The method 4000 also includes the scientific instrument support module 1000 checking the local file cache for the corresponding data file (block 4008). Such checking in block 4008 may include using applicable file naming conventions to search the directory structure of the local file cache for the corresponding file name. The method 4000 further includes determining whether the local file cache has the file identified by the file name (decision block 4010). If the local file cache has the file ("Yes" at decision block 4010), processing of the method 4000 is directed to block 4011. Otherwise ("No" at decision block 4010), processing of the method 4000 is directed to block 4012. The method 4000 further includes the scientific instrument support module 1000 causing the local file copy to be read and added to the LRU cache (block 4011).
[0041] The method 4000 includes the scientific instrument support module 1000 querying Redis to check whether another service instance has a copy of the corresponding file (block 4012). The method 4000 further includes the scientific instrument support module 1000 receiving a response from Redis and determining, based on the received response, whether another service instance has such a copy (decision block 4014). If the other service instance has the file ("Yes" at decision block 4014), processing of the method 4000 is directed to block 4018. Otherwise ("No" at decision block 4014), processing of the method 4000 is directed to block 4016.
[0042] The method 4000 further includes the scientific instrument support module 1000 requesting a copy of the file (object) from object storage and receiving the requested copy (block 4016). The method 4000 further includes the scientific instrument support module 1000 updating Redis to indicate that a local copy of the object is current (block 4016). The method 4000 further includes the scientific instrument support module 1000 adding the received file copy to the LRU and local file caches (block 4018).
[0043] The method 4000 further includes the scientific instrument support module 1000 applying data filtering to the corresponding files (block 4020). The applied filtering is in accordance with the request received at block 4002. The method 4000 further includes the scientific instrument support module 1000 directing the resulting filtered data to the client device (block 4020).
[0044] 5 is a block diagram illustrating communication and data flow between various components of a distributed computing system 5000 according to one embodiment. In at least some examples, communication and data flow in the system 5000 may follow the method 4000. In some examples, the scientific instrument support module 1000 may be in operative communication with the system 5000. In some other examples, some portions of the scientific instrument support module 1000 may be implemented within the system 5000.
[0045] In the depicted example, system 5000 includes API server instances 5004 and 5012, a local LRU memory cache 5006, a local disk cache 5008, a Redis in-memory data store 5010, and object storage 5014. The various transmissions within system 5000 shown in Figure 5 are triggered by a request 5022 for data received by API server instance 5004 from client device 5002. Request 5022 may follow the operations of block 4002 of method 4000.
[0046] In response to request 5022, API server instance 5004 sends a corresponding query 5024 to local LRU memory cache 5006. Query 5024 may follow the operations of block 4004 of method 4000. Local LRU memory cache 5006 sends a response 5026 to query 5024, which either returns the requested data or notifies API server instance 5004 that the requested data is not found. Using the latter result of query 5024, API server instance 5004 sends a corresponding query 5028 to local disk cache 5008. Query 5028 may follow the operations of block 4008 of method 4000. Local disk cache 5008 sends a response 5030 to query 5028, which either returns the requested data or notifies API server instance 5004 that the requested data is not found. Using the latter results of query 5028, API server instance 5004 sends a subsequent corresponding query 5032 to Redis in-memory data store 5010. Query 5032 may follow the operations of block 4012 of method 4000. Redis in-memory data store 5010 sends a response 5034 to query 5032, which either identifies other API server instances that have the requested data or notifies API server instance 5004 that no such server instances exist.
[0047] In one example result of query 5032, response 5034 identifies API server instance 5012. Using such result, API server instance 5004 sends request 5036 for corresponding data to API server instance 5012. In response to request 5036, API server instance 5012 returns response 5038 with the requested data.
[0048] In another example result of query 5032, response 5034 informs API server instance 5004 that no server instances have the corresponding data. With such a result, API server instance 5004 operates to access object storage 5014 via command 5040 and receive a corresponding data stream 5042 having a corresponding data file (e.g., object 2000). Command 5040 and data read 5042 may follow the operations of block 4016 of method 4000. Once the object is received, API server instance 5004 stores the fetched object and / or its relevant portions in local LRU memory cache 5006 and local disk cache 5008. Corresponding data writes 5044, 5046 may follow the associated operations of blocks 4011, 4018. The API server instance 5004 further operates to send updates 5048 to the Redis in-memory data store 5010, which provides the appropriate information about the fetched object to the database therein.
[0049] As noted above, the scientific instrument support module 1000 can be implemented by one or more computing devices. Figure 6 is a block diagram of a computing device 6000 capable of implementing some or all of the scientific instrument support functions and / or methods disclosed herein, according to various embodiments. In some embodiments, the scientific instrument support module 1000 can be implemented by a single computing device 6000 or by multiple computing devices 6000. Furthermore, as discussed below, the computing device 6000 (or multiple computing devices 6000) implementing the scientific instrument support module 1000 can be part of one or more of the scientific instrument 7010, user local computing device 7020, service local computing device 7030, or remote computing device 7040 of Figure 7.
[0050] 6 is shown as having several components, any one or more of which may be omitted or duplicated as appropriate for the application and configuration. In some embodiments, some or all of the components included in computing device 6000 may be mounted on one or more motherboards and enclosed in a housing (e.g., comprising plastic, metal, and / or other materials). In some embodiments, several of these components may be fabricated on a single system-on-a-chip (SoC) (e.g., an SoC may include one or more processing devices 6002 and one or more storage devices 6004). 6, but may include interface circuitry (not expressly shown) for coupling to one or more components using any suitable interface (e.g., a Universal Serial Bus (USB) interface, a High-Definition Multimedia Interface (HDMI) interface, a Controller Area Network (CAN) interface, a Serial Peripheral Interface (SPI) interface, an Ethernet interface, a wireless interface, or any other suitable interface). For example, computing device 6000 may not include display device 6010, but may include display device interface circuitry (e.g., connectors and driver circuits) to which display device 6010 may be coupled.
[0051] The computing device 6000 may include a processing device 6002 (e.g., one or more processing devices). As used herein, the term "processing device" may refer to any device or portion of a device that processes electronic data from registers and / or memory and converts the electronic data into other electronic data that may be stored in registers and / or memory. The processing device 6002 may include one or more digital signal processors (DSPs), application specific integrated circuits (ASICs), central processing units (CPUs), graphics processing units (GSMUs), or other processors. The processing device may include a graphics processing unit (GPU), a cryptographic processor (a dedicated processor that executes cryptographic algorithms in hardware), a server processor, or any other suitable processing device.
[0052] The computing device 6000 may include a storage device 6004 (e.g., one or more storage devices). The storage device 6004 may be a random access memory (RAM) device (e.g., static RAM). The storage device 6004 may include one or more memory devices, such as a static RAM (SRAM) device, a magnetic RAM (MRAM) device, a dynamic RAM (DRAM) device, a resistive RAM (RRAM) device, or a conductive-bridging RAM (CBRAM) device, a hard drive-type memory device, a solid-state memory device, a networked drive, a cloud drive, or any combination of memory devices. In some embodiments, the storage device 6004 may include memory that shares a die with the processing device 6002. In such embodiments, the memory may be used as cache memory and may include, for example, embedded dynamic random access memory (eDRAM) or spin transfer torque magnetic random access memory (STT-MRAM). In some embodiments, the storage device 6004 may include a non-transitory computer-readable medium having instructions, which, when executed by one or more processing devices (e.g., the processing device 6002), cause the computing device 6000 to perform any suitable method or any suitable portion of the methods disclosed herein.
[0053] The computing device 6000 may include an interface device 6006 (e.g., one or more interface devices 6006). The interface device 6006 may include one or more communication chips, connectors, and / or other hardware and software for managing communications between the computing device 6000 and other computing devices. For example, the interface device 6006 may include circuitry for managing wireless communications for data transfer to and from the computing device 6000. The term "wireless" and its derivatives may be used to describe circuits, devices, systems, methods, techniques, communication channels, etc. that may communicate data through the use of modulated electromagnetic radiation over a non-solid medium. This term does not imply that the associated devices do not include any wiring, although in some embodiments they may not. The circuitry included in the interface device 7006 for managing wireless communications may implement any of several wireless standards or protocols, including, but not limited to, Wi-Fi (IEEE 802.11 family), the IEEE 802.16 standard (e.g., the IEEE 802.16-2005 amendment), Institute for Electrical and Electronic Engineers (IEEE) standards including the Long-Term Evolution (LTE) project with any amendments, updates, and / or revisions (e.g., the Advanced LTE project, the Ultra Mobile Broadband (UMB) project (also referred to as "3GPP2"), etc.).In some embodiments, the circuitry included in the interface device 7006 for managing wireless communications is compatible with Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Universal Mobile Telecommunications System (UMTS), High Speed Packet Access (HSPA), Evolved HSPA, and the like. HSPA, E-HSPA, or LTE networks. In some embodiments, the circuitry included in interface device 7006 for managing wireless communications may operate in accordance with Enhanced Data for GSM Evolution (EDGE), GSM EDGE Radio Access Network (GERAN), Universal Terrestrial Radio Access Network (UTRAN), or Evolved UTRAN (E-UTRAN). In some embodiments, the circuitry included in interface device 7006 for managing wireless communications may operate in accordance with Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Digital Enhanced Cordless Telecommunications (DECT), Evolution-Data Optimized (EV-DO), and derivatives thereof, as well as any other wireless protocols designated as 3G, 4G, 5G, and beyond. In some embodiments, the interface device 7006 may include one or more antennas (eg, one or more antenna arrays) for receiving and / or transmitting wireless communications.
[0054] In some embodiments, the interface device 6006 may include circuitry for managing wired communications, such as electrical, optical, or any other suitable communications protocol. For example, the interface device 6006 may include circuitry supporting communications according to Ethernet technology. In some embodiments, the interface device 6006 may support both wireless and wired communications and / or multiple wired and / or wireless communications protocols. For example, a first set of circuits in the interface device 6006 may be dedicated to short-range wireless communications, such as Wi-Fi or Bluetooth, while a second set of circuits in the interface device 6006 may be dedicated to long-range wireless communications, such as global positioning system (GPS), EDGE, GPRS, CDMA, WiMAX, LTE, EV-DO, or others. In some embodiments, the first set of circuits in the interface device 6006 may be dedicated to wireless communications, while the second set of circuits in the interface device 6006 may be dedicated to wired communications.
[0055] Computing device 6000 may include battery / power circuitry 6008. Battery / power circuitry 6008 may include one or more energy storage devices (e.g., batteries or capacitors) and / or circuitry for coupling components of computing device 6000 to an energy source separate from computing device 6000 (e.g., AC line power).
[0056] The computing device 6000 may include a display device 6010 (e.g., multiple display devices). The display device 6010 may include any visual indicator, such as a heads-up display, a computer monitor, a projector, a touchscreen display, a liquid crystal display (LCD), a light-emitting diode display, or a flat panel display.
[0057] The computing device 6000 may include other input / output (I / O) devices 6012. The other I / O devices 6012 may include, for example, one or more audio output devices (e.g., speakers, headsets, earphones, alarms, etc.), one or more audio input devices (e.g., microphones or microphone arrays), a location device (e.g., a GPS device that communicates with a satellite-based system to receive the location of the computing device 6000, as known in the art), an audio codec, a video codec, a printer, sensors (e.g., thermocouples or other temperature sensors, humidity sensors, pressure sensors, vibration sensors, accelerometers, gyroscopes, etc.), an image capture device such as a camera, a keyboard, a cursor control device (e.g., a mouse, stylus, trackball, or touchpad, etc.), a barcode reader, a Quick Response (QR) code reader, or a radio frequency identification (RFID) sensor. The device may include a radio frequency identification (RFID) reader.
[0058] The computing device 6000 may have any form factor suitable for its application and configuration, such as a handheld or mobile computing device (e.g., a mobile phone, smartphone, mobile internet device, tablet computer, laptop computer, netbook computer, ultrabook computer, personal digital assistant (PDA), ultra-mobile personal computer, etc.), a desktop computing device, or a server computing device or other networked computing component.
[0059] In some examples, computing device 6000 is implemented using multiple pods in a Kubernetes cluster. A typical Kubernetes cluster includes multiple computer nodes that can be configured to host multiple pods, each of which functions as a virtual machine. In various deployments, several instances of a microservice may run on a single pod or on multiple pods (see also FIG. 8). In some examples, better performance is achieved when several of the pods are distributed across different computer nodes.
[0060] One or more computing devices implementing any of the scientific instrument support modules or methods disclosed herein may be part of a scientific instrument support system. Figure 7 is a block diagram of an exemplary scientific instrument support system 7000 in which some or all of the scientific instrument support methods disclosed herein may be implemented, according to various embodiments. The scientific instrument support modules and methods disclosed herein (e.g., scientific instrument support module 1000 of Figure 1 and method 4000 of Figure 4) may be implemented by one or more of the scientific instrument 7010, user local computing device 7020, service local computing device 7030, and remote computing device 7040 of the scientific instrument support system 7000.
[0061] Any of the scientific instrument 7010, the user local computing device 7020, the service local computing device 7030, or the remote computing device 7040 may include any of the embodiments of computing device 6000 discussed herein with reference to FIG. 6, and any of the scientific instrument 7010, the user local computing device 7020, the service local computing device 7030, or the remote computing device 7040 may take the form of any suitable embodiment of the embodiments of computing device 6000 discussed herein with reference to FIG. 6.
[0062] The scientific instrument 7010, the user local computing device 7020, the service local computing device 7030, and / or the remote computing device 7040 may each include a respective processing device 6002, a respective storage device 6004, and a respective interface device 6006. The processing device 6002 may take any suitable form, including any of the forms of the processing device 6002 described herein with reference to Figure 6, and the processing devices 6002 included in different ones of the scientific instrument 7010, the user local computing device 7020, the service local computing device 7030, or the remote computing device 7040 may take the same form or different forms. The storage device 6004 may take any suitable form, including any of the forms of the storage devices 6004 discussed herein with reference to Figure 6, and the storage devices 6004 included in different ones of the scientific instrument 7010, the user local computing device 7020, the service local computing device 7030, or the remote computing device 7040 may take the same form or different forms. The interface device 6006 may take any suitable form, including any of the forms of the interface device 6006 discussed herein with reference to FIG. 7, and the interface devices 6006 included in different ones of the scientific instrument 7010, the user local computing device 7020, the service local computing device 7030, or the remote computing device 7040 may take the same or different forms.
[0063] The scientific instrument 7010, the user local computing device 7020, the service local computing device 7030, and the remote computing device 7040 may communicate with other elements of the scientific instrument support system 7000 via communication paths 7008. The communication paths 7008 may communicatively couple the interface devices 6006 of different ones of the elements of the scientific instrument support system 7000, as shown, and may be wired or wireless communication paths (e.g., according to any of the communication techniques discussed herein with reference to the interface device 6006 of the computing device 6000 of FIG. 6 ). While the particular scientific instrument support system 7000 depicted in FIG. 7 includes communication paths between each pair of the scientific instrument 7010, the user local computing device 7020, the service local computing device 7030, and the remote computing device 7040, this “fully connected” implementation is merely illustrative, and in various embodiments, various ones of the communication paths 7008 may not be present. For example, in some embodiments, the service local computing device 7030 may not have a direct communication path 7008 between its interface device 6006 and the interface device 6006 of the scientific instrument 7010, but may instead communicate with the scientific instrument 7010 via the communication path 7008 between the service local computing device 7030 and the user local computing device 7020, and the communication path 7008 between the user local computing device 7020 and the scientific instrument 7010. The scientific instrument 7010 may include any suitable scientific instrument, such as, for example, a TOFMS instrument.
[0064] 8 is a block diagram illustrating a cloud-hosted deployment 8000 for use with the scientific instrument support module 1000, according to one embodiment. The deployment 8000 is a K8s deployment with horizontal scaling and the use of local volumes. As used herein, K8s (also known as Kubernetes or "kube") refers to an open-source container orchestration platform that automates many of the manual processes involved in deploying, managing, and scaling containerized applications. Horizontal scaling occurs when the deployment 8000 deploys data pods 8010 in response to increased workloads. n This means increasing the number N of data pods, where n=1, 2, ..., N. This response will not affect the data pods 8010 already running for the workload. n This differs from vertical scaling, which typically involves allocating more resources (e.g., memory or CPU) to the deployed data pods 80101-8010. N When the number N of data pods exceeds a preset minimum, the number N is scaled back down. n Each has a configurable first-in-first-out (FIFO) in-memory cache.
[0065] The deployment 8000 executes a data service 8002 configured to process data requests, such as request 5022 (FIG. 5), received from various client devices, such as client device 5002 (FIG. 5). The deployment 8000 includes object storage 5014 (see also FIG. 5). The deployment 8000 further includes a Redis cluster 8020 used for, among other things, data bookkeeping, pod discovery, and publish / subscribe (pub / sub) event coordination. Pub / sub messaging is a form of asynchronous inter-service communication that can be used in serverless and microservices architectures. In the pub / sub model, any message published to a topic is received by all subscribers to the topic. The deployment 8000 uses pub / sub messaging to distribute change events from the corresponding database. These events can be used to build views of database state and state history for parallel processing and workflow.
[0066] During operation, chunks of the stream transferred from the object storage 5014 are stored in the corresponding data pods 8010. n Cached on local volumes of data pods 80101 to 8010 N The different ones query the Redis cluster 8020 and which data pod 8010 n If so, it may check whether it has the requested data. Such a query may follow the operations of block 4012 of method 4000. Query 5032 shown in FIG. 5 is an example of such a query. Redis cluster 8020 may check whether data pod 8010 has the requested data. n Identifying data pods 80101 to 8010 N In response to a particular response from the Redis cluster 8020, the querying data pod responds to the query 5032 by either informing the querying data pod that none of the identified data pods 8010 have the requested data. n5 ) via one or more lateral links 8012 from the object storage 5014 (see also elements 5036, 5038 in FIG. 5 ) or to retrieve data from the object storage 5014 (see also elements 5040, 5042 in FIG. 5 ).
[0067] According to exemplary embodiments disclosed above, for example with reference to any one or any combination of some or all of FIGS. 1-8 , there is provided a support apparatus for a scientific instrument, the support apparatus including: first logic configured to acquire a first data file via one or more detectors of the scientific instrument, the first data file including unseparated data from multiple scans or channels of the one or more detectors; and second logic configured to apply automated processing to the first data file to generate a corresponding plurality of second data files based on one or more configuration settings, each of the second data files being smaller than the first data file and named using a file naming convention that references content of a respective one of the second data files. and third logic configured to process a data request received from a client device, the data request being for a data portion of the first data file, the third logic further configured to provide the data portion back to the client device by accessing one or more of the first memory cache, the second memory cache, and the object storage to retrieve a corresponding portion of a corresponding one of the second data files identified based on a file naming convention, the first memory cache and the second memory cache having different eviction policies for data loaded therein from the object storage. In some examples, configuration settings are derived by the corresponding file converter based on metadata and time series data recorded in the first data file. In one specific example, the derived settings include a maximum number of centroid scans for the second data file or a maximum number of log records for the second data file. In some examples, different configuration settings may apply to different portions of the first data file, for example, as specified in the corresponding metadata portion thereof.In some examples, the metadata specifies or is derived from one or more instrument and / or detector settings used to acquire the first data file.
[0068] In some embodiments of the above apparatus, the scientific instrument comprises at least one of a mass spectrometer and a chromatography system.
[0069] In some embodiments of any of the above apparatus, the third logic is configured to access the second cache in response to a cache miss for the data portion in the first cache.
[0070] In some embodiments of any of the above apparatus, the third logic is further configured to query the deployment database to determine whether another service instance has a copy of the corresponding portion of the corresponding one of the second data files.
[0071] In some embodiments of any of the above apparatus, the third logic is further configured to, when the service instance does not have a copy of the corresponding portion of the corresponding one of the second data files, obtain a copy of the corresponding one of the second data files from the object storage.
[0072] In some embodiments of any of the above devices, the second cache is larger than the first cache.
[0073] In some embodiments of any of the above devices, the first cache is configured to evict the least recently accessed files when the first cache reaches a memory limit, and the second cache is a disk cache that includes a local directory structure that is asynchronously updated based on the total number of files cached therein or the total disk space used by them.
[0074] In some embodiments of any of the above devices, the third logic is communicatively connected to a plurality of data pods, each of the data pods having a respective first-in, first-out cache for storing data received directly or indirectly from the object storage.
[0075] In some embodiments of any of the above apparatus, the total number of data pods in the plurality of data pods is variable in response to changes in workload.
[0076] In some embodiments of any of the above devices, the data in each of the second files is organized into a plurality of row groups, each of the row groups having a corresponding plurality of data columns having different respective data types stored therein, each of the data columns being individually readable by the support device.
[0077] In some embodiments of any of the above apparatus, at least two of the first logic, the second logic, and the third logic are implemented by a common user computing device.
[0078] In some embodiments of any of the above apparatus, at least one of the first logic, the second logic, and the third logic is implemented by a computing device remote from the scientific instrument.
[0079] In some embodiments of any of the above apparatus, at least one of the first logic, the second logic, and the third logic is implemented in a scientific instrument.
[0080] In some embodiments of any of the above apparatus, the automated processing is configured to cause each of the plurality of second data files to have data for a fixed number of detector scans or output data sequences.
[0081] In some embodiments of any of the above apparatus, the plurality of second data files comprises more or less than 100 files.
[0082] According to another exemplary embodiment disclosed above, for example with reference to any one of FIGS. 1-8 or any combination of some or all thereof, there is provided an automated method implemented via a computing device for providing scientific instrument support, the method including: acquiring a first data file via one or more detectors of a scientific instrument, the first data file including unsegregated data from multiple scans or channels of the one or more detectors; and applying automated processing to the first data file to generate a corresponding plurality of second data files based on one or more configuration settings, each of the second data files being smaller than the first data file and each of the second data files including a file that references content of a respective one of the second data files. and processing a data request received from a client device, the data request being for a data portion of the first data file, the processing including providing the data portion back to the client device by accessing one or more of the first memory cache, the second memory cache, and the object storage to obtain a corresponding portion of a corresponding one of the second data files identified based on the file naming convention, the first memory cache and the second memory cache having different respective eviction policies for data loaded therein from the object storage.
[0083] In some embodiments of the above method, the processing includes accessing the second cache in response to a cache miss for the data portion in the first cache.
[0084] In some embodiments of any of the above methods, the processing includes querying the deployment database to determine whether another service instance has a copy of the corresponding portion of the corresponding one of the second data files, and when the service instance does not have a copy of the corresponding portion of the corresponding one of the second data files, obtaining a copy of the corresponding one of the second data files from object storage.
[0085] Some embodiments provide one or more non-transitory computer-readable media having instructions that, when executed by one or more computing devices for providing scientific instrument support, cause the one or more computing devices to perform any of the methods described above.
[0086] According to yet another exemplary embodiment disclosed above, for example with reference to any one or any combination of some or all of FIGS. 1-8 , there is provided a scientific instrument comprising at least one of a mass spectrometer and a chromatography system including one or more detectors; and a computing device configured to acquire a first data file via the one or more detectors, the first data file including unseparated data from multiple scans or channels of the one or more detectors; and apply automated processing to the first data file based on one or more configuration settings to generate a corresponding plurality of second data files, each of the second data files being smaller than the first data file and using a file naming convention that references the content of each of the second data files differently. a computing device configured to: apply a file naming convention to a plurality of second data files, the plurality of second data files being stored in an object storage; process a data request received from a client device, the data request being for a data portion of the first data file; and provide the data portion back to the client device by accessing one or more of the first memory cache, the second memory cache, and the object storage to obtain a corresponding portion of a corresponding one of the second data files identified based on the file naming convention, the first memory cache and the second memory cache having different respective eviction policies for data loaded therein from the object storage.
Claims
1. 1. A support device for a scientific instrument, said support device comprising: first logic configured to acquire a first data file via one or more detectors of the scientific instrument, the first data file including unseparated data from multiple scans or channels of the one or more detectors; second logic configured to apply automated processing to the first data file to generate a corresponding plurality of second data files based on one or more configuration settings of the first data file, each of the second data files being smaller than the first data file and named using a file naming convention that references content of each different one of the second data files, and the corresponding plurality of second data files being stored in object storage; and third logic configured to process a data request received from a client device, the data request being for a data portion of the first data file, the third logic further configured to provide the data portion back to the client device by accessing one or more of a first memory cache, a second memory cache, and the object storage to obtain a corresponding portion of a corresponding one of the second data files identified based on the file naming convention, the first memory cache and the second memory cache having different respective eviction policies for data loaded therein from the object storage.
2. The support device of claim 1 , wherein the scientific instrument comprises at least one of a mass spectrometer and a chromatography system.
3. 3. The support device of claim 1, wherein the third logic is configured to access the second memory cache in response to a cache miss for the data portion in the first memory cache.
4. 4. The assist device of claim 3, wherein the third logic is further configured to query a deployment database to determine whether another service instance has a copy of the corresponding portion of the corresponding one of the second data files.
5. 5. The assistance device of claim 4, wherein the third logic is further configured to obtain a copy of the corresponding one of the second data files from the object storage if there is no service instance having the copy of the corresponding portion of the corresponding one of the second data files.
6. 3. The support device of claim 1, wherein the second memory cache is larger than the first memory cache.
7. the first memory cache is configured to evict the least recently accessed file when the first memory cache reaches a memory limit; 3. The support device of claim 1, wherein the second memory cache is a disk cache that includes a local directory structure that is asynchronously updated based on the total number of files cached therein or the total disk space used thereby.
8. 3. The support device of claim 1 or 2, wherein the third logic is communicatively connected to a plurality of data pods, each of the data pods having a respective first-in, first-out cache for storing data received directly or indirectly from the object storage.
9. The support device of claim 8 , wherein the total number of data pods in the plurality of data pods is variable in response to changes in workload.
10. 3. The support device of claim 1, wherein the data in each of the second data files is organized into a plurality of row groups, each of the row groups having a corresponding plurality of data columns storing different respective data types therein, each of the data columns being individually readable by the support device.
11. The assistance apparatus of claim 1 or 2, wherein at least two of the first logic, the second logic, and the third logic are implemented by a common computing device.
12. 3. The support apparatus of claim 1, wherein at least one of the first logic, the second logic, and the third logic is implemented by a computing device remote from the scientific instrument.
13. 3. The support device of claim 1, wherein at least one of the first logic, the second logic, and the third logic is implemented within the scientific instrument.
14. 3. The assistance device of claim 1 or 2, wherein the automatic processing is configured to cause each of the plurality of second data files to have data for a fixed number of detector scans or output data sequences.
15. The support device according to claim 1 or 2, wherein the plurality of second data files comprises at least 100 files.
16. 1. An automated method implemented via a computing device for providing scientific instrument support, the method comprising: acquiring a first data file via one or more detectors of a scientific instrument, the first data file including unseparated data from multiple scans or channels of the one or more detectors; applying an automated process to the first data file to generate a corresponding plurality of second data files based on one or more configuration settings, each of the second data files being smaller than the first data file and named using a file naming convention that references content of each different one of the second data files, and the corresponding plurality of second data files being stored in object storage; and processing a data request received from a client device, the data request being for a data portion of the first data file, the processing including providing the data portion back to the client device by accessing one or more of a first memory cache, a second memory cache, and the object storage to obtain a corresponding portion of a corresponding one of the second data files identified based on the file naming convention, the first memory cache and the second memory cache having different respective eviction policies for data loaded therein from the object storage.
17. 17. The automated method of claim 16, wherein the processing includes accessing the second memory cache in response to a cache miss for the data portion in the first memory cache.
18. The processing comprises: querying a deployment database to determine whether another service instance has a copy of the corresponding portion of the corresponding one of the second data files; retrieving a copy of the corresponding one of the second data files from the object storage if no service instance has the copy of the corresponding portion of the corresponding one of the second data files.
19. 20. One or more non-transitory computer-readable media having instructions that, when executed by one or more computing devices for providing scientific instrument support, cause the one or more computing devices to perform the automated method of claim 16.
20. A scientific instrument, at least one of a mass spectrometer and a chromatography system including one or more detectors; 1. A computing device, comprising: acquiring a first data file via the one or more detectors, the first data file including unseparated data from multiple scans or channels of the one or more detectors; applying an automated process to the first data file to generate a corresponding plurality of second data files based on one or more configuration settings, each of the second data files being smaller than the first data file and named using a file naming convention that references content of each different one of the second data files, and the corresponding plurality of second data files being stored in object storage; processing a data request received from a client device, the data request being for a data portion of the first data file; and providing the data portion back to the client device by accessing one or more of a first memory cache, a second memory cache, and the object storage to obtain a corresponding portion of a corresponding one of the second data files identified based on the file naming convention, wherein the first memory cache and the second memory cache have different respective eviction policies for data loaded therein from the object storage.