Data storage device for scalable processing of large files generated by scientific instruments
By converting large data files generated by scientific instruments into multiple smaller files and adopting column-oriented storage formats and optimized cache management, the problem of inefficient data storage and access in scientific instruments is solved, achieving fast, flexible and low-cost data access and processing.
Patent Information
- Application Number
- CN202380081462.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-29
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, large data files generated by scientific instruments are inefficient in storage and access, resulting in high data access latency and storage costs, especially in cloud storage environments, making it difficult to achieve fast and flexible data access and processing.
Fast access and processing of data by converting large raw data files into multiple smaller data files and using column-oriented storage formats and optimized cache management strategies, combining object storage devices and distributed caches.
It significantly reduces data access latency, reduces storage costs, and supports parallel access to multiple client devices, improves data query performance and system flexibility, and adapts to different types of storage devices and deployment methods.
Smart Images

Figure CN120266102A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to U.S. Non - Provisional Application No. 18 / 148,248, filed on December 29, 2022, the entire content of which is incorporated herein by reference. Background of the Invention
[0003] This application generally relates to data storage devices, and more particularly but not exclusively to data processing and cache management.
[0004] Scientific instruments such as imaging instruments and spectrometers typically include a complex arrangement of components, sensors, detectors, input and output ports, energy sources, and consumable elements. Some such scientific instruments generate a relatively large amount of data when operating, which can affect memory requirements and efficient data access. Brief Description of the Drawings
[0005] Embodiments will be readily understood by the following detailed description in conjunction with the accompanying drawings. For ease of description, the same reference numerals denote the same structural elements. Embodiments are illustrated by way of example and not limitation in the figures of the accompanying drawings.
[0006] Figure 1 is a block diagram of an example scientific instrument support module for performing support operations according to various embodiments.
[0007] Figure 2 is a block diagram illustrating an example data structure (object) generated by a scientific instrument support module by applying an automated process to an original instrument data file via Figure 1 of.
[0008] Figure 3 is illustrative of Figure 2 a specific example of the object illustrated.
[0009] Figure 4 is a flowchart of a data delivery method applied by a scientific instrument support module according to various embodiments via Figure 1 of.
[0010] Figure 5 is a block diagram illustrating communication and data flow between various components of a distributed computing system according to various embodiments.
[0011] Figure 6 is a block diagram of an example computing device capable of performing some or all of the scientific instrument support methods and / or functions disclosed herein according to various embodiments.
[0012] Figure 7is a block diagram illustrating an example scientific instrument support system in which some or all of the scientific instrument support methods and / or functions disclosed herein may be implemented according to various embodiments.
[0013] Figure 8 is an illustration of a cloud-hosted deployment used with a Figure 1 scientific instrument support module according to various embodiments. DETAILED DESCRIPTION
[0014] Disclosed herein are scientific instrument support systems and related methods, computing and storage devices, and computer-readable media. For example, in some embodiments, a support apparatus for a scientific instrument includes a first logic component, a second logic component, and a third logic component. The first logic component is configured to obtain a first data file via one or more detectors of the scientific instrument, the first data file including unseparated data from a plurality of scans or channels of the one or more detectors. The second logic component is configured to apply an automated process to the first data file to generate a corresponding plurality of second data files based on one or more configuration settings, each of the second data files being smaller than the first data file and being named using a file naming convention that references corresponding content of different ones of the second data files. The plurality of second data files are stored in an object storage device. The third logic component is configured to process a data request received from a client device that is directed to a data portion of the first data file. The third logic component is further configured to provide the data portion back to the client device by accessing one or more of a first memory cache, a second memory cache, and the object storage device to obtain a corresponding portion of a corresponding one of the second data files identified based on the file naming convention. The first memory cache and the second memory cache have different respective eviction policies for data loaded into them from the object storage device.
[0015] The scientific instrument support embodiments disclosed herein may achieve improved performance relative to conventional methods and, for example, may achieve improved performance for time-of-flight (TOF) mass spectrometers, quadrupole mass spectrometers, ion trap mass spectrometers, and instruments including multiple spectrometers and / or detectors, such as ultraviolet (UV) detectors, diode array detectors (DADs, such as photodiode array detectors, PDAs), and chromatographic system detectors. For purposes of illustration and without any implied limitation, some example embodiments are described below with reference to a TOF mass spectrometer. Based on the provided description, one of ordinary skill in the relevant art will be able to make and use additional support embodiments for scientific instruments employing other types of spectrometers, detectors, devices, and various combinations thereof without any undue experimentation.
[0016] Time-of-flight mass spectrometry (TOFMS) is a mass spectrometry method in which the mass-to-charge ratio of ions is determined by measuring the time of flight. Ions are accelerated by an electric field of known intensity. This acceleration causes the ions to have the same kinetic energy as any other ion with the same charge. The velocity of the ions depends on the mass-to-charge ratio, such that heavier ions with the same charge reach a lower velocity than lighter ions. The time it takes for the ions to reach a downstream detector is measured. This time depends on the velocity of the ions and thus provides a measure of the mass-to-charge ratio of the ions. Based on the measured TOF mass spectrum, the mass-to-charge ratio of its various components, and other known experimental parameters, the composition of the analyte can usually be determined.
[0017] TOF mass spectrometers typically include a mass analyzer and a detector. An ion source (pulsed or continuous) is used to generate ions from the analyte. The TOF mass analyzer can be a linear flight tube or a reflector. In various examples, the ion detector is a microchannel plate (MCP) detector or a secondary electron multiplier (SEM). The electrical signal from the detector is digitized using a time-to-digital converter (TDC) or an analog-to-digital converter (ADC). The TDC is a counting detector, and the ion counting performed thereby is typically accompanied by summing a large number (e.g., hundreds) of individual mass spectra, which is sometimes referred to as histogramming. The corresponding TOF mass analyzer typically operates at a repetition rate of 5 kHz to 20 kHz to generate a sufficiently large number of mass spectra to be summed. The ADC typically operates at a speed of about 10 gigasamples per second to digitize the pulsed ion current from the MCP detector at discrete time intervals. In various examples, the ADC has a dynamic range of 8 bits to 12 bits. The use of an ADC (as opposed to a TDC) is more beneficial for some specific types of TOF mass spectrometers, such as matrix-assisted laser desorption / ionization (MALDI)-TOF instruments with relatively high peak currents.
[0018] The raw data detected by the mass spectrometer is typically in the form of signals distributed across various m / z (mass-to-charge ratio) values of the detected ions. Centroid data includes the raw data that has been processed via a suitable algorithm to retain only the local maximum in each mass range of the detected ions. Such centroid data is commonly referred to as a "centroid scan".
[0019] In some examples, the raw data file generated by a TOFMS instrument has a size of approximately 1 GB to 100 GB and includes data from multiple scans and / or channels, which is in an unseparated binary form of a file format optimized for the recording, transmission, and encapsulation of experimental data corresponding to a single analyte injection. In a typical data storage system, such a binary data file takes a relatively long time to be transferred from an object storage device to a file system that enables a corresponding file reader associated with the instrument (e.g., provided by the instrument's manufacturer) to extract selected scan and / or channel data for further processing, e.g., in response to a request from a client device. For larger data files, the corresponding data access latency typically increases non-linearly and can pose a significant obstacle to the users and operators of the corresponding scientific instrument and / or system.
[0020] In some use cases, an important consideration is the cost of the data storage device. For example, the service fee associated with an object storage device is typically lower than the service fee associated with a general-purpose data storage device (e.g., about one-fifth in some specific cases). Given the fact that for typical practical applications of various scientific instruments, the raw data is accessed relatively infrequently and usually in a partially selective manner after its initial acquisition, an object storage device is generally considered to be the preferred option.
[0021] Furthermore, the inventors recognize that in some examples, by adapting the data structure used for the object storage device to the way the corresponding data is used after initial acquisition, significant performance improvements can be achieved using the object storage device. In many use cases, such data structures are functionally different from the data structures that facilitate fast recording during data acquisition. For example, the inventors recognize that generating an extracted ion chromatogram (XIC) using a column-oriented storage device can be approximately an order of magnitude faster than when using the raw data format for the same purpose.
[0022] In a network environment, various cache memory systems can be used to accelerate data access in response to requests for data from client devices. However, given the size and structure of typical raw data files, cache memory systems may not be able to provide fast enough data access for client devices when dealing with large data files.
[0023] Using the various examples, aspects, features, and implementations of the systems and methods for data processing and cache management disclosed herein can advantageously address the above and possibly some other related problems in the prior art. In a representative example, a cache management system is designed and configured to consider various factors related to the type of measurements performed by an instrument and / or the data analysis criteria specified in a data access request received from a client device. In some examples, the cache management system operates in a cloud or enterprise environment to provide fast data access and data processing based on mass spectrometry criteria. In various examples, such mass spectrometry criteria include, but are not limited to, scans, scan types, instrument types, scan fragments corresponding to a specified mass range around the requested mass, and the like. At least some implementations advantageously reduce data access latency relative to the typical latency associated with prior data access solutions. At least some implementations support horizontal scaling, which advantageously mitigates latency fluctuations associated with changes to the workload of a data delivery service.
[0024] Accordingly, the implementations disclosed herein provide improvements to scientific instrument technology (e.g., improvements in the computer technology aspects that support such scientific instruments, among other improvements). For example, the various implementations disclosed herein can achieve improved (e.g., faster and more selective) access to scientific instrument data relative to conventional methods. In various examples, the improvements are aimed at achieving one or more of the following objectives: (a) the ability to store petabytes of raw data in inexpensive cloud storage devices while having relatively fast and selective access to any selected portion of the data; (b) the ability to scale caches and other data operations across multiple servers; (c) support for parallel (e.g., substantially simultaneous) access to data from many (e.g., up to a thousand) client devices; (d) multi-level caches for improved query performance; (e) using column-oriented file formats to quickly traverse time series data; (f) support for Representational State Transfer (REST) Application Programming Interfaces (APIs), gRPC (Unary General Remote Procedure Call), and streaming calls for better in-cluster performance; (g) compatibility with different types of object storage devices; (h) portability between local and cloud deployments through relatively straightforward configuration changes; and (i) flexibility in data access patterns, e.g., having Functions as a Service (FaaS) capabilities, various programming languages, "big data" solutions, and the like. In this document, the acronym "FaaS" stands for Functions as a Service, which is a class of cloud computing services that allows customers to develop, run, and manage application functionality without the complexity of building and maintaining the infrastructure typically associated with developing and launching applications.
[0025] According to some embodiments, a data storage method is provided, the data storage method comprising the steps of: (i) converting a large raw instrument data file into a corresponding plurality of smaller data files more suitable for storing time series data; (ii) generating a common metadata file for the plurality of such smaller data files, wherein the metadata details relevant parameters of the conversion process; (iii) assembling data portions from different regions of the raw instrument data file into records; (iv) dividing different streams stored in the raw instrument data file into one or more record streams; (v) further dividing the streams into groups, e.g., a fixed number (e.g., 100) of centroid scans selected for each stream; and (vi) using a suitable object storage device naming convention (e.g., conceptually similar to a file path) to quickly locate a desired group of a stream, and subsequently transferring the desired group of the stream from the object storage device to a cache memory to reduce the frequency of repeated transfers. In various examples, a cache management strategy that manages cache memory operations provides transfers of only a relatively small portion of data objects sufficient to address pending data requests. Example benefits of such cache management strategies include (i) significantly reducing the amount of data stored in the more expensive cache memory, and (ii) faster data access due to the smaller size of data portions transferred to the cache memory in the event of a cache miss. In some examples, the raw data REST and gRPC API services are designed to scale horizontally as demand increases, where the cache hierarchy is shared across multiple computing instances. In such examples, a distributed cache is utilized to enable synchronization of transferring data from the object storage device to the cache memory. The distributed cache allows one instance to initiate replication into the cache and allows other instances to wait until the transfer is complete when access to the same cache data is needed.
[0026] The various embodiments in the implementations disclosed herein can improve conventional methods to achieve the technical advantages of improved data operations performed via file conversion and data separation and with an optimized cache management strategy. Such technical advantages may not be achievable by routine and conventional methods, and all users of systems including such embodiments can benefit from these advantages (e.g., by assisting users in accelerating technical tasks such as the processing and analysis of experimental data). Thus, the technical features of the embodiments disclosed herein are clearly unconventional in the field of instrument-related data storage devices, as are the various combinations of the features disclosed herein. As further discussed herein, aspects of the embodiments disclosed in this document can, for example, improve the functionality of the computer itself by operating an instrument-related data storage device in an optimized manner, resulting in a higher level of productivity. The computing features disclosed herein not only involve the collection and comparison of information, but also apply new analysis and technical tools to transform the operation of instrument-related data storage devices. Thus, the present disclosure introduces functionality that cannot be performed by conventional computing devices or humans.
[0027] Accordingly, the embodiments of the present disclosure can serve any one of a number of technical purposes, such as controlling a particular technical system or method; determining how to control or configure a machine based on measurement results; or increasing the throughput of a data pipeline. Some of the examples disclosed herein provide solutions to technical problems, including but not limited to improvements to TOFMS instruments, e.g., improvements to the computer technology that supports TOFMS instruments, and other improvements.
[0028] In the following detailed description, reference is made to the accompanying drawings, which form a part of the detailed description, wherein like reference numerals always indicate like parts, and in which embodiments that may be practiced are shown by way of illustration. It should be understood that other embodiments may be utilized and structural or logical changes may be made without departing from the scope of the present disclosure. Accordingly, the following detailed description should not be taken as limiting.
[0029] The various operations may be described sequentially in a manner that is most helpful in understanding the subject matter disclosed herein as a number of discrete actions or operations. However, the described order should not be construed as implying that these operations must be order-dependent. Specifically, these operations may not be performed in the order presented. The described operations may be performed in an order different from that of the described embodiments. Various additional operations may be performed, and / or the described operations may be omitted in additional embodiments.
[0030] For purposes of this disclosure, the phrases "A and / or B" and "A or B" mean (A), (B), or (A and B). For purposes of this disclosure, the phrases "A, B, and / or C" and "A, B or C" mean (A), (B), (C), (A and B), (A and C), (B and C), or (A, B, and C). Although some elements may be expressed in the singular (e.g., "processing device"), any suitable element may be represented by multiple instances of that element and vice versa. For example, a set of operations described as being performed by a processing device may be implemented as different operations of these operations being performed by different processing devices.
[0031] This specification uses the phrases "an embodiment", "various embodiments", and "some embodiments", each of which may refer to one or more embodiments in the same or different embodiments. Additionally, terms such as "comprising", "including", "having", etc. as used with respect to embodiments of the present disclosure are synonymous. When used to describe a range of values, the phrase "between X and Y" means a range that includes X and Y. As used herein, "device" may refer to any individual device, a collection of devices, a part of a device, or a collection of parts of a device. The drawings are not necessarily drawn to scale.
[0032] Figure 1 is a block diagram illustrating a scientific instrument support module 1000 for performing support operations according to various embodiments. The scientific instrument support module 1000 may be implemented by a circuit (e.g., including electrical and / or optical components) such as a programmed computing device. The logical components of the scientific instrument support module 1000 may be included in a single common computing device or distributed across multiple computing devices that communicate with each other as the case may be. Reference is made herein to Figure 6 the computing device 6000 for examples of computing devices that may implement the scientific instrument support module 1000 individually or in combination, and reference is made herein to Figure 7 the scientific instrument support system 7000 for examples of systems of interconnected computing devices (wherein the scientific instrument support module 1000 may be implemented across one or more of these computing devices).
[0033] As Figure 1As shown, the scientific instrument support module 1000 includes a first logic component 1002, a second logic component 1004, and a third logic component 1006 for performing support methods for scientific instruments (such as TOFMS instruments) as described herein. As used herein, the term "logic component" may include a device configured to perform a set of operations associated with the logic component. For example, any of the logic elements included in the scientific instrument support module 1000 may be implemented by one or more computing devices programmed with instructions to cause one or more processing devices of the computing device to perform a related set of operations. In a particular embodiment, the logic element may include one or more non-transitory computer-readable media having instructions thereon that, when executed by one or more processing devices in one or more computing devices, cause the one or more computing devices to perform a related set of operations. As used herein, the term "module" may refer to a collection of one or more logic elements that together perform functions associated with the module. Different logic elements in the module may take the same form or may take different forms. For example, some of the logic components in the module may be implemented by a programmed general-purpose processing device, while other logic components in the module may be implemented by an application-specific integrated circuit (ASIC). In another example, different logic elements in the module may be associated with different instruction sets executed by one or more processing devices. A module may not include all of the logic elements depicted in the associated drawings; for example, when the module is to perform a subset of the operations discussed herein with reference to the module, the module may include a subset of the logic elements depicted in the associated drawings.
[0034] The first logic component 1002 may obtain one or more raw instrument data files corresponding to one or more analytes via one or more detectors of the TOFMS instrument. As indicated above, to perform such an acquisition, ions generated by the ion source of the TOFMS instrument are directed through the mass analyzer and detected via one or more detectors of the TOFMS instrument. Thus, the first logic component 1002 may obtain data via the TOFMS instrument under one or more settings of the mass analyzer and the detector. The first logic component 1002 may then include one or more parameters associated with the settings into the data file.
[0035] The second logic component 1004 may apply automated processing to one or more raw instrument data files obtained via the first logic component 1002. In various examples, the automated processing includes one or more of the following: (i) converting each data file in a large raw instrument data file into a corresponding plurality of smaller data files; (ii) generating a metadata file for the plurality of smaller data files based on one or more parameters associated with settings of a TOFMS instrument included in the raw data by the first logic component 1002 and further based on relevant parameters of the conversion from the large raw data file to the plurality of smaller data files; and (iii) naming each of the plurality of files using a naming convention suitable for referencing its content. In some examples, the processing steps of the conversion include some or all of the following sub-steps: (a) assembling different data portions of the raw data file into records; (b) dividing different data streams of the raw data file into one or more record streams; and (c) further dividing the streams into groups.
[0036] The third logic component 1006 can process data requests received from various client devices. Accordingly, based on the data requests, the third logic component 1006 can perform the following operations. First, the third logic component 1006 can map the received data requests to corresponding data files among a plurality of smaller data files generated via the second logic component 1004, for example, based on a naming convention. Once the name of the smaller data file is identified by the mapping, the third logic component 1006 can check the least recently used (LRU) cache in the memory for the identified data file. If the LRU cache has the identified data file, the third logic component 1006 causes the file to be retrieved from there, filtered to obtain the requested scan data, and sent back to the requesting client device. If the LRU cache does not have the identified data file, the third logic component 1006 initiates a search for the file in the local file cache. If the file is found in the local file cache, the third logic component 1006 causes the required data to be read from there and further causes a copy of the data file to be loaded into the LRU cache. If the file is not found in the local file cache, the third logic component 1006 can query a corresponding database (e.g., Redis) to see if another service instance has the corresponding object. If another service instance has a copy of the object, the third logic component 1006 can request the copy from that instance, for example, using a high-speed binary serialization protocol for the request. If no service instance has the object, the third logic component 1006 can request a copy of the object from the object storage of the local file system and update Redis to indicate that a local copy of the object now exists. The third logic component 1006 can then cause the local file copy to be read, added to the LRU cache in the memory, and directed to the requesting client device after appropriate data filtering as indicated above.
[0037] As used herein, the term "Redis" represents a remote dictionary server. Redis is an advanced key-value repository that can be used as a non-relational structured query language (NoSQL) database or as a memory-cache repository to improve performance when serving data stored in system memory. Redis particularly supports various data structures, such as strings, hashes, sets, lists, sorted sets, bitmaps, and geospatial indexes with radius queries. In various additional examples, other (besides Redis) distributed in-memory data repositories can also be used.
[0038] Figure 2FIG. 0 is a block diagram illustrating a data structure (object) 2000 that may be generated by applying an automated process to an original instrument data file via a second logic component 1004. In the example shown, object 2000 is a column-oriented file that is divided into row groups. For purposes of illustration and without any implied limitation, object 2000 is shown in Figure 2 as having two row groups, labeled row group 0 and row group 1, respectively. In other examples, the corresponding object 2000 may have a different (not two) number of row groups (see, e.g., Figure 3 ). Each row group in a row group has columns, and each column in a column is an array of data of the same data type. For purposes of illustration and without any implied limitation, each of row group 0 and row group 1 is shown as having five corresponding columns, labeled column A through column E, respectively. In a representative example, the columns of a row group may be read independently from disk into other memory.
[0039] In operation, the scientific instrument support module 1000 may cause multiple objects 2000 to be stored in an object storage device. When needed, the scientific instrument support module 1000 may cause a particular object 2000 required to access data stored therein to be copied into a local file cache of a service instance. The scientific instrument support module 1000 may further cause the respective columns of each row group to be read individually from the local file cache and cached into an in-memory LRU cache for faster access. In a representative example, the in-memory LRU cache is smaller than the local file cache, with example sizes of approximately 1 Gb and 100 Gb, respectively. The LRU cache and the local file cache have different respective eviction policies. In one example, the LRU cache is configured to evict the least recently accessed data when the LRU cache has reached its memory limit and additional data needs to be added to the cache. In contrast, the local file cache includes a local directory structure that may be cleaned asynchronously based on the total number of files cached therein and / or the total disk space used. For example, when individual files have metadata (such as a “created” timestamp and / or a “last accessed” timestamp), files may be evicted from the local file cache based on duration, e.g., by first deleting the oldest files, or based on the time elapsed since last access, e.g., by first deleting the least recently accessed files.
[0040] Figure 3A block diagram showing a specific example of the exemplary object 2000. In the example shown, the object 2000 has one hundred row groups, labeled RG1 through RG100 respectively. Each row group in row group RGn has six columns labeled 3001 through 3006. The data types of columns 3001 - 3006 are indicated in the header row 3010. In this example, the data types are: centroid scan number, mass, peak intensity, baseline, noise, and resolution. The number of rows in row group RGn can vary from row group to row group. In different specific examples, the object 2000 can have a different number of corresponding row groups RGn or the same fixed number of corresponding row groups RGn. In some examples, the mass range covered by the row groups can be reduced, where the number of row groups for each object 2000 can be increased to optimize the scan traversal for certain analytes and / or data analysis types.
[0041] In a representative example, a row group RGn in a column - oriented storage device is a set of arrays of manageable size. In some cases, row group RGn can have columns for scan number, mass, intensity, etc. Each of these columns can have a large number (e.g., 100,000) of entries (rows). However, some smaller auxiliary files may have only one row group, where each row contains a centroid peak, such that the entire auxiliary file has centroid peaks for a range of 100 scans. Each column in each row group RGn is compressed individually, such that the row group is more suitable as a data block that can be decompressed and loaded into memory. In some examples, there is no direct correspondence between centroid scans and row groups. In one example of such cases, an Optimized Row Columnar (ORC) file contains one hundred centroid scans, but limits each row group RGn to 10 centroid scans, where each row in the row group represents a centroid peak.
[0042] In one embodiment, the LRU cache key is assembled from context information. Examples of such context information are provided by the following string:
[0043] {injection identifier} / {column - oriented group} / {row group} / {column identifier}(1)
[0044] The naming convention can be such that if the initial raw data file is named "03_lumos_prg_sa_r1.raw", then the second logical component 1004 can be configured to use this file name as the injection identifier in the above string. If the values are for centroid scan numbers 1 through 100, and the mass values for row group 0 are cached, then the corresponding LRU key can be in the form of the following string:
[0045] 03_lumos_prg_sa_r1 / centroid_00001-00100 / 0 / mass(2)
[0046] The corresponding intensities for these masses can be accessed with another LRU key, for example, represented by the following strings:
[0047] 03_lumos_prg_sa_r1 / centroid_00001-00100 / 0 / intensity(3)
[0048] When retrieving a value from the LRU cache, the value can be in the form of an array of double-precision floating-point numbers, for example, because that particular data type is specified for the corresponding column in the column-oriented storage format. In other embodiments, other suitable naming conventions, as well as data types and formats, may also be used.
[0049] Figure 4 is a flowchart of a data delivery method 4000 according to one embodiment. In one example, method 4000 is implemented using a scientific instrument support module 1000. Continuing to refer below to Figures 1 to 3 to describe method 4000.
[0050] Method 4000 includes the scientific instrument support module 1000 receiving a request for data from a client device (at block 4002). In a typical example, the request identifies a particular data slice that needs to be returned to the client device in response to the request. Such identification can be performed using an applicable naming convention, exemplary examples of which have been described above. For example, the request received at block 4002 can specify a centroid scan for the injection "03_lumos_prg_sa_r1", scan number 42. Thus, the corresponding object 2000 can be Figure 3 an example object 2000.
[0051] Method 4000 also includes the scientific instrument support module 1000 checking for the requested data slice in the LRU cache in memory (at block 4004). Such checking at block 4004 can include generating a corresponding LRU cache lookup key, which can be in the form indicated by the above strings (2) and (3), for example. Method 4000 also includes determining whether the LRU cache has the (smaller) data file identified by the key (at decision block 4006). If the LRU cache has the data file ("yes" at decision block 4006), the processing of method 4000 is directed to block 4020. Otherwise ("no" at decision block 4006), the processing of method 4000 is directed to block 4008.
[0052] Method 4000 also includes the Scientific Instrument Support Module 1000 checking for a corresponding data file in the local file cache (at block 4008). Such checking at block 4008 can include searching for the corresponding file name in the directory structure of the local file cache using an applicable file naming convention. Method 4000 also includes determining whether the local file cache has a file identified by the file name (at decision block 4010). If the local file cache has the file (the "yes" at decision block 4010), the processing of Method 4000 is directed to block 4011. Otherwise (the "no" at decision block 4010), the processing of Method 4000 is directed to block 4012. Method 4000 also includes the Scientific Instrument Support Module 1000 causing a local file copy to be read and added to the LRU cache (at block 4011).
[0053] Method 4000 includes the Scientific Instrument Support Module 1000 querying Redis to check whether another service instance has a copy of the corresponding file (at block 4012). Method 4000 also includes the Scientific Instrument Support Module 1000 receiving a response from Redis and determining based on the received response whether another service instance has such a copy (at decision block 4014). If another service instance has the file (the "yes" at decision block 4014), the processing of Method 4000 is directed to block 4018. Otherwise (the "no" at decision block 4014), the processing of Method 4000 is directed to block 4016.
[0054] Method 4000 also includes the Scientific Instrument Support Module 1000 requesting a copy of the file (object) from the object storage device and receiving the requested copy (at block 4016). Method 4000 also includes the Scientific Instrument Support Module 1000 updating Redis to indicate that a local copy of the object now exists (at block 4016). Method 4000 also includes the Scientific Instrument Support Module 1000 causing the received file copy to be added to the LRU and local file caches (at block 4018).
[0055] Method 4000 also includes the Scientific Instrument Support Module 1000 applying data filtering to the corresponding file (at block 4020). The filtering applied is according to the request received at block 4002. Method 4000 also includes the Scientific Instrument Support Module 1000 directing the resulting filtered data to the client device (at block 4020).
[0056] Figure 5FIG. is a block diagram illustrating communication and data flow among various components of a distributed computing system 5000 according to one embodiment. In at least some examples, communication and data flow in system 5000 may be according to method 4000. In some examples, a scientific instrument support module 1000 may communicate operatively with system 5000. In some other examples, some portions of scientific instrument support module 1000 may be implemented within system 5000.
[0057] In the example shown, system 5000 includes API server instances 5004 and 5012, a local LRU memory cache 5006, a local disk cache 5008, a Redis in-memory data repository 5010, and an object storage device 5014. Figure 5 Various transmissions within the shown system 5000 are triggered by a request 5022 for data received by API server instance 5004 from a client device 5002. Request 5022 may be according to the operation of block 4002 of method 4000.
[0058] In response to request 5022, API server instance 5004 sends a corresponding query 5024 to local LRU memory cache 5006. Query 5024 may be according to the operation of block 4004 of method 4000. Local LRU memory cache 5006 sends a response 5026 to query 5024, which returns the requested data or notifies API server instance 5004 that the requested data was not found. Using the latter result of query 5024, API server instance 5004 sends a corresponding query 5028 to local disk cache 5008. Query 5028 may be according to the operation of block 4008 of method 4000. Local disk cache 5008 sends a response 5030 to query 5028, which returns the requested data or notifies API server instance 5004 that the requested data was not found. Using the latter result of query 5028, API server instance 5004 sends the next corresponding query 5032 to Redis in-memory data repository 5010. Query 5032 may be according to the operation of block 4012 of method 4000. Redis in-memory data repository 5010 sends a response 5034 to query 5032, which identifies other API server instances having the requested data or notifies API server instance 5004 that no such server instances exist.
[0059] In one example result of query 5032, response 5034 identifies API server instance 5012. Using such a result, API server instance 5004 sends a request 5036 for the corresponding data to API server instance 5012. In response to request 5036, API server instance 5012 returns a response 5038 having the requested data.
[0060] In another example result of query 5032, response 5034 notifies the API server instance 5004 that there is no server instance with corresponding data. Using this result, the API server instance 5004 operates to access the object storage device 5014 via command 5040 and receive back the corresponding data stream 5042 with the corresponding data file (e.g., object 2000). Command 5040 and data read 5042 may operate according to block 4016 of method 4000. Once the object is received, the API server instance 5004 saves the obtained object and / or its relevant parts in the local LRU memory cache 5006 and the local disk cache 5008. The corresponding data writes 5044, 5046 may operate according to the relevant operations of blocks 4011, 4018. The API server instance 5004 also operates to send an update 5048 to the Redis in-memory data repository 5010, which provides appropriate information about the obtained object to the databases therein.
[0061] As noted above, the scientific instrument support module 1000 may be implemented by one or more computing devices. Figure 6 is a block diagram of a computing device 6000 that can perform some or all of the scientific instrument support functions and / or methods disclosed herein according to various embodiments. In some embodiments, the scientific instrument support module 1000 may be implemented by a single computing device 6000 or by multiple computing devices 6000. Additionally, as discussed below, the computing device 6000 (or multiple computing devices 6000) implementing the scientific instrument support module 1000 may be Figure 7 part of one or more of the scientific instrument 7010, the user local computing device 7020, the service local computing device 7030, or the remote computing device 7040.
[0062] Figure 6 The computing device 6000 is illustrated as having multiple components, but any one or more of these components may be omitted or duplicated depending on the applicability to the application and settings. In some embodiments, some or all of the components included in the computing device 6000 may be attached to one or more motherboards and encapsulated in a housing (e.g., including plastic, metal, and / or other materials). In some embodiments, some of these components may be fabricated onto a single system-on-chip (SoC) (e.g., the SoC may include one or more processing devices 6002 and one or more storage devices 6004). Additionally, in various embodiments, the computing device 6000 may not include Figure 6One or more of the illustrated components, but may include interface circuitry (not explicitly shown) for coupling to one or more components using any suitable interface (e.g., Universal Serial Bus (USB) interface, High-Definition Multimedia Interface (HDMI) interface, Controller Area Network (CAN) interface, Serial Peripheral Interface (SPI) interface, Ethernet interface, wireless interface, or any other suitable interface). For example, computing device 6000 may not include display device 6010, but may include display device interface circuitry (e.g., connectors and driver circuitry) to which display device 6010 may be coupled.
[0063] Computing device 6000 may include a processing device 6002 (e.g., one or more processing devices). As used herein, the term "processing device" may refer to any device or portion of a device that processes electronic data from registers and / or memory to transform that electronic data into other electronic data that may be stored in registers and / or memory. Processing device 6002 may include one or more digital signal processors (DSPs), application specific integrated circuits (ASICs), central processing units (CPUs), graphics processing units (GPUs), cryptographic processors (specialized processors that execute cryptographic algorithms in hardware), server processors, or any other suitable processing device.
[0064] Computing device 6000 may include a storage device 6004 (e.g., one or more storage devices). Storage device 6004 may include one or more memory devices, such as random access memory (RAM) devices (e.g., static RAM (SRAM) devices, magnetic RAM (MRAM) devices, dynamic RAM (DRAM) devices, resistive RAM (RRAM) devices, or conductive-bridging RAM (CBRAM) devices), hard-drive-based memory devices, solid-state memory devices, network drives, cloud drives, or any combination of memory devices. In some embodiments, storage device 6004 may include a memory that shares a die with processing device 6002. In such embodiments, the memory may be used as a cache memory and may include, for example, embedded dynamic random access memory (eDRAM) or spin-transfer torque magnetic random access memory (STT-MRAM). In some embodiments, storage device 6004 may include a non-transitory computer-readable medium having instructions that, when executed by one or more processing devices (e.g., processing device 6002), cause computing device 6000 to perform any suitable method or portion of the methods disclosed herein.
[0065] The computing device 6000 may include interface device 6006 (e.g., one or more interface devices 6006). The interface device 6006 may include one or more communication chips, connectors, and / or other hardware and software to manage communication between the computing device 6000 and other computing devices. For example, the interface device 6006 may include circuitry for managing wireless communication that is used to transfer data to and from the computing device 6000. The term “wireless” and its derivatives may be used to describe circuits, devices, systems, methods, techniques, communication channels, etc., that can convey data by using modulated electromagnetic radiation through a non-solid medium. The term does not imply that the associated devices do not contain any wires, although in some embodiments they may not contain any wires. The circuitry for managing wireless communication included in interface device 7006 may implement any one of a number of wireless standards or protocols, including but not limited to Institute of Electrical and Electronics Engineers (IEEE) standards, including Wi-Fi (IEEE 802.11 family), IEEE 802.16 standards (e.g., IEEE 802.16-2005 amendment), Long Term Evolution (LTE) project, and any amendments, updates, and / or revisions thereof (e.g., LTE-Advanced project, Ultra Mobile Broadband (UMB) project (also known as “3GPP2”), etc.). In some embodiments, the circuitry for managing wireless communication included in interface device 7006 may operate in accordance with Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Universal Mobile Telecommunications System (UMTS), High Speed Packet Access (HSPA), Evolved HSPA (E-HSPA), or LTE networks. In some embodiments, the circuitry for managing wireless communication included in interface device 7006 may operate in accordance with Enhanced Data GSM Evolution (EDGE), GSM EDGE Radio Access Network (GERAN), Universal Terrestrial Radio Access Network (UTRAN), or Evolved UTRAN (E-UTRAN). In some embodiments, the circuitry for managing wireless communication included in interface device 7006 may operate in accordance with Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Digital Enhanced Cordless Telecommunications (DECT), Evolution-Data Optimized (EV-DO), and their derivative protocols, and any other wireless protocols designated as 3G, 4G, 5G, and higher generations, etc. In some embodiments, the interface device 7006 may include one or more antennas (e.g., one or more antenna arrays) for receiving and / or transmitting wireless communication.
[0066] In some embodiments, the interface device 6006 may include circuitry for managing wired communications, such as electrical communication protocols, optical communication protocols, or any other suitable communication protocol. For example, the interface device 6006 may include circuitry that supports communication according to Ethernet technology. In some embodiments, the interface device 6006 may support both wireless and wired communications, and / or may support multiple wired communication protocols and / or multiple wireless communication protocols. For example, a first set of circuitry of the interface device 6006 may be dedicated to short-range wireless communication such as Wi-Fi or Bluetooth, while a second set of circuitry of the interface device 7006 may be dedicated to long-range wireless communication such as Global Positioning System (GPS), EDGE, GPRS, CDMA, WiMAX, LTE, EV-DO, etc. In some embodiments, a first set of circuitry of the interface device 6006 may be dedicated to wireless communication, while a second set of circuitry of the interface device 6006 may be dedicated to wired communication.
[0067] The computing device 6000 may include a battery / power circuit 6008. The battery / power circuit 6008 may include one or more energy storage devices (e.g., batteries or capacitors), and / or circuitry for coupling components of the computing device 6000 to an energy source separate from the computing device 6000 (e.g., an AC line power).
[0068] The computing device 6000 may include a display device 6010 (e.g., multiple display devices). The display device 6010 may include any visual indicator, such as a head-up display, a computer monitor, a projector, a touchscreen display, a liquid crystal display (LCD), a light-emitting diode display, or a flat panel display.
[0069] The computing device 6000 may include other input / output (I / O) devices 6012. The other I / O devices 6012 may include, for example, one or more audio output devices (e.g., speakers, headphones, earbuds, sirens, etc.), one or more audio input devices (e.g., a microphone or a microphone array), a positioning device (e.g., a GPS device known in the art that communicates with a satellite-based system to receive the location of the computing device 6000), an audio codec, a video codec, a printer, sensors (e.g., a thermocouple or other temperature sensor, a humidity sensor, a pressure sensor, a vibration sensor, an accelerometer, a gyroscope, etc.), an image capture device such as a camera, a keyboard, a cursor control device such as a mouse, a stylus, a trackball, or a touchpad, a barcode reader, a quick response (QR) code reader, or a radio frequency identification (RFID) reader.
[0070] The computing device 6000 can have any suitable form factor appropriate for its applications and settings, such as a handheld or mobile computing device (e.g., a cellular phone, smartphone, mobile Internet device, tablet computer, laptop computer, netbook computer, ultrabook computer, personal digital assistant (PDA), ultra-mobile personal computer, etc.), a desktop computing device, or a server computing device or other networked computing components.
[0071] In some examples, the computing device 6000 is implemented using multiple containers (pods) in a Kubernetes cluster. A representative Kubernetes cluster includes multiple computer nodes configured to be able to host multiple containers, with each container acting as a virtual machine. In various deployments, several instances of a microservice can run on a single container or multiple containers (see also Figure 8 ). In some examples, better performance is achieved when some of the multiple containers are distributed across different computer nodes.
[0072] One or more computing devices implementing any of the scientific instrument support modules or methods disclosed herein can be part of a scientific instrument support system. Figure 7 is a block diagram of an example scientific instrument support system 7000 in which some or all of the scientific instrument support methods disclosed herein can be executed according to various embodiments. The scientific instrument support modules and methods disclosed herein (e.g., Figure 1 's scientific instrument support module 1000 and Figure 4 's method 4000) can be implemented by one or more of the scientific instrument 7010, user local computing device 7020, service local computing device 7030, and remote computing device 7040 of the scientific instrument support system 7000.
[0073] Any one of the scientific instrument 7010, user local computing device 7020, service local computing device 7030, or remote computing device 7040 can include any of the embodiments of the computing device 6000 discussed herein with reference to Figure 6 , and any one of the scientific instrument 7010, user local computing device 7020, service local computing device 7030, or remote computing device 7040 can take the form of any suitable embodiment of the computing device 6000 discussed herein with reference to Figure 6 .
[0074] The scientific instrument 7010, the user local computing device 7020, the service local computing device 7030, and / or the remote computing device 7040 may each include a corresponding processing device 6002, a corresponding storage device 6004, and a corresponding interface device 6006. The processing device 6002 may take any suitable form, including the form of any of the processing devices 6002 discussed herein with reference to Figure 6 any of the processing devices 6002 discussed herein, and the processing devices 6002 included in different devices among the scientific instrument 7010, the user local computing device 7020, the service local computing device 7030, or the remote computing device 7040 may take the same form or different forms. The storage device 6004 may take any suitable form, including the form of any of the storage devices 6004 discussed herein with reference to Figure 6 any of the storage devices 6004 discussed herein, and the storage devices 6004 included in different devices among the scientific instrument 7010, the user local computing device 7020, the service local computing device 7030, or the remote computing device 7040 may take the same form or different forms. The interface device 6006 may take any suitable form, including the form of any of the interface devices 6006 discussed herein with reference to Figure 7 any of the interface devices 6006 discussed herein, and the interface devices 6006 included in different devices among the scientific instrument 7010, the user local computing device 7020, the service local computing device 7030, or the remote computing device 7040 may take the same form or different forms.
[0075] The scientific instrument 7010, the user local computing device 7020, the service local computing device 7030, and the remote computing device 7040 may communicate with other elements of the scientific instrument support system 7000 via a communication path 7008. The communication path 7008 may communicatively couple the interface devices 6006 of different elements among the elements of the scientific instrument support system 7000 (as shown), and may be a wired or wireless communication path (e.g., according to any of the communication technologies discussed for the interface device 6006 of the computing device 6000 herein with reference to Figure 6 any of the communication technologies). Figure 7The specific scientific instrument support system 7000 depicted therein includes communication paths between each pair of devices among the scientific instrument 7010, the user local computing device 7020, the service local computing device 7030, and the remote computing device 7040. However, this specific implementation of "fully connected" is merely illustrative, and in various embodiments, various communication paths in the communication paths 7008 may not exist. For example, in some embodiments, the service local computing device 7030 may not have a direct communication path 7008 between its interface device 6006 and the interface device 6006 of the scientific instrument 7010, but may communicate with the scientific instrument 7010 via the communication path 7008 between the service local computing device 7030 and the user local computing device 7020 and the communication path 7008 between the user local computing device 7020 and the scientific instrument 7010. The scientific instrument 7010 may include any suitable scientific instrument, such as a TOFMS instrument.
[0076] Figure 8 is a block diagram illustrating a cloud-hosted deployment 8000 for use with the scientific instrument support module 1000 according to one embodiment. The deployment 8000 is a K8s deployment with horizontal scaling and the use of local volumes. Herein, K8s (also known as Kubernetes or "kube") refers to an open-source container orchestration platform that automates many of the manual processes involved in deploying, managing, and scaling containerized applications. Horizontal scaling means that in response to an increased workload, the deployment 8000 increases the number N of the deployed databases 8010 n where n = 1, 2,..., N. This response is different from vertical scaling, which typically involves allocating more resources (e.g., memory or CPU) to the database 8010 that is already running for the workload n . When the workload decreases and the number N of the deployed data containers 80101 - 8010 N is higher than a preset minimum value, the number N scales down proportionally. In various examples, each data container 8010 in the deployment 8000 n has a corresponding configurable first-in-first-out (FIFO) memory cache.
[0077] The deployment 8000 runs a data service 8002 that is configured to process data requests received from various client devices (such as the client device 5002 ( Figure 5 ))), such as the request 5022 ( Figure 5 ). The deployment 8000 includes an object storage device 5014 (also see Figure 5)。Deployment 8000 also includes Redis cluster 8020, which is particularly used for data bookkeeping, container discovery, and orchestrating publish / subscribe (pub / sub) events. Pub / sub messaging is a form of asynchronous service-to-service communication that can be used in serverless and microservices architectures. In the pub / sub model, any message published to a topic is received by all subscribers of that topic. In deployment 8000, pub / sub messaging is used to distribute change events from the corresponding database. These events can be used to construct views of the database state and state history for parallel processing and workflows.
[0078] In operation, blocks of the stream that have been transferred from the object storage device 5014 are cached on the local volume of the corresponding data container 8010 n . Different data containers among data containers 80101 - 8010 N can query the Redis cluster 8020 to see which data container 8010 n has the requested data, if any. Such queries can be in accordance with the operation of block 4012 of method 4000. Figure 5 The query 5032 shown is an example of such a query. The Redis cluster 8020 responds to the query 5032 by identifying the data container 8010 n that has the requested data or notifying the querying data container that no data container 80101 - 8010 N has the requested data. Depending on the specific response from the Redis cluster 8020, the querying data container is used to request and receive data from the identified data container 8010 n via one or more lateral links 8012 (also see Figure 5 components 5036, 5038 therein), or retrieve data from the object storage device 5014 (also see Figure 5 components 5040, 5042 therein).
[0079] According to the example implementation schemes disclosed above, for example, with reference to Figures 1 to 8In any one or any combination of some or all of the accompanying drawings, there is provided a support device for a scientific instrument, the support device comprising: a first logic component configured to obtain a first data file via one or more detectors of the scientific instrument, the first data file comprising unseparated data from a plurality of scans or channels of the one or more detectors; a second logic component configured to apply an automated process to the first data file to generate a corresponding plurality of second data files based on one or more configuration settings, each of the second data files being smaller than the first data file and being named using a file naming convention that references the respective content of a different second data file in the second data files, the corresponding plurality of second data files being stored in an object storage device; and a third logic component configured to process a data request received from a client device, the data request being directed to a data portion of the first data file, the third logic component further configured to provide the data portion back to the client device by accessing one or more of a first memory cache, a second memory cache, and the object storage device to obtain a corresponding portion of the corresponding one second data file in the second data files identified based on the file naming convention, the first memory cache and the second memory cache having different respective eviction policies for data loaded into them from the object storage device. In some examples, the configuration settings are derived by a corresponding file converter based on metadata and time series data recorded in the first data file. In a specific example, the derived settings include the maximum number of centroid scans per second of the data file or the maximum number of log records per second of the data file. In some examples, different configuration settings may apply to different portions of the first data file, for example, as specified in their corresponding metadata portions. In some examples, the metadata specifies or is derived from one or more instrument and / or detector settings used to obtain the first data file.
[0080] In some embodiments of the above device, the scientific instrument includes at least one of a mass spectrometer and a chromatography system.
[0081] In some embodiments of any one of the above devices, the third logic component is configured to access the second cache in response to a cache miss for the data portion in the first cache.
[0082] In some embodiments of any one of the above devices, the third logic component is further configured to query a deployment database to determine whether another service instance has a copy of the corresponding portion of the corresponding one second data file in the second data files.
[0083] In some embodiments of any of the above devices, the third logic component is further configured to obtain a copy of the corresponding one of the second data files in the second data file from the object storage device when no service instance has a copy of the corresponding portion of the corresponding one of the second data files in the second data file.
[0084] In some embodiments of any of the above devices, the second cache is larger than the first cache.
[0085] In some embodiments of any of the above devices, the first cache is configured to evict the least recently accessed file when the first cache reaches the memory limit; and wherein the second cache is a disk cache, the disk cache including a local directory structure that is updated asynchronously based on the total number of files cached therein or the total disk space used thereby.
[0086] In some embodiments of any of the above devices, the third logic component is communicatively coupled to a plurality of data containers, each of the data containers having a corresponding first-in / first-out cache for storing data received directly or indirectly from the object storage device.
[0087] In some embodiments of any of the above devices, the total number of data containers among the plurality of data containers is variable in response to a change in the workload.
[0088] In some embodiments of any of the above devices, the data in each of the second files is organized into a plurality of row groups, each of the row groups having a corresponding plurality of data columns, different corresponding data types being stored in the corresponding plurality of data columns, and each of the data columns being readable separately by the support device.
[0089] In some embodiments of any of the above devices, at least two of the first logic component, the second logic component, and the third logic component are implemented by a common computing device.
[0090] In some embodiments of any of the above devices, at least one of the first logic component, the second logic component, and the third logic component is implemented by a computing device remote from the scientific instrument.
[0091] In some embodiments of any of the above devices, at least one of the first logic component, the second logic component, and the third logic component is implemented in the scientific instrument.
[0092] In some embodiments of any of the above devices, the automated process is configured to cause each of the plurality of second data files to have a fixed number of detector scans or output data sequences.
[0093] In some embodiments of any of the above devices, the plurality of second data files have more or fewer than one hundred files.
[0094] According to another exemplary embodiment disclosed above, for example, with reference to Figures 1 to 8 any one or any combination of some or all of the figures in, there is provided an automated method for providing scientific instrument support performed via a computing device, the method comprising: obtaining a first data file via one or more detectors of a scientific instrument, the first data file comprising unseparated data from a plurality of scans or channels of the one or more detectors; applying an automated process to the first data file to generate a corresponding plurality of second data files based on one or more configuration settings, each of the second data files being smaller than the first data file and being named using a file naming convention that references corresponding content of different second data files in the second data files, the corresponding plurality of second data files being stored in an object storage device; and processing a data request received from a client device that is directed to a data portion of the first data file, the processing comprising providing the data portion back to the client device by accessing one or more of a first memory cache, a second memory cache, and the object storage device to obtain a corresponding portion of a corresponding one of the second data files identified based on the file naming convention, the first memory cache and the second memory cache having different respective eviction policies for data loaded into them from the object storage device.
[0095] In some embodiments of the above method, the processing comprises accessing the second cache in response to a cache miss for the data portion in the first cache.
[0096] In some embodiments of any of the above methods, the processing comprises: querying a deployment database to determine whether another service instance has a copy of the corresponding portion of the corresponding one of the second data files; and when no service instance has a copy of the corresponding portion of the corresponding one of the second data files, obtaining a copy of the corresponding one of the second data files from the object storage device.
[0097] Some embodiments provide one or more non-transitory computer-readable media having instructions thereon that, when executed by one or more computing devices to provide scientific instrument support, cause the one or more computing devices to perform any of the above methods.
[0098] According to yet another example embodiment disclosed above, for example, with reference to Figures 1 to 8 any one or any combination of some or all of the figures in, there is provided a scientific instrument comprising: at least one of a mass spectrometer and a chromatography system including one or more detectors; and a computing device configured to: obtain a first data file via the one or more detectors, the first data file including unseparated data from a plurality of scans or channels of the one or more detectors; apply an automated process to the first data file to generate a corresponding plurality of second data files based on one or more configuration settings, each of the second data files being smaller than the first data file and named using a file naming convention that references corresponding content of different second data files in the second data files, the corresponding plurality of second data files being stored in an object storage device; process a data request received from a client device that is directed to a data portion of the first data file; and provide the data portion back to the client device by accessing one or more of a first memory cache, a second memory cache, and the object storage device to obtain a corresponding portion of a corresponding one of the second data files identified based on the file naming convention, the first memory cache and the second memory cache having different respective eviction policies for data loaded into them from the object storage device.
Claims
1. A support device for a scientific instrument, the support device comprising: A first logic component configured to obtain a first data file via one or more detectors of the scientific instrument, the first data file including unseparated data from a plurality of scans or channels of the one or more detectors; A second logic component configured to apply an automated process to the first data file to generate a corresponding plurality of second data files based on one or more configuration settings, each of the second data files being smaller than the first data file and named using a file naming convention that references corresponding content of different second data files among the second data files, the corresponding plurality of second data files being stored in an object storage device; And A third logic component configured to process a data request received from a client device, the data request being for a data portion of the first data file, the third logic component further configured to provide the data portion back to the client device by accessing one or more of a first memory cache, a second memory cache, and the object storage device to obtain a corresponding portion of a corresponding one of the second data files identified based on the file naming convention, the first memory cache and the second memory cache having different respective eviction policies for data loaded into them from the object storage device.
2. The support device according to claim 1, wherein the scientific instrument includes at least one of a mass spectrometer and a chromatography system.
3. The support device according to any one of claims 1 to 2, wherein the third logic component is configured to access the second cache in response to a cache miss for the data portion in the first cache.
4. The support device according to claim 3, wherein the third logic component is further configured to query a deployment database to determine whether another service instance has a copy of the corresponding portion of the corresponding one of the second data files.
5. The support device according to claim 4, wherein the third logic component is further configured to obtain a copy of the corresponding one of the second data files from the object storage device when no service instance has a copy of the corresponding portion of the corresponding one of the second data files.
6. The support device according to any one of claims 1 to 2, wherein the second cache is larger than the first cache.
7. The support device according to any one of claims 1 to 2, wherein the first cache is configured to evict the least recently accessed file when the first cache reaches a memory limit; and wherein the second cache is a disk cache, the disk cache including a local directory structure that is updated asynchronously based on the total number of files cached therein or the total disk space used thereby.
8. The support apparatus according to any one of claims 1 to 2, wherein the third logic component is communicatively coupled to a plurality of data containers, each of the data containers having a respective first-in / first-out cache to store data received directly or indirectly from the object storage device.
9. The support apparatus according to any one of claims 8, wherein the total number of data containers in the plurality of data containers is variable in response to a change in the workload.
10. The support apparatus according to any one of claims 1 to 2, wherein the data in each of the second files is organized in a plurality of row groups, each of the row groups having a corresponding plurality of data columns, different respective data types being stored in the corresponding plurality of data columns, and each of the data columns being individually readable by the support apparatus.
11. The support apparatus according to any one of claims 1 to 2, wherein at least two of the first logic component, the second logic component, and the third logic component are implemented by a common computing device.
12. The support apparatus according to any one of claims 1 to 2, wherein at least one of the first logic component, the second logic component, and the third logic component is implemented by a computing device remote from the scientific instrument.
13. The support apparatus according to any one of claims 1 to 2, wherein at least one of the first logic component, the second logic component, and the third logic component is implemented in the scientific instrument.
14. The support apparatus according to any one of claims 1 to 2, wherein the automated process is configured to cause each of the plurality of second data files to have a fixed number of detector scans of data or output data sequences.
15. The support apparatus according to any one of claims 1 to 2, wherein the plurality of second data files have at least one hundred files.
16. An automated method for providing scientific instrument support, performed via a computing device, the method comprising: obtaining, via one or more detectors of a scientific instrument, a first data file that includes unseparated data from a plurality of scans or channels of the one or more detectors; applying an automated process to the first data file to generate a corresponding plurality of second data files based on one or more configuration settings, each of the second data files being smaller than the first data file and being named using a file naming convention that references the respective content of different second data files among the second data files, the corresponding plurality of second data files being stored in an object storage device; and Process a data request received from a client device, the data request being for a data portion of the first data file, the processing including providing the data portion back to the client device by accessing one or more of a first memory cache, a second memory cache, and the object storage device to obtain a corresponding portion of a corresponding one of the second data files identified based on the file naming convention, the first memory cache and the second memory cache having different respective eviction policies for data loaded into them from the object storage device.
17. The automated method according to claim 16, wherein the processing includes accessing the second cache in response to a cache miss for the data portion in the first cache.
18. The automated method according to claim 17, wherein the processing includes: Querying a deployment database to determine whether another service instance has a copy of the corresponding portion of the corresponding one of the second data files in the second data file; And When no service instance has a copy of the corresponding portion of the corresponding one of the second data files in the second data file, obtaining a copy of the corresponding one of the second data files in the second data file from the object storage device.
19. One or more non-transitory computer-readable media having instructions thereon that, when executed by one or more computing devices to provide scientific instrument support, cause the one or more computing devices to perform the automated method according to claim 16.
20. A scientific instrument, the scientific instrument comprising: At least one of a mass spectrometer and a chromatography system including one or more detectors; And A computing device configured to: Obtain a first data file via the one or more detectors, the first data file including unseparated data from a plurality of scans or channels of the one or more detectors; Apply an automated process to the first data file to generate a corresponding plurality of second data files based on one or more configuration settings, each of the second data files in the plurality of second data files being smaller than the first data file and being named using a file naming convention that references corresponding content of different second data files in the plurality of second data files, the corresponding plurality of second data files being stored in an object storage device; Process a data request received from a client device, the data request being for a data portion of the first data file; And Provide the data portion back to the client device by accessing one or more of a first memory cache, a second memory cache, and the object storage device to obtain a corresponding portion of a corresponding one of the second data files identified based on the file naming convention, the first memory cache and the second memory cache having different respective eviction policies for data loaded into them from the object storage device.