Systems and methods for reproducibility verification
The system addresses the challenge of verifying research reproducibility by recording and graphically representing all steps in a research analysis, enabling transparent verification and reproduction of results.
Patent Information
- Application Number
- PCT/US2024/059360
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-15
- Filing Date
- 2024-12-10
- Publication Date
- 2025-06-19
AI Technical Summary
The reproducibility of research results is challenging to verify and authenticate, as existing methods rely on consumers to verify the reliability of the results without providing a clear audit trail of the production process.
A computer-implemented method and system that records all events in a collaborative computational analysis, generates a directed reproducibility graph showing all steps leading to a research result, and provides a certificate of reproducibility that allows others to step-through and reproduce the analysis pathway.
This approach ensures that research results are verifiable and reproducible by providing a transparent and auditable record of the data analysis process, shifting the onus of reproducibility from the consumer to the producer.
Smart Images

Figure US2024059360_19062025_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHODS FOR REPRODUCIBILITY VERIFICATIONField of the Invention
[0001] The present disclosure is related to systems and methods for verifying the reproducibility of research results. In particular, the present disclosure relates to methods and systems that provide an auditable research environment and tools to verify and authenticate results produced in such an environment.Background of the Invention
[0002] The debate on solving the reproducibility crisis has been very much focused on delivering the materials needed alongside a publication so that a reader is able to reproduce the findings. In many cases the focus is simply on delivering the raw data.
[0003] This approach makes the consumer responsible for verifying that what is produced is reliable. However, in most consumer-producer relationships, the onus is on the producer to prove that its product meets the standards. This is done for various good reasons, among which that the producer knows the exact steps of productions and can verify the accuracy / legality / 'goodness' of those steps. The consumer often does not know what was involved in making this product, and is left guessing.
[0004] The certificate of reproducibility shifts the onus of the reproducibility determination from the consumer to the producer. The producing scientists must show the full process by which they obtained the results as part of the responsibility of publishing the results.Summary’ of the Invention
[0005] In some embodiments provided herein is computer implemented method of auditing a data analysis process to be performed by at least one processor. The method includes performing a data ingest of a data set according to ingestion instructions; recording an ingestion event associated with the data ingest; performing a processing step on the data set according to processing instructions to produce an analyzed data set; recording a processing event associated with the processing step; generating an output according to the analyzed data set according to output instructions; recording an output generation event associated with generating the output; generating a certificate of reproducibility according to certificate generation instructions bygenerating a directed reproducibility graph of events contributing to the output, the events including at least the ingestion event, the processing event, and the output generation event; and publishing the output and the certificate of reproducibility according to publication instructions.Brief Description of the Figures
[0006] FIG. 1 is a schematic of a system configured for the generation of reproducibility certificates;
[0007] FIG. 2 is a flow chart illustrating a method for the generation of reproducibility certificates;
[0008] FIG. 3 illustrates an example research environment for use with reproducibility certificates consistent with embodiments hereof;
[0009] FIG. 4 illustrates a directed reproducibility' graph consistent with embodiments hereof;
[0010] FIG. 5 illustrates a method of generating a directed reproducibility graph consistent with embodiments hereof.Detailed Description of the Invention
[0011] The present disclosure provides systems and computer implemented methods for generating, managing, and otherwise implementing certificates of reproducibility. The presently disclosed technology represents an improvement in computer technology in the field of collaborational computational analysis. Specifically, the claims are directed to a method that records all events in a collaborational computational analysis, and, when a result is published, generates a directed reproducibility graph showing all steps leading to that result. Further, the claims provide an executable means of reproducing the analysis pathway.
[0012] Systems and methods described herein solve a technical problem associated with the production of research results. Specifically, the generation and production of research results may be difficult to verify, authenticate, and / or reproduce. In a research environment, a user may publish an output file set with associated visualizations and descriptions. Often, raw data may be provided in conjunction with such a publication. However, it may be difficult or impossible for a consumer of such a publication to go from the raw data to the output file setaccording to only a set of directions. Systems and methods described herein solve this technical problem by tracking and tracing steps, operations, processes, data sets. etc. involved in the development of a research result. These steps, operations, processes, data sets, etc. may then be systematized into a certificate of reproducibility that provides a later user with an ability to step-through or individually recreate (e.g., reproduce) all steps in a process to arrive at the research result, so as to verify the process and the result. Accordingly, the systems and methods described herein provide a technical improvement to computational environments used for research, e.g., for the processing, analysis, and output of data, by generating a verified “road map’’ that permits others to directly reproduce and verify the data analysis process.
[0013] FIG. 1 illustrates a reproducibility certificate system interacting with one or more research clients. In FIG. 1, an embodiment of a network environment is depicted. The network environment may include one or more reproducibility certificate systems (RCS) 102 in communication with one or more research clients 104 and one or more data retention systems 190 via one or more networks 199. In embodiments, the RCS 102 may operate locally or remotely with respect to the research clients. In embodiments, the RCS 102 may operate on the same hardware and / or as integrated software with the one or more research clients 104.
[0014] The RCS 102 may be configured as a server (e.g., having one or more server blades, processors, etc.), a personal computer (e.g., a desktop computer, a laptop computer, etc.), a smartphone, a tablet computing device, and / or other device that can be programmed to interface with the one or more research clients 104. In an embodiment, any or all of the functionality of the RCS 102 may be performed as part of a cloud computing platform. The RCS 102 is further discussed below with respect to FIG. 2.
[0015] The one or more clients 104 may be configured as a personal computer (e.g., a desktop computer, a laptop computer, etc.), a smartphone, a tablet computing device, and / or other device that can be programmed with a user interface for accessing a research environment. In embodiments, the RCS 102 and a client 104 may reside within a single system, such as a laptop, desktop, tablet, or other computing device with a user interface.
[0016] The network environment depicted in FIG. 1 represents an example embodiment of an RCS 102 configured to interface with one or more research clients 104. Although depicted as connected via network 199, any suitable series of individual or network connections may beemployed to permit an RCS 102 to control an automated cell engineering system installation 111 and access required resources such as various data retention systems 190.
[0017] The network 199 may be connected via wired or wireless links. Wired links may include Digital Subscriber Line (DSL), coaxial cable lines, or optical fiber lines. Wireless links may include Bluetooth®. Bluetooth Low Energy (BLE), ANT / ANT+, ZigBee, Z-Wave, Thread, Wi-Fi®, Worldwide Interoperability for Microwave Access (WiMAX®), mobile WiMAX®, WiMAX®- Advanced, NFC, SigFox, LoRa, Random Phase Multiple Access (RPMA), Weightless-N / P / W, an infrared channel or a satellite band. The wireless links may also include any cellular network standards to communicate among mobile devices, including standards that qualify as 2G, 3G. 4G, or 5G. Wireless standards may use various channel access methods, e.g., FDMA, TDMA, CDMA, or SDMA. In some embodiments, different types of data may be transmitted via different links and standards. In other embodiments, the same types of data may be transmitted via different links and standards. Network communications may be conducted via any suitable protocol, including, e.g.. http, tcp / ip, udp, ethemet, ATM, etc.
[0018] The network 199 may be any type and / or form of network. The geographical scope of the network may vary’ widely and the network 199 can be a body area network (BAN), a personal area network (PAN), a local-area network (LAN), e.g.. Intranet, a metropolitan area network (MAN), a wide area network (WAN), or the Internet. The topology of the network 199 may be of any form and may include, e.g., any of the following: point-to-point, bus, star, ring, mesh, or tree. The network 199 may be of any such network topology as known to those ordinarily skilled in the art capable of supporting the operations described herein. The network 199 may utilize different techniques and layers or stacks of protocols, including, e.g., the Ethemet protocol, the internet protocol suite (TCP / IP), the ATM (Asynchronous Transfer Mode) technique, the SONET (Synchronous Optical Networking) protocol, or the SDH (Synchronous Digital Hierarchy) protocol. The TCP / IP internet protocol suite may include application layer, transport layer, internet layer (including, e.g., IPv4 and IPv4), or the link layer. The network 199 may be a type of broadcast network, a telecommunications network, a data communication network, or a computer network.
[0019] The data retention systems 190 may include any type of computer readable storage medium (or media) and / or a computer readable storage device. Such computer readable storage medium or device may be configured to store and provide access to data. Examples ofcomputer readable storage medium or device may include, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof, for example, such as a computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick.
[0020] FIG. 2 illustrates an automated process control system consistent with embodiments hereof. The RCS 102 is a system for collaborational data analysis and reproducibility. The RCS 102 includes one or more processors 110 (also interchangeably referred to herein as processors 110, processor(s) 110, or processor 110 for convenience), one or more storage device(s) 120, and / or other components. In other embodiments, the functionality of the processor may be perfomied by hardware (e.g., through the use of an application specific integrated circuit (“ASIC”), a programmable gate array (“PGA”), a field programmable gate array (“FPGA”), etc.), or any combination of hardware and software. The storage device 120 includes any type of non-transitory computer readable storage medium (or media) and / or non- transitory computer readable storage device. Such computer readable storage media or devices may store computer readable program instructions for causing a processor to carry out one or more methodologies described here. Examples of the computer readable storage medium or device may include, but is not limited to an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof, for example, such as a computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, but not limited to only those examples.
[0021] The processor 110 is programmed by one or more computer program instructions stored on the storage device 120. For example, the processor 110 is programmed by anetwork manager 252, a user interface manager 254, a data storage manager 256, a research manager 258, and a certificate manager 260. It will be understood that the functionality of the various managers as discussed herein is representative and not limiting. Additionally, the storage device 120 may act as a data retention system 190 to provide data storage. As used herein, forconvenience, the various “managers"’ will be described as performing operations, when, in fact, the managers program the processor 110 (and therefore the RCS 102) perform the operation.
[0022] The various components of the RCS 102 work in concert to provide access to a research environment and provide reproducibility certificate functionality'.
[0023] The network manager 252 is a software protocol operating on the RCS 102. The network manager 252 is configured to establish a network communication between the RCS 102, data retention systems 190, and research clients 104. The established communications pathway may utilize any appropriate network transfer protocol and provide for one way or two way data transfer. The network manager 252 may establish as many network communications as required to communicate with one or more research clients 104. The network manager 252 provides access to the storage and functionality' of the RCS 102 to a user (e.g., a researcher).
[0024] The user interface manager 254 is a software protocol operating on the RCS 102. The user interface manager 254 is configured to provide a user interface to allow user interaction with the RCS 102. The user interface manager 254 is configured to provide a user interface (e.g., via hardware associated with the RCS 102 and / or a research client 104) configured to receive input from any user input source, including but not limited to touchscreens, keyboards, mice, controllers, joysticks, voice control. The user interface manager 254 is configured to provide a user interface, such as a text based user interface, a graphical user interface, or any other suitable user interface. The user interface manager 254 is configured to use the network manager 252 to provide such user interface services through the one or more research clients 104 or via hardware (e.g., a display device) associated with the RCS 102. The user interface manager 254 may be configured to provide different user interface services depending on a type of client device. For example, a laptop or desktop computer may be provided with a user interface including a full suite of interface options, while a smartphone or tablet may be provided with a user interface limited to status updates.
[0025] The user interface manager 254 is configured to provide user authentication services. Users may be authenticated via, for example, passwords, biometric scanning (retina scans, fingerprints, voice prints, facial recognition, etc.), key cards, token access, and any other suitable means of user authentication. User authentication services may be provided to control access to a research environment.
[0026] The data storage manager 256 is a software protocol operating on the RCS 102. The data storage manager 256 is configured to manage the receipt, storage, and retrieval of data (such as research data) associated with the RCS 102. The data storage manager 256 is further configured to access one or more data retention systems 190 to store and / or receive research data. The data storage manager 256 may provide data to a user via user interface manager 254. In embodiments, the data storage manager 256 is further configured to provide access tools to the user to manage and manipulate research data.
[0027] The research manager 258 is a software protocol operating on the RCS 102. The research manager 258 is configured to provide a research environment (such as the HISE illustrated in and described with respect to FIG. 3) to users for performing research, such as data analysis. The research manager 258 is configured to operate in conjunction with the certificate manager 260 to audit and / or reproduce a data analysis process. The research manager 258 is configured to perform any and all functions related to the functioning of a research environment. Broadly, the research manager 258 is configured for the ingestion of one or more data sets, the analysis of the ingested data, and the output of results based on the analysis.
[0028] The certificate manager 260 is a software protocol operating on the RCS 102. The certificate manager 260 is configured to audit and to facilitate replication of a data analysis process performed by the research manager 258. In embodiments, the certificate manager 260 is configured to record the details of a research process, including the steps executed to go from data ingestion to result publication. The certificate manager 260 operates in conjunction with the research manager 258 to record such steps and to generate a certificate of reproducibility accordingly. In embodiments, the certificate manager 260 is configured to execute a certificate of reproducibility by stepping through, in conjunction with the research manager 258, the recorded steps of the certificate to verify that the recorded steps are able to successfully produce the recorded output.
[0029] FIG. 3 illustrates an example research environment operable on the RCS 102, as instantiated by the software instructions associated with the research manager 258. The research environment 300 represents an example of a research environment compatible with the process auditing embodiments described herein. The research environment 300 as described herein is provided by way of example only, and the process auditing systems, methods, and techniques described herein may be applied to other environments and researchdomains that provide complex analysis tools. The research environment 300 is specifically applied to the field of immunology in that the system assumes basic metadata concepts such as human subjects and samples (tissue, PBMC). A different metadata model may be employed for a different research domain. Similarly, the multi-step automated analysis services may be general purpose but the specific pipelines implemented are based on the data generated for immunology research. In other arenas, different specific implementations of research tools may be employed without departing from the scope of this disclosure.
[0030] In embodiments, the research environment 300 may be implemented using specific computer based tools. For example, databases discussed herein may be implemented via Mongo DB or any other suitable database software. Services described herein may be Kubemetes (k8s) services, or any other suitable services. The file store may be a Google Storage bucket or other suitable storage tool. The IDE may be Jupyter Notebook or other suitable environment. The system may run on cloud servers, such as Google Cloud or AWS, but may also be implemented in local frameworks. All systems discussed herein may be abstracted such that alternate implementations can be accomplished including a different cloud provider, server-less services, R Studio or another form of IDE. The presently discussed invention provides a technical improvement to the computing systems described (e g., to the computer itself) by providing enhanced research auditing capabilities, as discussed herein.
[0031] The research environment 300 may include the following features and aspects and may represent an instantiation of the RCS 102 as well as additional features. The following description refers to assay data as a type of data ingested and analyzed by the research environment 300. It is understood that this is by way of example only and that any type of research data may be ingested and analyzed and that the tools described herein may
[0032] Ingest Components. The ingest components control the external data that enters the research environment 300. This includes LIMS data, human metadata 301, and “raw” (unanalyzed) assay data 302. The ingest components may include an ingest service 303. The Ingest Service 303 is an on-demand service that detects new data, and places ingested data either in a File Store 310 or a Ledger database 311. In embodiments, the File Store 310 may be used to store large, coherent blocks of data, including, for example, assay results. The Ledger database 311 may be used to stored metadata, e.g., relationship data, and may include references to the File Store 310. For example, the Ledger database 311 may store metadata describing a test, assay, or other experiment, and may include a reference or pointer to a file inthe File Store 310 that contains the data associated with the metadata. Thus, the Ledger database 311 may be used for quick access to information about tests, assays, etc., while the File Store 310 may be used for bulk storage of raw or analyzed data.
[0033] In embodiments, multistep Automated Analysis 306 may be provided. A number of services collaborate to deliver automated analysis of raw data, specific to the type of assay (e.g. Flow Cytometry, scRNA-seq, and so on). These may include a quality control sendee 307, an analysis service 308, and a labeling service 309.
[0034] The quality control service 307 performs basic verifications on file-based data and tags it with metadata denoting the origin of the assay (subject, sample, assay type, and so on). Quality control verifications may include various actions, such as verifying readability, completeness, and presence of identified data. The quality control service 307 may also perform assay specific QC checks.
[0035] The analysis service 308 executes the scientific analyses to interpret the data. Depending on the assay type this may involve a complex series of steps with interdependencies. Scientific analyses conducted by the analysis service 308 may include various analyses specific to the type of data collected. For example, for single cell sequencing data, the analysis sendee308 may examine sequenced data to arrive at an understanding of what cell types may be found in the sample and with what proportions using deterministic analysis. For flow cy tometry’ data, the analysis service 308 may try' to determine the cell types composition of a sample, using physical characteristics of each cell (heigh, with, transparency, etc) and then use a machine learning generated reference set to determine cell types. For proteomics analysis, the analysis service 308 may try to determine the proteins found in a cell. The above are presented by way of example only, and the analysis service 308 may be configured to provide any appropriate scientific analysis library according to the requirements of the input data.
[0036] The labeling service 309 is configured to interpret the results of the scientific analysis and to do categorization to reduce data complexity. For example, the labeling sendee309 may e configured to perform cell labeling, e.g., identifying, labeling, and / or classifying cells as monocytes, lymphocytes, T cells, subclasses of T cells (CD4, CD8. etc), B cell, NK cells, etc.
[0037] Each of these services may operate in communication with the file store 310 and / or the ledger database 311 to access and store data as it is used. For each step in the multistepAutomated Analysis 306 that produces a new or transformed data set, the file store 310 may be updated with the new data set and the ledger database 311 may be updated with metadata related to the newly produced data set. Such metadata may include, for example, information about the tests or assays to which the data pertains (e.g., subjects, samples, etc.), information about analysis performed, and information about file storage locations.
[0038] In embodiments, a suite of exploratory analysis tools 312 may be provided. A variety of follow-up analyses may be needed to further interpret results and make comparisons with other data. Exploratory' analysis tools 312 may be provided by visualization tools 313, an integrated development environment (IDE) 314, and an exploration tracking service 315. Interactive data visualization tools 313 are provided to permit non-coders to compare and interpret data. An IDE 314 is provided for coders to engage in more complex analyses. An exploration tracking service 315 is responsible for streaming data to the visualization tools and to the IDE; storing new data produced during complex analyses in an IDE as well as the IDE setups (including code) that are of interest to scientists, so these can be recalled at a later time; and storing visualization that are of interest to scientists so these can be recalled at a later time. Streaming the data to the visualization tools and visualizing the data as it is received may provide an improved user experience with respect to loading all data and then visualizing, because a user will have the opportunity to watch the process proceed and thus be reassured that the process is occurring appropriately. In some embodiments, the exploration tracking service 315 may transfer the necessary' data to the visualization tools and to the IDE in discrete blocks, rather than by streaming. New data or transformed data produced during use of the data visualizing tools 313 or the IDE 314 may be stored to the file store 310 while appropriate metadata is stored to the ledger database 311, as described above. The ledger database 311 persists all the data flow- and data processing, so that it may be recalled at a later time. The ledger database 311 may use MongoDB, SQL, or any ty pe of database technology7as the persistence mechanism. The file store 310, which stores large blocks of data, may use appropriate file storage tools, such as a Google storage bucket.
[0039] In embodiments, the research environment may include reproducibility7systems. The reproducibility systems may include a variety of services configured to generate certificates of reproducibility and make these available to anyone for inspection and re- execution. These include a certificate generation service 316, a certificate database 317, and a public access portal 318. The certificate generation service 316 is configured to produce acertificate of reproducibility for a publication. This includes constructing a reproducibility graph and storing its associated data and analysis processes. The certificate database 317 persists the certificate of reproducibility. This includes persisting the reproducibility graph, its data and its analysis processes. It may use MongoDB, SQL, or any type of database technology as the persistence mechanism. The public access portal 318 provides access to any end user to inspect and re-execute the certificate of reproducibility. It may be a web site or any other user interface that provides access.
[0040] FIG. 4 illustrates a directed reproducibility' graph of events that contribute to a processed data output by the RCS 102. The directed reproducibility graph 400 represents an example of a directed graph that represents information related to the processes carried out within the research environment to advance from ingested data to file set output. The directed reproducibility7graph 400 may represent a map of the connections between the file and process information and metadata, as stored in the ledger database 311. As illustrated in FIG. 4, the directed reproducibility graph 400 includes a directed series of nodes (vertices) 401 and links 402 that provide information regarding connections between nodes. Each node 401 represents an “event” within the research environment and may' correspond or associate with a data set, process action, or process state within the research environment. The various nodes 401 may include, but are not limited to ingestion receipt nodes, file nodes, process nodes, and file set nodes.
[0041] The directed reproducibility graph 400 may be constructed according to the data flow and data processing information stored in the ledger database 311. The directed reproducibility graph 400 may be represented by any suitable data structure, for example, as a list of nodes 401 (that represent the vertices of the graph) and a series of tuples storing connections betw een the nodes that represent the edges of the graph by <parent node, child node>. In embodiments, the directed reproducibility graph 400 is derived from the data flow and processing information stored in the ledger database 311 and is representative of that information. For example, each node of the directed reproducibility graph 400 is representative of an event related to a file or process, and is associated with respective node metadata that is stored in the ledger database 311 and node data that is stored in the file store 310. In embodiments, the directed reproducibility graph 400 may be two-way. and may include links and / or references to the represented information in the ledger database 31 1. For example, a given node 401 in the directed reproducibility graph 400 may be used to retrieve the nodemetadata in the ledger database 311 and / or the node data in the file store 310. In embodiments, the directed reproducibility’ graph 400 may be one-way. and may be representative of the node metadata in the ledger database 31 1 and the node data in the file store 310 information without including any specific links and / or references.
[0042] Ingestion receipt nodes (IR) represent information indicative of and associated with a data ingestion event. Each ingestion receipt node represents the ingestion of raw data (ingestion event) and represents or refers to node metadata (e.g., as stored in the ledger database 311) of relevant information, e.g., file ty pe, related research projects, associated users, etc., related to the ingested data and the ingestion event. Ingestion receipt nodes may also refer to and represent node data, e.g., the ingested data itself, as stored in the file store 310. As shown in FIG. 4, IR1 and IR2 represent ingestion receipt nodes. As discussed above, ingestion receipt nodes may include information suitable for locating the represented ingestion metadata.
[0043] File nodes (F) represent stored data, which may be raw data (e.g.. ingested data) or processed data. The node data (e.g., as stored in the file store 310) represented by each file node (F) may point to and / or otherwise represent a snapshot of the stored data at a point in time during the research process. File nodes (F) may also represent node metadata stored in the ledger database 311 and containing information about the stored data. As shown in FIG. 4, Fl, F2, and F3 represent raw data associated with data ingestion node IR1 and F4 represents raw data associated with data ingestion node IR2. F5, F6, F7, F8, F9, F10, and Fl l represent processed data resulting from processes at process node PR1 (F5), process node PR2 (F6, F7, F8), process node PR3 (F9), process node PR4 (F10), and process node PR5 (Fl l). As discussed above, file nodes may include information suitable for locating the represented file metadata and / or file data.
[0044] Process nodes (PR) represent data analysis or transformation process events. Actions or events represented by process nodes serve to act on one or more data sets to generate processed data sets. For example, the actions associated with the process node PR2 operate on the data associated with the file node F4 to generate new processed data associated with nodes F6, F7, and F8. Process nodes represent all the necessary node metadata (e.g., as stored in the ledger database 311) to describe and re-execute a process, such as a quality control process, analysis process, labeling process, or a code analysis in an IDE. As discussed above, process nodes may’ include information suitable for locating the represented process metadata and / or process data.
[0045] File Set nodes (FS1) are representative of a group of files representing the final results of an analysis. File set nodes may refer to or represent node metadata stored in the ledger database 31 1 describing the file set and also may refer to or represent node data as stored in the file store 310. As discussed above, file set nodes may include information suitable for locating the represented file set metadata and / or file set data.
[0046] Metadata nodes may represent and refer to all metadata associated with a file that is relevant to the analysis such as details of a subject such as demographics. In some embodiments, metadata may be associated with each of the other nodes described herein.
[0047] The directed reproducibility graph 400 represents an ordered record of operations conducted within a research environment and may be employed to validate and verify the operations.
[0048] FIG. 5 illustrates a computer-implemented method of auditing a data analysis process. The auditing method 500 may be performed or instantiated by the certificate manager 260 (e.g., operating the certificate generation service 316 ) in communication with the research manager 258 (e.g., operating multi-step automated analysis services 306 and the exploratory analysis tools 312). The method 500 may be performed to audit, verify, and / or reproduce a data analysis process performed within a research environment (e.g., research environment 300) by the research manager 258. This process may result in the generation of a certificate of reproducibility. Although presented in sequential order, it is not necessary or required for the steps and operations of the method 500 to be performed in a sequential order unless otherwise stated.
[0049] In embodiments, the certificate manager 260 and the research manager 258 may operate substantially simultaneously, wherein the certificate manager 260 operates to record the operations of the research manager 258 and build the certificate of reproducibility as the various operations of the research manager 258 are being completed. In embodiments, the certificate manager 260 may operate to build or generate the certificate of reproducibility at any point in the process analysis of the research manager 258. For example, a user may request that the certificate manager 260 build a certificate of reproducibility at the time that a publication output is generated by the research manager 258. In such an embodiment, the certificate manager 260 may access a record of processing steps recorded within the research environment 300 to build the reproducibility certificate.
[0050] FIG. 5 illustrates the operations of the auditing method 500 as parallel sequences performed by either the research manager 258 or the certificate manager 260. This arrangement is provided by way of example for illustrative purposes only and does not limit the various ways in which the steps of the method may be carried out.
[0051] In a data ingestion operation 502. a data ingest of a data set may be performed according to ingestion instructions. The data ingestion operation 502 may be performed according to ingestion instructions provided by the research manager 258 operating the ingest sendee 303. The data ingestion operation 502 may be performed by the ingest service 303, handling the data ingestion operation 502. The data ingestion operation 502 may operate to store the ingested data in the file store 310 or in the ledger database 311 and / or may pass the ingested data to the multi-step automated analysis services 306. The ingest service 303 may operate to generate an ingestion receipt recording the details of the data ingestion event.
[0052] In an ingestion event recording operation 512, an ingestion event associated with the data ingestion is recorded according to the information of the ingestion receipt. The ingestion event recording operation 512 may be performed by the certificate manager 260 operating the ingestion service 303 responsive to, subsequent to, or in conjunction with the ingestion event. In embodiments, the ingestion event may be recorded in the ledger database 311 and represented by an ingestion receipt node (e.g., IR1 as shown in FIG. 4) of a directed reproducibility graph that refers to the ingestion event. The information associated with the ingestion event may be stored as node metadata related to the ingestion node. The ingested data may stored in the file store 310 as node data and may represented by a file node of the directed reproducibility graph that is associated with the ingestion receipt node.. For instance, in Fig. 4, IR1 represents an ingest receipt containing a data ingestion event that ingests three files, represented by the file nodes Fl, F2 and F3.
[0053] In a data processing operation 504, a processing step is performed on the ingested data set (or sets) according to processing instructions to produce an analyzed data set. The data processing operation 504 may be performed by the research manager 258 operating any of the multi-step automated analysis services 306 or the exploratory' analysis tools 312 on the ingested data set to generate the analyzed data set.
[0054] In a data processing event recording operation 514, a data processing event associated with the data processing is recorded. The data processing event recording operation514 may be performed by the certificate manager 260 responsive to or subsequent to the data processing event. In embodiments, the data processing event may be recorded in the ledger database 311 as node metadata and represented by a processing node of a directed reproducibility graph. For example, in Fig 4, PR1 represents a data operation that included 2 input files Fl and F2, and generated one output file F5. Data sets produced by a data processing operation may be stored as node data represented by a file node and may also have metadata associated therewith stored as node metadata associated with the file node. Any operation that receives input data and produces output data may be considered a data processing event and recorded as such. Such operations may be simple and / or may be time consuming and complex.
[0055] In embodiments, additional instructions may be provided via the research manager 258 and associated recording events may be performed by the certificate manager 260. Such additional events may include, for example, additional ingestion, processing, and output generation events, as well as additional events associated with the previously described nodes 401 of the directed reproducibility graph 400. Such events may also be associated and / or represented with File Set nodes, and Metadata nodes.
[0056] In an output generation operation 506, an output according to the analyzed data set is generated according to output instructions. The output generation operation 506 may be performed by the research manager 258 to generate an output. Output generation operations produce new files. Any of the multi-step automated analysis services 306 or the exploratory analysis tools 312 may operate to generate an output data set.
[0057] In an output generation event recording operation 516, an output generation event associated with the output generation may be recorded. The output generation event recording operation 512 may be performed by the certificate manager 260 responsive to the output generation event. In embodiments, the output generation event may be recorded in the ledger database 311 with node metadata and may be represented as an output fileset node of the directed reproducibility graph 400, as a visualization node, or any other ty pe of output node. Data sets associated with the output generation operation may be stored in the file store 310 as node data associated with the fileset node. In embodiments, data sets associated with the output generation operation may be stored in the file store 310 as node data associated with the file nodes represented in the fileset. In Figure 4, for example, the File Set represented by FS1 was created containing the Files associated with F3, F5, F9, F10, and Fl 1.
[0058] Each of the operations and event recording operations described above are described in the singular. In practice, however, each operation and associated event recording operation may be performed many times over in any order or combination during use of the research environment.
[0059] In a certificate of reproducibility generation operation 518, a certificate of reproducibility may be generated according to certificate generation instructions by generating a directed reproducibility graph representing events contributing to the output of the output generation operation 506. The events may include at least the ingestion event, the processing event, and the output generation event. The certificate of reproducibility may further include the process flow information stored in the ledger database 311 and the appropriate file sets stored in the file store 310 related to each node in the directed reproducibility' graph of events.
[0060] Generating the certificate of reproducibility' from the directed reproducibility graph may include walking or tracing the graph backwards from an output node to identify links between nodes to arrive at an ingestion node. During operations within the research environment, many different operations may be performed, some of which may contribute to an output and some of which may not. Thus, all of the operations that occur may generate a very large reproducibility graph. In embodiments, only portions of the direct reproducibility graph may be relevant or of interest with respect to the output file set. Operations that do not contribute to the output may be excluded from the certificate of reproducibility produced for output. Accordingly, during a walk back process a modified directed reproducibility graph to be associated with the output file set may be produced. The original unmodified directed reproducibility graph may remain stored, for example, if a researcher wishes to revisit the research environment to perform further data analysis and generate additional results.
[0061] Referring now to FIG. 4, walking or tracing a graph backwards may start, for example, from the to-be-published file set node 1 (fsl). As shown in FIG. 4. file nodes 13. f5, fl 1, flO, and © are packaged together as the output fileset (fsl) and are thus included in the directed reproducibility7graph. Each of these file nodes is further walked back or traced to introduce the additional nodes that led to each file node until a valid ingest receipt (e.g., ir2, irl) is reached.
[0062] During the walk back that captures the nodes of the directed reproducibility7graph to generate the reproducibility' certificate, any process node that is traversed represents theexact details of the data operations used to analyze the data, as stored in the ledger database 311 as node metadata. The process nodes refer to and / or represent all of the necessary information, as stored in the ledger database 311, to describe and re-execute a process, such as a quality control process, analysis process, labeling process, or a code analysis in an IDE. The necessary information (e.g., file sets and process information) for reproduction may be stored in the reproducibility certificate from the relevant nodes during the walk back process.
[0063] Upon completion of the walk back and generation of the modified directed reproducibility7graph for the certificate of reproducibility7including the relevant nodes of interest, the certificate of reproducibility may be persisted and / or published with the modified directed reproducibility graph, all the data files that are represented by the graph, and the fully re-executable process details as represented by the process nodes.
[0064] In an operation 508, the method 500 may include a publishing operation. In the publishing operation, the generated output for publication and the associated certificate of reproducibility may be published, e g., to the public access portal 318 according to publication instructions. Publication may include the public provision of the output file set including visualizations, graphs, publication data, etc. that are associated with the directed reproducibility graph along with the certificate of reproducibility. As shown in in Figure 3. the certificate generation service 316 operated by the certificate manager 260 generates and stores all details of the certification in the certificate database 317. In addition to the certificate of reproducibility, the certificate database 317 may store or provide access to all file sets relevant to the certificate of reproducibility, including input, output, and interim file sets. In embodiments, the certificate of reproducibility7may be stored in the certificate database 317 in a non-editable or otherwise '‘locked” format. The public access portal 318 displays the full details of the certificate, allows access to all data sets, supports re-execution of any data operation process, and, in some embodiments, permits users to download certificates and associated data sets.
[0065] In embodiments, the certificate of reproducibility7generation operation 518 may be performed in conjunction with the publishing operation 508. When a user intends to publish an output file set, the certificate of reproducibility generation operation 518 may be performed such that the certificate of reproducibility may be published at a same time as the output file set.
[0066] In embodiments, the ingestion instructions, processing instructions, and output instructions may be provided by one or more users working in collaboration with one another. For example, the one or more users may be part of a research team, members of the same organization or laboratory, etc. In embodiments, the one or more users (e.g., first one or more users) may further provide additional instructions through the research manager 258 to perform additional operations within the research environment consistent with embodiments hereof.
[0067] Embodiments consistent with the present disclosure may include verifying a certificate of reproducibility by executing the certificate of reproducibility. The certificate of reproducibility may be executed by the first one or more users or by a second one or more users associated with a different organization, laboratory, research group, etc., than the first one or more users. Executing the certificate of reproducibility may include performing any or all of the research analysis operations (e.g., 502, 504, 560) of method 500. Executing the certificate of reproducibility may include perfomiing any additional operations performed within the research environment 300 by the first one or more users. Each operation may be performed according to the recorded processing event associated therewith and stored as a node of the directed reproducibility graph.
[0068] As discussed above, access to the certificate of reproducibility and all associated file sets (inputs, outputs, and interim file sets) may be provided by the public access portal 318 accessing the certificate database 317. The public access portal 318 provides a user with the ability to access the input file set and to perform a forward walk through the modified directed reproducibility graph by re-executing all of the processes associated with the nodes of the graph until the output file set is reached. In this manner, the user can verify that the output file set(s) was achieved from the input file set(s) according to the published steps of the certificate of reproducibility. In embodiments, the public access portal 318 may also permit a backwards walk through of the modified directed reproducibility graph from the output file set. In embodiments, the public access portal 318 may permit download of the certificate of reproducibility and the associated file sets to permit a user to independently verify the certificate with their own analysis tools.
[0069] Thus, for example, executing the certificate of reproducibility may include performing the processing steps on the data set according to the processing event to produce a second analyzed data set, and generating a second output according to the analyzed data set and the output generation event. Upon execution of the various data processing and outputgeneration steps of the method 500, the second output may be compared to the original output to verify reproducibility.
[0070] In some embodiments, an error state may occur within the research environment 300 that does not permit generation of a certificate of reproducibility. If one or more potential nodes traced back from a file output reach a dead end node, e.g., a node that cannot be further traced back through valid nodes to a valid ingest receipt, then the directed reproducibility graph and / or certificate of reproducibility may be considered invalid or incomplete. In embodiments, it may also be required that there be at least one processing node within the directed reproducibility graph (e.g., indicating that the graph does not simply pass data from ingest to output).
[0071] In embodiments, generating the certificate of reproducibility7may include determining a reproducibility score. The reproducibility score may be determined, for example, according to a completeness of the directed reproducibility graph, according to whether the certificate is executable or not, according to whether there is a processing pipeline or not, and / or according to whether there is an interactive visualization. In some embodiments, each of these factors may add one or more “points” out of a total point score. In some embodiments, the factors may interact. For example, some factors may be worth more points in the presence or absence of other factors. An example scoring method is shown below, in Table 1. For example, according to Table 1, an executable certificate with a processing pipeline and an interactive visualization may have a score of 5.
[0072] Further specific embodiments may include the following.
[0073] Embodiment 1 is a computer implemented method of auditing a data analysis process to be performed by at least one processor, the method comprising: performing a dataingest of a data set according to ingestion instructions; recording an ingestion event associated with the data ingest; performing a processing step on the data set according to processing instructions to produce an analyzed data set; recording a processing event associated with the processing step; generating an output according to the analyzed data set according to output instructions; recording an output generation event associated with generating the output; generating a certificate of reproducibility according to certificate generation instructions by generating a directed reproducibility graph of events contributing to the output, the events including at least the ingestion event, the processing event, and the output generation event; and publishing the output and the certificate of reproducibility according to publication instructions.
[0074] Embodiment 2 is the method of Embodiment 1, wherein the ingestion instructions, the processing instructions, the output instructions, and the certificate generation instructions are provided by a first one or more users.
[0075] Embodiment 3 is the method of Embodiment 2, further comprising executing the certificate of reproducibility according to reproduction instructions provided by a second user, different from the first one or more users, w herein executing the certificate of reproducibility includes: performing the processing step on the data set according to the processing event to produce a second analyzed data set; generating a second output according to the analyzed data set and the output generation event; and comparing the output to the second output to verify reproducibility.
[0076] Embodiment 4 is the method of any of Embodiments 1-3, wherein generating the certificate of reproducibility further includes determining a reproducibility score according to the directed reproducibility' graph of events.
[0077] Embodiment 5 is the method of Embodiment 4. wherein the reproducibility score indicates an incomplete certificate of reproducibility.
[0078] Embodiment 6 is the method of any of Embodiment 4 or 5, wherein the reproducibility score provides a measure of reproducibility of the output.
[0079] Embodiment 7 is the method of any of Embodiments 1-6, wherein generating the directed reproducibili ty graph includes tracing a first link between the output generation eventand the processing event and tracing a second link between the processing event and the ingestion event.
[0080] Embodiment 8 is the method of any of Embodiments 1-7, further comprising: performing a plurality of data ingests of a plurality of data sets; recording a plurality of ingestion events associated with the plurality of data ingests; performing a plurality of processing steps on the plurality’ of data set to produce a plurality of analyzed data sets; and recording a plurality’ of processing events associated with the plurality of processing steps; wherein: generating the output is further performed according to the plurality of analyzed data sets according to the output instructions, and generating the directed reproducibility graph of events contributing to the output is further based on the plurality’ of ingestion events and the plurality of processing events.
[0081] Embodiment 9 is the method of any of Embodiments 1-8, further comprising storing the certificate of reproducibility in a non-editable format.
[0082] Embodiment 10 is the method of any of Embodiments 1-9, further comprising: providing the certificate of reproducibility and the output for publication review; and receiving approval for each event in the directed reproducibility graph of events, wherein the publication instructions are provided responsive to receiving approval.
[0083] Embodiment 11 is a system for collaborational data analysis and reproducibility comprising at least one processor configured to execute software instructions, the software instructions configuring the at least one processor for auditing a data analysis process to be performed by at least one processor by: performing a data ingest of a data set according to ingestion instructions; recording an ingestion event associated with the data ingest; performing a processing step on the data set according to processing instructions to produce an analyzed data set; recording a processing event associated with the processing step; generating an output according to the analyzed data set according to output instructions; recording an output generation event associated with generating the output; generating a certificate of reproducibility according to certificate generation instructions by generating a directed reproducibility graph of events contributing to the output, the events including at least the ingestion event, the processing event, and the output generation event; and publishing the output and the certificate of reproducibility according to publication instructions.
[0084] Embodiment 12 is the system of Embodiment 11, wherein the ingestion instructions, the processing instructions, the output instructions, and the certificate generation instructions are provided by a first one or more users.
[0085] Embodiment 13 is the system of Embodiment 12, wherein the at least one processor is further configured for: executing the certificate of reproducibility according to reproduction instructions provided by a second user, different from the first one or more users, wherein executing the certificate of reproducibility includes: performing the processing step on the data set according to the processing event to produce a second analyzed data set; generating a second output according to the analyzed data set and the output generation event; and comparing the output to the second output to verily reproducibility.
[0086] Embodiment 14 is the system of any of Embodiments 11-13, wherein generating the certificate of reproducibility further includes determining a reproducibility score according to the directed reproducibility graph of events.
[0087] Embodiment 15 is the system of Embodiment 14, wherein the reproducibility score indicates an incomplete certificate of reproducibility.
[0088] Embodiment 16 is the system of any of Embodiments 14 or 15, wherein the reproducibility score provides a measure of reproducibility of the output.
[0089] Embodiment 17 is the system of any of Embodiments 11-16, wherein generating the directed reproducibility graph includes tracing a first link between the output generation event and the processing event and tracing a second link between the processing event and the ingestion event.
[0090] Embodiment 18 is the system of any of Embodiments 11 -17, wherein the at least one processor is further configured for: performing a plurality of data ingests of a plurality of data sets; recording a plurality of ingestion events associated with the plurality of data ingests; performing a plurality of processing steps on the plurality of data sets to produce a plurality of analyzed data sets; and recording a plurality of processing events associated with the plurality of processing steps; wherein: generating the output is further performed according to the plurality' of analyzed data sets according to the output instructions, and generating the directed reproducibility graph of events contributing to the output is further based on the plurality of ingestion events and the plurality of processing events.
[0091] Embodiment 19 is the sy stem of any of Embodiments 11-18, wherein the at least one processor is further configured for storing the certificate of reproducibility in a non-edi table format.
[0092] Embodiment 20 is the system of any of Embodiments 11-19. wherein the at least one processor is further configured for: providing the certificate of reproducibility’ and the output for publication review; and receiving approval for each event in the directed reproducibility graph of events, wherein the publication instructions are provided responsive to receiving approval.
[0093] Embodiment 21 is a computer implemented method of auditing a data analysis process to be performed by at least one processor, the method comprising: accessing a collaboration work space, the collaboration work space including at least one raw data set, at least one ingested data set. at least one analyzed data set, at least one data output, at least one data ingest event, at least one processing event, and one or more output generation events associated with the data output; identifying the at least one raw data set, the at least one ingested data set, the at least one analyzed data set, and the at least one data output as a plurality of nodes in a directed reproducibility graph; identifying a plurality of links between the plurality of nodes, wherein the plurality of links represent the at least one data ingest event, the at least one processing event, and the one or more output generation events; generating a directed reproducibility graph according to the plurality of nodes and the plurality of links measuring a degree of reproducibility of the directed reproducibility graph; generating a certificate of reproducibility including the at least one data output, the degree of reproducibility, and the directed reproducibility graph; and publishing the output and the certificate of reproducibility according to publication instructions.
[0094] While various embodiments of the methods and systems have been described, these embodiments are illustrative and in no way limit the scope of the described methods or systems. Those having skill in the relevant art can effect changes to form and details of the described methods and systems without departing from the broadest scope of the described methods and systems. Thus, the scope of the methods and systems described herein should not be limited by any of the illustrative embodiments and should be defined in accordance with the accompanying claims and their equivalents.
Claims
CLAIMS:
1. A computer implemented method of auditing a data analysis process to be performed by at least one processor, the method comprising: performing a data ingest of a data set according to ingestion instructions; recording an ingestion event associated with the data ingest; performing a processing step on the data set according to processing instructions to produce an analyzed data set; recording a processing event associated with the processing step; generating an output according to the analyzed data set according to output instructions; recording an output generation event associated with generating the output; generating a certificate of reproducibility according to certificate generation instructions by generating a directed reproducibility graph of events contributing to the output, the events including at least the ingestion event, the processing event, and the output generation event; and publishing the output and the certificate of reproducibility according to publication instructions.
2. The method of claim 1, wherein the ingestion instructions, the processing instructions, the output instructions, and the certificate generation instructions are provided by a first one or more users.
3. The method of claim 2, further comprising executing the certificate of reproducibility according to reproduction instructions provided by a second user, different from the first one or more users, wherein executing the certificate of reproducibility includes: performing the processing step on the data set according to the processing event to produce a second analyzed data set; generating a second output according to the analyzed data set and the output generation event; and comparing the output to the second output to verify reproducibility.
4. The method of claim 1, wherein generating the certificate of reproducibility further includes determining a reproducibility score according to the directed reproducibility graph of events.
5. The method of claim 4, wherein the reproducibility score indicates an incomplete certificate of reproducibility.
6. The method of claim 4, wherein the reproducibility score provides a measure of reproducibility of the output.
7. The method of claim 1, wherein generating the directed reproducibility graph includes tracing a first link between the output generation event and the processing event and tracing a second link between the processing event and the ingestion event.
8. The method of claim 1, further comprising: performing a plurality of data ingests of a plurality of data sets; recording a plurality of ingestion events associated with the plurality of data ingests; performing a plurality of processing steps on the plurality7of data set to produce a plurality of analyzed data sets; and recording a plurality of processing events associated with the plurality of processing steps; wherein: generating the output is further performed according to the plurality of analyzed data sets according to the output instructions, and generating the directed reproducibility graph of events contributing to the output is further based on the plurality7of ingestion events and the plurality7of processing events.
9. The method of claim 1, further comprising storing the certificate of reproducibilityin a non-editable format.
10. The method of claim 1. further comprising: providing the certificate of reproducibility and the output for publication review; and receiving approval for each event in the directed reproducibility graph of events, wherein the publication instructions are provided responsive to receiving approval.
11. A system for collaborational data analysis and reproducibility comprising at least one processor configured to execute software instructions, the software instructions configuring the at least one processor for auditing a data analysis process to be performed by at least one processor by: performing a data ingest of a data set according to ingestion instructions; recording an ingestion event associated with the data ingest; performing a processing step on the data set according to processing instructions to produce an analyzed data set; recording a processing event associated with the processing step; generating an output according to the analyzed data set according to output instructions; recording an output generation event associated with generating the output; generating a certificate of reproducibility according to certificate generation instructions by generating a directed reproducibility graph of events contributing to the output, the events including at least the ingestion event, the processing event, and the output generation event; and publishing the output and the certificate of reproducibility according to publication instructions.
12. The system of claim 1 1, wherein the ingestion instructions, the processing instructions, the output instructions, and the certificate generation instructions are provided by a first one or more users.
13. The system of claim 12, wherein the at least one processor is further configured for: executing the certificate of reproducibility according to reproduction instructions provided by a second user, different from the first one or more users, wherein executing the certificate of reproducibility includes: performing the processing step on the data set according to the processing event to produce a second analyzed data set; generating a second output according to the analyzed data set and the output generation event; and comparing the output to the second output to verify reproducibility.
14. The system of claim 11, wherein generating the certificate of reproducibility further includes determining a reproducibility score according to the directed reproducibility graph of events.
15. The system of claim 14, wherein the reproducibility score indicates an incomplete certificate of reproducibility.
16. The system of claim 14, wherein the reproducibility7score provides a measure of reproducibility of the output.
17. The system of claim 11, wherein generating the directed reproducibility graph includes tracing a first link between the output generation event and the processing event and tracing a second link between the processing event and the ingestion event.
18. The system of claim 11. wherein the at least one processor is further configured for: performing a plurality of data ingests of a plurality of data sets; recording a plurality of ingestion events associated with the plurality of data ingests; performing a plurality of processing steps on the plurality of data sets to produce a plurality of analyzed data sets; and recording a plurality of processing events associated with the plurality of processing steps; wherein: generating the output is further performed according to the plurality of analyzed data sets according to the output instructions, and generating the directed reproducibility graph of events contributing to the output is further based on the plurality of ingestion events and the plurality of processing events.
19. The system of claim 1 1, wherein the at least one processor is further configured for storing the certificate of reproducibility in a non-editable format.
20. The system of claim 11, wherein the at least one processor is further configured for:providing the certificate of reproducibility and the output for publication review; and receiving approval for each event in the directed reproducibility graph of events, wherein the publication instructions are provided responsive to receiving approval.
21. A computer implemented method of auditing a data analysis process to be performed by at least one processor, the method comprising: accessing a collaboration work space, the collaboration work space including at least one raw data set, at least one ingested data set, at least one analyzed data set, at least one data output, at least one data ingest event, at least one processing event, and one or more output generation events associated with the data output; identifying the at least one raw data set, the at least one ingested data set, the at least one analyzed data set, and the at least one data output as a plurality of nodes in a directed reproducibility graph; identifying a plurality of links between the plurality of nodes, wherein the plurality of links represent the at least one data ingest event, the at least one processing event, and the one or more output generation events; generating a directed reproducibility graph according to the plurality of nodes and the plurality7of links measuring a degree of reproducibility of the directed reproducibility graph; generating a certificate of reproducibility including the at least one data output, the degree of reproducibility, and the directed reproducibility graph; and publishing the output and the certificate of reproducibility according to publication instructions.
Citation Information
Patent Citations
Distributed data processing method with complete provenance and reproducibility
US11275726B1