Systems and methods for analyzing, storing, and sharing genomic data using blockchains
By constructing a decentralized network and statistical compression methods using blockchain technology, the issues of data scale and privacy protection in genomic data sharing are solved, achieving efficient and secure genomic data storage and access control, and promoting multi-party participation and governance.
Patent Information
- Application Number
- CN202480017749.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-18
- Filing Date
- 2024-01-18
- Publication Date
- 2025-11-21
AI Technical Summary
Existing genomic data sharing faces challenges such as massive data volume, privacy protection issues, complex data ownership and access control, and low efficiency of traditional centralized storage platforms, which hinder the open and responsible sharing of genomic data.
By employing blockchain technology to construct a decentralized distributed network, combined with statistical compression methods and encryption technology, genome sequencing data is compressed and stored through the blockchain, enabling secure sharing and access control of genome data. Non-transitory computer-readable storage devices and processors are used to execute relevant methods, supporting the comparison and analysis of genome data.
It enables efficient storage and secure sharing of genomic data, ensures data privacy protection, simplifies access control, promotes multi-party participation and governance, and improves the transparency and fairness of data sharing.
Smart Images

Figure CN121002575A_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 439,700, filed January 18, 2023, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This disclosure generally relates to systems and methods for analyzing, storing and / or sharing genomic data, such as genomic data of humans and other life forms (such as animals), and more specifically, to systems and methods for analyzing, storing and / or sharing genomic data as compressed archives and / or using blockchain. Background Technology
[0004] Human genome data carries unique information about individuals and offers unprecedented opportunities for healthcare. Genome data from other life forms, such as animals, are also useful across various industries.
[0005] Clinical interpretations derived from large genomic datasets could significantly improve healthcare and pave the way for personalized medicine. Genomic data, due to its massive size (e.g., a single human genome occupies 100 gigabytes of storage), typically requires computer systems and storage devices for processing and storage. With the rapid advancements in genome sequencing, the volume of genomic data generated is growing exponentially. Several large-scale sequencing projects targeting humans and other species / life forms are expected to further increase the volume of such data.
[0006] However, sharing genomic datasets via computer systems can present significant challenges because, unlike traditional medical data, genomic data may indirectly reveal information about the data owner's descendants and relatives, and may still carry valuable information even after the owner's death. Therefore, strict data ownership and control measures are required for computerized genomic data storage and sharing.
[0007] Promoting open and responsible sharing of genomic data has become a core principle of many national and international initiatives, such as the U.S. All of Us research program and the Global Genome & Health Alliance. Despite widespread recognition of the importance of genomic data sharing, technological and governance bottlenecks hinder its progress. Recent surveys of genome sequencing projects indicate that a lack of consistency and interoperability in bioinformatics pipelines, insufficient financial support, and legal, consent, and privacy-related issues are major challenges facing genomic data sharing. Summary of the Invention
[0008] According to one aspect of this disclosure, a computerized first method for compressing genome sequencing data is provided, the first method comprising: aligning genome sequencing data with reference sequencing data; obtaining one or more differential read sequences, each of the one or more differential read sequences being a read sequence of genome sequencing data that is different from a corresponding read sequence of the reference sequencing data; and obtaining compressed genome sequencing data by compressing the one or more differential read sequences using a statistical compression method or an assembly method using a probabilistic data structure.
[0009] In some implementations, the first method further includes: processing one or more differential read sequences using one or more statistical modeling methods; wherein, compressing one or more differential read sequences using a statistical compression method includes: compressing the processed one or more differential read sequences using a statistical compression method.
[0010] In some implementations, the first method further includes assembling multiple reads to form reference sequencing data.
[0011] In some implementations, processing one or more differential read sequences includes: resolving a first read identifier (ID) of the genome sequencing data into one or more first fields (tokens); storing one or more fields of the first read ID; resolving a second read ID of the genome sequencing data into one or more second fields; and obtaining one or more differential fields, each differential field representing a field in one or more second fields that differs from a corresponding field in one or more first fields.
[0012] In some implementations, one or more difference fields include a first difference field that corresponds to a numeric field in one or more second fields that is different from a corresponding field in one or more first fields; and the first difference field includes the numeric field, or includes an offset between the numeric field and a corresponding field in one or more first fields.
[0013] In some implementations, one or more difference fields include a second difference field that corresponds to a non-numeric field in one or more second fields that is different from a corresponding field in one or more first fields; and the second difference field includes a portion of the non-numeric field that is different from a corresponding portion of a corresponding field in one or more first fields.
[0014] In some implementations, processing one or more differential read sequences includes compressing one or more nucleotide sequences in one or more differential read sequences using a Markov-chain-based statistical model.
[0015] In some implementations, the statistical model is based on a 12th-order Markov chain.
[0016] In some implementations, the first method further includes assembling multiple contigs into reference sequencing data using multiple reads of the genome sequencing data.
[0017] In some implementations, assembling multiple contigs into reference sequencing data using multiple reads from genome sequencing data includes constructing a de Bruijn graph using a Bloom filter-based data structure.
[0018] In some implementations, assembling multiple contigs into reference sequencing data using multiple reads from genome sequencing data includes constructing a de Brouin graph using a data structure based on a d-left count Bloom filter (dlCBF).
[0019] In some implementations, the comparison of genome sequencing data with reference sequencing data includes: using a sparse bitmap representation of the de Bruin graph with Elias-Fano encoded compression.
[0020] In some implementations, the first method further includes storing the compressed genome sequencing data in a blockchain.
[0021] According to one aspect of this disclosure, one or more non-transitory computer-readable storage devices are provided, including computer-executable instructions for compressing genome sequencing data, wherein, when executed, the instructions cause a processing structure to perform the first method described above.
[0022] According to one aspect of this disclosure, one or more processors are provided, the processors being configured to perform the first method described above.
[0023] According to one aspect of this disclosure, a computerized second method is provided, comprising: storing at least a first portion of an entity’s genome sequencing data as transaction data in network nodes of a blockchain system.
[0024] In some implementations, the second method further includes converting at least a first portion of the entity’s genome sequencing data into a unique first hash string.
[0025] In some implementations, the second method further includes combining a first hash string with one or more first hash strings from at least a first portion of the genome sequencing data of one or more other entities from the blockchain system to obtain a second hash string.
[0026] In some implementations, at least a first portion of the entity’s genome sequencing data is stored as one or more binary alignment and mapping (BAM) files.
[0027] In some implementations, one or more binary comparison and mapping (BAM) files are associated with a first hash string.
[0028] In some implementations, the second method further includes storing at least a first portion of an entity’s genome sequencing data as one or more non-fungible tokens (NFTs) in network nodes of the blockchain system.
[0029] In some implementations, the second method further includes: using pseudo-anonymity and public key infrastructure (PKI) to protect at least a first portion of the entity’s genome sequencing data.
[0030] In some implementations, the second method further includes: denying unauthorized users access to at least a first portion of the entity's genome sequencing data.
[0031] In some implementations, the second method further includes presenting metadata of at least a first portion of the entity’s genome sequencing data to an unauthorized user.
[0032] In some implementations, at least a first portion of the entity’s genome sequencing data is stored in a blockchain system, while a second portion of the entity’s genome sequencing data is not stored.
[0033] According to one aspect of this disclosure, one or more non-transitory computer-readable storage devices are provided, including computer-executable instructions for compressing genome sequencing data, wherein, when executed, the instructions cause a processing structure to perform the second method described above.
[0034] According to one aspect of this disclosure, one or more processors are provided, which are configured to perform the second method described above.
[0035] According to one aspect of this disclosure, a computerized third method is provided, comprising: aligning genomic data of a DNA sample to data of a reference genome via secondary genomic analysis; and performing pan-omics sequencing by comparing the aligned genomic data to a local reference dataset to identify disease-related genomic fragments.
[0036] In some implementations, the third method further includes storing the genomic data of the DNA sample and / or the identified genomic fragments in FASTQ format.
[0037] In some implementations, the third method further includes storing, in BAM format, the alignment data obtained by comparing the genomic data of the DNA sample with data from a reference genome.
[0038] In some implementations, performing pan-omics sequencing includes performing whole exome sequencing (WES) or whole genome sequencing (WGS).
[0039] According to one aspect of this disclosure, one or more non-transitory computer-readable storage devices are provided, including computer-executable instructions for compressing genome sequencing data, wherein, when executed, the instructions cause a processing structure to perform the third method described above.
[0040] According to one aspect of this disclosure, one or more processors are provided, which are configured to perform the third method described above.
[0041] According to one aspect of this disclosure, a blockchain system is provided, comprising: multiple nodes for distributed storage of genomic sequencing data of entities as transaction data within the nodes.
[0042] In some implementations, multiple nodes are configured to compress genome sequencing data using the first method described above.
[0043] In some implementations, multiple nodes are configured to perform the second method described above for distributed storage of an entity's genome sequencing data within the nodes.
[0044] In some implementations, multiple nodes are configured to perform the third method described above. Attached Figure Description
[0045] For a more complete understanding of this disclosure, please refer to the following description and accompanying drawings, in which:
[0046] Figure 1 This is a schematic diagram of a computer network system for data sharing according to some embodiments of the present disclosure;
[0047] Figure 2 It is shown Figure 1 A simplified hardware structure diagram of the computing device in the computer network system shown.
[0048] Figure 3 It is shown Figure 1 A simplified software architecture diagram of the computing device in the computer network system shown.
[0049] Figure 4 This illustrates some embodiments according to the present disclosure. Figure 1 The flowchart shown illustrates the process performed by a computer network system for storing and manipulating (such as inserting and querying) genomic data using blockchain.
[0050] Figure 5 Examples of reads, quality scores, and read IDs for genomic data are shown;
[0051] Figure 6 This illustrates, according to some implementation methods, the... Figure 1 The flowchart shown illustrates a reference-based lossless genomic data compression process performed by a computer network system to compress genomic or sequence data while retaining all information from the original file.
[0052] Figure 7 This illustrates some other embodiments according to this disclosure. Figure 1 The flowchart shown illustrates a reference-based lossless genomic data compression process performed by a computer network system to compress sequence data.
[0053] Figure 8 This illustrates some other embodiments according to this disclosure. Figure 7 The flowchart shown illustrates the incremental encoding process used in the statistical modeling step of the reference-based lossless genomic data compression process.
[0054] Figure 9 This illustrates, according to some implementation methods, the... Figure 1 The flowchart shown is a process of assembly-based lossless genomic data compression performed by the assembler of a computer network system to compress genomic or sequence data while retaining all information from the original file;
[0055] Figure 10 This illustrates an initialization process performed by a computing device according to some embodiments of the present disclosure. Figure 1 A flowchart illustrating the steps of the initialization process of a computer network system.
[0056] Figure 11 This illustrates an initialization process performed by a computing device according to some embodiments of the present disclosure. Figure 1 A flowchart illustrating the steps of the pan-omics process in a computer network system; and
[0057] Figure 12 This illustrates an initialization process performed by a computing device according to some embodiments of the present disclosure. Figure 1 A flowchart illustrating the steps of the initialization process of a computer network system. Detailed Implementation
[0058] Some of the challenges associated with genomic data sharing stem from the adoption of a centralized approach to the storage, sharing, and access of genomic data. The success of centralized data sharing largely depends on the functionality and resources of a central data storage infrastructure and / or centralized data access control services. Recent research indicates that data custodians are facing limitations in establishing centralized data access management mechanisms. Furthermore, the non-automated nature of traditional data sharing and access increases the complexity of monitoring whether data sharing and use comply with consent forms and data access agreements.
[0059] Furthermore, centralized platforms are not well-suited for fostering active participation from multiple stakeholders, such as individuals and patients, in the governance of data sharing.
[0060] Distributed networks can serve as a solution that enables approved queries to distributed, encrypted databases and allows each individual data contributor to manage data access. Therefore, data sharing based on distributed networks may be more advantageous than data sharing based on centralized third-party services, as distributed networks are designed to overcome the inefficiencies, costs, and security risks of transferring datasets to a central repository (often across international borders).
[0061] One emerging example of distributed networks is a blockchain-based platform for data sharing and access. A blockchain is a decentralized, peer-to-peer architecture with nodes that are network participants. Each member in the network stores the same copy of the blockchain and contributes to the network's collective process of verifying and authenticating digital transactions.
[0062] According to one aspect of this disclosure, a blockchain system is provided for storing and sharing genomic data (also referred to below as "genome sequencing data," "sequencing data," or "sequence data") among users. The system disclosed herein provides an alternative to traditional distributed systems with improved data security.
[0063] System Structure
[0064] Now go to Figure 1 This figure illustrates a computer network system for storing and sharing genomic data, and this computer network system is generally identified by reference numeral 100. As shown, the computer network system 100 includes one or more server computers 102 and multiple client computing devices 104, which are connected via suitable wired and / or wireless networks and functionally interconnected through networks 108 such as the Internet, local area network (LAN), wide area network (WAN), metropolitan area network (MAN).
[0065] Server computer 102 may be a computing device specifically designed for use as a server, and / or a general-purpose computing device used as a server computer but also by various users. Each server computer 102 may execute one or more server programs.
[0066] The client computing device 104 can be a portable and / or non-portable computing device, such as a laptop computer, tablet computer, smartphone, personal digital assistant (PDA), desktop computer, etc. Each client computing device 104 can execute one or more client applications, which are sometimes referred to as "apps".
[0067] Typically, computing devices 102 and 104 have similar hardware structures, such as Figure 2 The hardware structure 120 shown is illustrated. As shown, the computing device 102 / 104 includes a processing structure 122, a control structure 124, one or more non-transitory computer-readable storage devices 126, a network interface 128, an input interface 130, and an output interface 132, which are functionally interconnected via a system bus 138. The computing device 102 / 104 may also include other components 134 coupled to the system bus 138.
[0068] Processing architecture 122 can be one or more single-core or multi-core computing processors, such as Microprocessor (INTEL is a registered trademark of Intel Corp., Santa Clara, California, USA) Microprocessors (AMD is a registered trademark of Advanced Micro Devices Inc. of Sunnyvale, California, USA), manufactured by various manufacturers (such as Qualcomm of San Diego, California, USA) using Architecture Microprocessors (ARM is a registered trademark of Arm Ltd. of Cambridge, UK), etc. When the processing architecture 122 includes multiple processors, the processors may cooperate via dedicated circuitry (such as a dedicated bus) or via system bus 138.
[0069] The processing structure 122 may also include one or more real-time processors, programmable logic controllers (PLCs), microcontroller units (MCUs), μ controllers (UCs), dedicated / custom processors, and / or controllers using technologies such as field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs).
[0070] Typically, each processor of processing structure 122 includes the necessary circuitry implemented using technologies such as electrical and / or optical hardware components to execute one or more processes to perform various tasks. In many embodiments, one or more processes may be implemented as firmware and / or software stored in memory 126. Those skilled in the art will understand that in these embodiments, the one or more processors of processing structure 122 are generally useless without meaningful firmware and / or software.
[0071] The control structure 124 includes one or more control circuits, such as a graphics controller, an input / output chipset, etc., for coordinating the operation of various hardware components and modules of the computing device 102 / 104.
[0072] Memory 126 includes one or more non-transitory computer-readable storage devices or media accessible by processing structure 122 and control structure 124 for reading and / or storing instructions executable by processing structure 122, and for reading and / or storing data, including input data and data generated by processing structure 122 and control structure 124. Memory 126 may be volatile and / or non-volatile, non-removable or removable memory, such as RAM, ROM, EEPROM, solid-state memory, hard disk, CD, DVD, flash memory, etc. In use, memory 126 is typically divided into multiple parts for different purposes. For example, a portion of memory 126 (referred to herein as storage memory) may be used for long-term data storage, such as storing files or databases. Another portion of memory 126 may be used as system memory (referred herein as working memory) for storing data during processing.
[0073] Network interface 128 includes one or more network modules for connecting to other computing devices or networks, such as Ethernet, via network 108 using appropriate wired and / or wireless communication technologies. (WI-FI is a registered trademark of Wi-FiAlliance in Austin, Texas, USA) (BLUETOOTH is a registered trademark of Bluetooth Sig Inc., Kirkland, Washington, USA), Bluetooth Low Energy (BLE), Z-Wave, LoRa (long-range communication). (ZIGBEE is a registered trademark of ZigBee Alliance Corp., San Ramon, California, USA), wireless broadband communication technologies such as Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Universal Mobile Telecommunications System (UMTS), Global System for Microwave Access (WiMAX), CDMA2000, Long Term Evolution (LTE), 3GPP, 5G New Radio (5G NR), and / or other 5G networks. In some implementations, parallel ports, serial ports, USB connections, optical connections, etc., may also be used to connect other computing devices or networks, although they are generally considered to be input / output interfaces for connecting input / output devices.
[0074] Input interface 130 includes one or more input modules for one or more users to input data via, for example, a touchscreen, a touch-sensitive whiteboard, a touchpad, a keyboard, a computer mouse, a trackball, a microphone, a scanner, a camera, etc. Input interface 130 may be a physically integrated part of computing device 102 / 104 (e.g., a touchpad of a laptop computer or a touchscreen of a tablet computer), or it may be a device physically separate from but functionally coupled to other components of computing device 102 / 104 (e.g., a computer mouse). In some implementations, input interface 130 may be integrated with a display output to form a touchscreen or touch-sensitive whiteboard.
[0075] Output interface 132 includes one or more output modules for outputting data to a user. Examples of output modules include displays (such as monitors, LCD displays, LED displays, projectors, etc.), speakers, printers, virtual reality (VR) headsets, augmented reality (AR) glasses, etc. Output interface 132 may be a physically integrated part of computing device 102 / 104 (e.g., the display of a laptop computer or tablet computer), or it may be a device that is physically separate from but functionally coupled to other components of computing device 102 / 104 (e.g., the monitor of a desktop computer).
[0076] The computing device 102 / 104 may also include other components 134, such as one or more positioning modules, temperature sensors, barometers, inertial measurement units (IMUs), etc. Examples of positioning modules may be one or more Global Navigation Satellite System (GNSS) components (e.g., one or more components operating with the U.S. Global Positioning System (GPS), Russia's Global Navigation Satellite System (GLONASS), the European Union's Galileo positioning system, and / or China's BeiDou system).
[0077] The system bus 138 interconnects various components 122 to 134, enabling them to send and receive data and control signals to each other.
[0078] Figure 3 A simplified software architecture 160 of computing device 102 or 104 is shown. Software architecture 160 includes an application layer 162, an operating system 166, a logical input / output (I / O) interface 168, and logical memory 172. The application layer 162, operating system 166, and logical I / O interface 168 are typically implemented as computer-executable instructions or code, which are stored in logical memory 172 as software programs or firmware programs and can be executed by processing architecture 122.
[0079] In this document, software or firmware is a set of computer-executable instructions or code stored in one or more non-transitory computer-readable storage devices or media (such as memory 126) and readable and executable by processing structure 122 and / or other suitable components of computing device 102 / 104 for performing one or more processes. Those skilled in the art will understand that a program may be implemented as software or firmware depending on the design purpose and requirements. Therefore, for ease of description, the terms "software" and "firmware" may be used interchangeably below.
[0080] Return to reference Figure 3 The application layer 162 includes one or more applications 164 executed or performed by the processing structure 122 for performing various tasks.
[0081] Operating system 166 manages various hardware components of computing device 102 or 104, manages logical memory 172, and manages and supports application program 164 via logical I / O interface 168. Operating system 166 also communicates with other computing devices (not shown) via network 108 to allow application program 164 to communicate with programs running on other computing devices. As those skilled in the art will understand, operating system 166 can be any suitable operating system, such as... (MICROSOFT and WINDOWS are registered trademarks of Microsoft Corp. of Redmond, Washington, USA.) OS X iOS (APPLE is a registered trademark of Apple Inc. of Cupertino, California, USA), Linux, (ANDROID is a registered trademark of Google Inc. of Mountain View, California, USA.) The computing devices 102 and 104 of the computer network system 100 may all have the same operating system, or they may have different operating systems.
[0082] The logic I / O interface 168 includes one or more device drivers 170 for communicating with corresponding input interfaces 130 and output interfaces 132 to receive and send data to them. Received data can be sent to the application layer 162 for processing by one or more application programs 164. Data generated by the application programs 164 can be sent to the logic I / O interface 168 for output to various output devices (via output interface 132).
[0083] Logical memory 172 is a logical mapping of physical memory 126, used to facilitate access by application 164. In this embodiment, logical memory 172 includes a storage memory region that can be mapped to non-volatile physical memory, such as a hard disk, solid-state drive, flash drive, etc., typically used for long-term data storage therein. Logical memory 172 also includes a working memory region that is typically mapped to high-speed and volatile (in some implementations) physical memory (such as RAM), typically used by application 164 to temporarily store data during program execution. For example, application 164 can load data from the storage memory region into the working memory region, and can store data generated during its execution into the working memory region. Application 164 can also store some data into the storage memory region as needed or in response to user commands.
[0084] In server computer 102, application layer 162 typically includes one or more server-side applications 164 that provide server functions for managing network communication with client computing devices 104 and facilitating collaboration between server computer 102 and client computing devices 104. In this document, depending on the context, the term "server" can refer to server computer 102 from a hardware perspective or to a logical server from a software perspective.
[0085] As stated above, the processing structure 122 is generally useless without meaningful firmware and / or software. Similarly, while the computer system 100 may have the potential to perform a variety of tasks, it is incapable of performing any task and is useless without meaningful firmware and / or software. As will be described in more detail later, the computer system 100 described herein, as a combination of hardware and software, generally produces tangible results relating to the physical world, whereby tangible results such as those described herein can lead to improvements in the computer and system itself.
[0086] Blockchain system
[0087] In some implementations, the computer network system 100 is a blockchain system, which may be a public blockchain system that is accessible to the public, or a private blockchain system that spans one or more organizations or one or more industries.
[0088] Blockchain system 100 can be considered a decentralized and distributed database system, in which the database is managed by multiple users via any suitable computing device (also referred to as a "computer node," "network node," or "node"; such as multiple server computers 102 and / or multiple client computing devices 104). More specifically, blockchain system 100 maintains a decentralized, distributed digital ledger of transactions (the so-called "blockchain network" or simply "blockchain"). Therefore, there is no central control point in blockchain system 100, and the blockchain is stored across multiple physical computing devices 102 / 104 within the computer network system 100. Server computers 102 in blockchain system 100 typically play a similar role to client computing devices 104 in maintaining the blockchain, and in addition, server computers 102 can also be responsible for the computer network management of blockchain system 100.
[0089] The blockchain is replicated and distributed throughout the blockchain system 100 and shared peer-to-peer. The blockchain comprises multiple blocks linked by cryptographic hashes. Each block includes: an index, a timestamp, a cryptographic hash of the content of previous blocks, a hash of the block's own content, and transaction data about the network state stored at the root of a trie (a k-ary search tree; also known as a number tree or prefix tree). Furthermore, various blockchain technologies employ different cryptographic proof methods, such as proof-of-work and proof-of-stake, as part of the consensus mechanism of the blockchain technology to establish protocols regarding data values or network state between distributed processes within the blockchain system 100.
[0090] In these implementations, blockchain system 100 uses the Solana (SOL) blockchain technology developed by Solana Labs and the Solana Foundation. The state trie data structure in the Solana blockchain stores information about users and SOL contract accounts. The Solana blockchain uses proof-of-stake consensus enhanced with historical proof-of-stake.
[0091] The Solana blockchain includes conditional procedures (so-called "smart contracts") for automation, such as automating protocol execution, automating workflow execution, and automatically triggering the next action when conditions are met. A smart contract includes one or more predefined conditions, and the smart contract is executed when its predefined conditions are met. In these implementations, the smart contract's code is stored at an address in network storage, and it also maintains its own storage (such as for storing variables).
[0092] Figure 4 This is a flowchart illustrating a process 200 performed by a blockchain system 100 (or more specifically, by one or more processors 122 of the blockchain system) for using a blockchain to store and manipulate (such as insert and query) genomic data, according to some embodiments of this disclosure. For illustrative purposes, Figure 4 The example shown includes four nodes 202 (such as four client computing devices 104), each node synchronizing a copy of the blockchain.
[0093] like Figure 4 As shown, user 212 can log in to blockchain system 100 (step 214), where user credentials (such as username and password) can be stored at backend 216 (such as server 102) and managed by authorized users such as system administrators. New user 212 can also register by setting user credentials at backend 216.
[0094] After logging in, user 212 can upload their genomic data using the homepage (step 218). The uploaded genomic data is processed as transaction data, which can be converted into unique hash strings, for example, across four nodes 202. More specifically, each node 202 receives a copy of the transaction data (220A to 220D) and converts each node's copy of the transaction data into unique hash strings 222A to 222D. The hash strings 222A to 222D are then divided into one or more hash string pairs (e.g., hash string pair 222A / 222B and hash string pair 222C / 222D), and each pair of hash strings is combined (step 224) to obtain new hash strings (226A, 226B) for another security layer.
[0095] Pairings and combinations of hash strings can be repeated. For example, such as... Figure 4 As shown, the two hash strings 226A and 226B are further paired and combined (step 228) to obtain a top-level hash string 230, which can be used to retrieve genomic data stored in the blockchain. For example, genomic data of user 212 stored in blockchain system 100 can be sent to genome viewer 232 upon a viewing request from another user 242, such as a pharmacist. As another example, in the case of creating event 244 (such as needing 100 cancer reports), request 244 is sent (step 246) to homepage 218. The top-level hash string 230 can then be used to retrieve the requested genomic data from the blockchain (step 248) in response to event 244.
[0096] In some implementations, the retrieved genomic data is used only for viewing in the genomic viewer 232 or for responding to event 244 within a predefined time period (e.g., within a predefined number of days), and / or cannot be copied out of the blockchain without the permission of the owner of the genomic data.
[0097] The following text describes in detail the various technical features of the blockchain system 100.
[0098] Genomic data storage
[0099] In some implementations, genomic data is stored in a blockchain as one or more binary alignment and mapping (BAM) files. A BAM file is a binary file that stores sequence alignment data and can be used for specific analyses. BAM files are stored in a blockchain... Figure 4 The public hash string shown is associated with different transactions.
[0100] Computation of genomic data
[0101] In some implementations, the blockchain ledger uses the following for storage computation:
[0102] • All calculations are in actual bytes, i.e., 1 kilobyte (KB) = 1024 bytes.
[0103] • All ledger blocks are 1 megabyte (MB).
[0104] • Only hash, signature, or key data is stored in the blockchain ledger.
[0105] In some implementations, each block of the blockchain system 100 contains approximately 1,000 transactions. Using the parameters described above, the storage required for transactions per second (TPS) is approximately 6.75 gigabytes (GB) or 0.00659 terabytes (TiB) per transaction per year.
[0106] Genome Viewer
[0107] In some implementations, the genome viewer 232 provides the following features:
[0108] • Open genomic data in a human-readable format;
[0109] • Open and display the requested genome data, provided all licensing and date requirements are met;
[0110] • With the permission of the owner of the genomic data, download the displayed genomic data to the local computing device;
[0111] • Upload genome data; and
[0112] • Browser-based viewer.
[0113] Non-fungible tokens (NFTs)
[0114] Non-fungible tokens (NFTs) are unique, encrypted tokens linked to digital (and sometimes physical) content, serving as proof of ownership. NFTs have been used in many areas, such as artwork, digital collectibles, music, items in video games, and more.
[0115] In some implementations, the blockchain system 100 can store genomic data as one or more NFTs. For example, the blockchain system 100 can use an avatar with a background as a unique transaction identifier (ID) for a user to upload genomic data and create a new NFT (associated with a unique transaction ID). A user's photograph can be used to uniquely identify the user at the front end.
[0116] Participatory Access Control
[0117] One of the main challenges in managing genomic data sharing is access control. Because genomic data can reveal sensitive health- and non-health-related personal information, it is crucial to employ privacy-preserving mechanisms when sharing data. Traditionally, access control has been managed in a non-automated manner, primarily through access committees that vet data users and grant access to specific datasets or data available in databases. However, major drawbacks associated with this controlled access model include a lack of coordinated access policies, cumbersome or bureaucratic access procedures, resource-intensive monitoring, and a lack of appropriate tools for ongoing oversight.
[0118] In some implementations, blockchain system 100 uses the “permissioned” structure of blockchain to manage and control access to sensitive genomic data in an automated manner, wherein only pre-approved users are allowed to access the genomic data, and access to the genomic data by unapproved or unauthorized users is prohibited or denied. More specifically, blockchain system 100 provides metadata of the genomic dataset (which is a description of the genomic data rather than the genomic data itself) and allows all users of blockchain system 100 to access the metadata. On the other hand, blockchain system 100 uses pseudo-anonymity and public key infrastructure (PKI) to encrypt the content of the blockchain (including genomic data) and allows only authorized users to access it.
[0119] Therefore, blockchain system 100 allows for the discoverability of existing genomic datasets while protecting individual privacy by restricting access to actual genomic data to authorized users. Furthermore, to ensure the anonymity of genomic datasets when necessary, blockchain system 100 can utilize appropriate tools such as Software Protection Extensions (SGX) to achieve efficient and secure data storage and computational outsourcing using multiple cryptographic protocols.
[0120] In some implementations, the blockchain system 100 may also use other suitable methods, such as keeping sensitive data off-chain, to further protect the privacy of sensitive data.
[0121] In some implementations, the blockchain system 100 can manage access control participatoryly, where individuals (i.e., genomic data subjects), healthcare professionals, and researchers can participate in managing genomic data access controls. In these implementations, the blockchain system 100 provides the necessary infrastructure where various stakeholders can directly participate in the management of genomic data access. Therefore, the blockchain system 100 creates opportunities for maintaining the property of patients (i.e., genomic data subjects) with medical information and allows individuals to opt in or out of specific events (such as research and learning).
[0122] In some implementations, the blockchain system 100 can use the patient's biometric data as a private key, requiring other users to obtain the patient's approval and authorization before using their anonymized health data.
[0123] Data ownership and DNA coin
[0124] As understood by those skilled in the art, there is growing support among academics, patient groups, and the public for recognizing individual ownership of raw genomic data and for facilitating access to raw genomic data and related medical records.
[0125] In some implementations, the blockchain system 100 allows individuals to maintain ownership of their personal genomic and health-related data and decide how and under what conditions to share their data. Therefore, individuals can share their genomic data for a variety of purposes.
[0126] Therefore, blockchain systems 100 ensure that users can share their genomic data using privacy-preserving methods that promote compliance with legal and ethical standards, which could be beneficial in providing multiple approaches to address the governance challenges in genomic data sharing.
[0127] Therefore, Blockchain System 100 is capable of managing open networks where the potential of decentralized networks, industry demands, and consumer genetics is being explored. Blockchain System 100 aims to expand the volume of data while providing various ownership models and promoting active individual participation in data-sharing governance.
[0128] Specifically, blockchain system 100 can automate the data access control process and improve the transparency and fairness of genomic data access. Similarly, by using smart contracts of blockchain system 100, the enforceability of access agreements can be significantly improved, which can provide guarantees for different users such as researchers and data custodians that downstream data use can comply with the terms and conditions of data use.
[0129] Furthermore, leveraging blockchain systems can strengthen the role of patients and individuals in the data-sharing ecosystem and weaken the monopoly of public and private testing providers on the management of genomic data sharing. Blockchain has the potential to create new public resources, occupying the space between markets and public goods.
[0130] Compression of reference genomic data
[0131] As those skilled in the art will understand, genomic data is typically enormous. With the development of next-generation sequencing (NGS) technology, researchers have had to adapt rapidly to cope with the dramatic increase in raw genomic data. Experiments previously conducted with microarrays and producing several megabytes of data now generate many gigabytes when performed with sequencing, requiring significant investments in computing infrastructure. While the cost of disk storage has steadily decreased over time, it has not kept pace with the enormous changes in sequencing costs and data volumes. Storage technology may see revolutionary breakthroughs in the coming years, but the era of $1,000 genomes will inevitably precede the era of $100 petabyte hard drives.
[0132] As cloud computing and Software as a Service (SaaS) become increasingly integrated into molecular biology research, the time spent transferring NGS datasets round trip to remote servers for analysis delays meaningful results. More often than not, researchers are forced to maximize bandwidth by physically transporting storage media (so-called "sneakernets"), an expensive and logically complex option. With sequencing data growing exponentially, these difficulties will only intensify, meaning that even modest gains in domain-specific compression methods will translate into a significant reduction in the cost of managing these massive datasets over time.
[0133] Therefore, due to the massive scale of genomic data, compression techniques play a crucial role in achieving efficient storage and transmission of genomic data. However, traditional general-purpose data compression methods and procedures (such as Gzip) may not fully utilize the inherent redundancy in genomic data and thus may not achieve high compression ratios. Furthermore, in many cases, genomic data can be noisy, and lossy compression methods that can reduce storage space can be used without significantly adversely affecting the data quality for downstream analysis.
[0134] NGS data center storage and analysis primarily revolve around two formats that have recently emerged as de facto standards: the FASTQ format and the Sequence Alignment Map (SAM) format. In addition to the nucleotide sequence, FASTQ files store a unique ID for each read (represented as a "read ID"; it is the sequence of a fragment of DNA, i.e., a small segment of DNA) and a quality score, which encodes an estimate of the probability that each base is correctly called. The SAM format is much more complex but also more rigorously defined, and a reference implementation is provided in the form of SAMtools. Besides read IDs, sequences, and quality scores, it is also capable of storing alignment information. SAM files stored in plain text can also be converted to BAM format, a compressed binary version of SAM that is more compact and allows for relatively efficient random access. Figure 5 Examples of read segments, quality scores, and read segment IDs are shown.
[0135] Nucleotide sequence compression has long been a subject of interest, but compressing NGS data, which consists of millions of short fragments in a larger whole, combined with metadata in the form of read IDs and quality scores, presents a very different problem and requires new techniques. Breaking the data down into separate contexts of read IDs, sequences, and quality scores, and developing compression methods utilizing the Lempel-Zip algorithm and Huffman coding, demonstrates promise for domain-specific compression, offering significant improvements over general-purpose programs such as gzip and bzip2.
[0136] Reference-based compression methods leverage the redundancy of data by aligning reads to a known reference genome sequence and storing the genomic location instead of the nucleotide sequence. Decompression is then performed by copying the read sequence from the genome. Although any differences from the reference sequence must also be stored, reference-based methods achieve much higher compression ratios and are more efficient for long reads because the amount of space required to store the genomic location is the same, regardless of the read length.
[0137] For some applications, reference-based compression can be improved by storing only single nucleotide polymorphism (SNP) information, allowing sequencing experiments to be compiled into a few megabytes. However, discarding the original reads will prevent any reanalysis of the data.
[0138] While reference-based methods generally achieve better compression, they have several drawbacks. Most notably, a suitable reference sequence database is not always available, especially in the case of metagenomic sequencing. A reference sequence database can be constructed by compiling a collection of genomes from species expected to be present in the sample. However, a high level of expertise is required to curate and manage such a project-dependent database. Second, a practical problem is that files compressed using reference-based methods are not self-contained. Decompression requires the exact same reference database used for compression, and if the reference database is lost or forgotten, the compressed data becomes inaccessible.
[0139] Another approach to short read compression is lossy encoding of sequence quality scores. This stems naturally from the recognition that quality scores are particularly difficult to compress. Unlike highly redundant read IDs or nucleotide sequences containing certain structures, quality scores are encoded inconsistently across protocols and computational pipelines and often exhibit only high entropy. Disappointingly, metadata (such as quality scores) should take up more space than master data (such as nucleotide sequences). However, equally displeasing to many researchers is the idea of discarding information without a full understanding of its impact on downstream analysis.
[0140] Reducing the entropy of the quality score while maintaining accuracy is an important goal. However, successful lossy compression requires understanding what is being lost. For example, lossy audio compression (such as MP3) is based on psychoacoustic principles, prioritizing the discarding of the least perceptible sounds. Given that both the algorithms that generate NGS quality scores and the algorithms affected by them are moving targets, it is difficult to devise a similar principled approach for NGS quality scores.
[0141] The following describes a reference-based lossless compression method that utilizes various techniques to achieve very high compression rates for many types of sequencing data while remaining efficient and practical. Given aligned reads in SAM or BAM format and their corresponding reference sequences (FASTA format), the reference-based lossless compression method compresses the reads while preserving all information in the SAM / BAM file (including header information, read IDs, alignment information, and all optional fields allowed in the SAM format). Unaligned reads are preserved and compressed using a Markov chain model.
[0142] As will be described in more detail later, reference-based lossless compression methods store only unaligned reads, not the entire genome. Using initial compression based on statistical modeling and statistical compression methods, reference-based lossless compression can effectively compress very large sequence data to less than 15% of their original size without losing information.
[0143] Figure 6This is a flowchart illustrating a reference-based lossless genomic data compression process 300 performed by a blockchain system 100 (or more specifically, by one or more processors 122 of the blockchain system) according to some embodiments, for compressing genomic or sequence data while retaining all information from the original file. Sequence data can be stored in any suitable format, such as in FASTQ and SAM / BAM formats (e.g., as NGS data in FASTQ and SAM / BAM formats).
[0144] After the process begins (step 302), the blockchain system 100 compares the sequence data with reference sequence data (step 304) and obtains one or more differential read sequences, which include sequence data that differs from the reference sequence data (step 306). The differential read sequences are then stored as locations within an assembled contiguous group. In this document, a contiguous group is a collection of overlapping DNA fragments or sequences that provide a continuous representation of a genomic region.
[0145] In step 310, a statistical compression method is used to compress the read ID, quality score, alignment information, read sequence, etc. of one or more differential read sequences. Then process 300 ends (step 312).
[0146] In some implementations, the statistical compression method used in process 300 is arithmetic coding, which is a form of entropy coding close to optimal. Typically, when encoding strings using arithmetic coding, characters that appear more frequently (or have a higher probability of occurrence) are stored with fewer bits, while characters that appear less frequently (or have a lower probability of occurrence) are stored with more bits, resulting in a reduction in the total number of bits. Unlike Huffman coding (another form of entropy coding), arithmetic coding methods have the advantage of being able to allocate codes of non-integer lengths. For example, if a symbol has a probability of occurrence of 0.1, that symbol can be encoded with a code length close to its optimal length -log2(0.1) ≈ 3.3 bits.
[0147] Figure 7 This is a flowchart illustrating a reference-based lossless genomic data compression process 300 for compressing sequence data, performed by a blockchain system 100 (or more specifically, one or more processors 122 of the blockchain system) according to some other embodiments of this disclosure. The process 300 in these embodiments is similar to... Figure 6The process is illustrated in the diagram. However, by leveraging the advantage of arithmetic coding allowing for complete separation between statistical modeling and coding, process 300 in these embodiments also includes a statistical modeling step 308 prior to coding step 310. In other words, the blockchain system 100 processes the differential read sequence using a corresponding statistical modeling method for initial compression of the sequence data (step 308), and then further compresses the sequence data using a statistical compression method (step 310), thereby achieving an increased compression ratio.
[0148] In some implementations, the blockchain system 100 uses the same statistical compression method (such as the same arithmetic encoder) to encode the various fields of the sequence data (such as quality score, read ID, nucleotide sequence, alignment information, etc.), and the blockchain system 100 can use different statistical models to perform initial compression on different fields.
[0149] Those skilled in the art will understand that, in various implementations, each field of the sequence data may not necessarily require initial compression using a different statistical model. Instead, depending on the implementation, some fields of the sequence data may be initially compressed using a different statistical model at step 308, some other fields of the sequence data may be initially compressed using the same statistical model at step 308, and / or some still other fields of the sequence data may not be initially compressed at step 308.
[0150] By employing statistical modeling at step 308 and statistical compression at step 310, process 300 achieves significant advantages over generic compression methods that concentrate all content into a single context. In some implementations, the statistical model is an adaptive model with parameters trained and updated as the data is compressed, enabling increasingly tighter fits and high compression ratios on large files.
[0151] Statistical modeling for read segment identifiers
[0152] A read segment ID uniquely identifies a read segment. While integers can be used as read segment IDs, each read segment is typically associated with a complex string containing the instrument name, operation identifier, flow cell identifier, and tile coordinates. Much of this information is the same for each read segment and is simply repeated, introducing redundancy.
[0153] In some implementations, incremental coding methods can be used to remove this redundancy. Figure 8 This is a flowchart illustrating the incremental encoding process 340.
[0154] After process 340 begins (step 342), the parser (which may be an application or program module 164) parses the read segment ID into one or more individual fields (step 344) and stores the fields of the read segment ID (step 346). The parser then parses another read segment ID into fields (step 348) and compares the fields of the current read segment ID with the fields of the previous read segment ID (i.e., fields in the same position) (step 350). At step 352, dissimilar fields (i.e., fields with values different from the previous read segment ID) are stored, while only identical fields (i.e., fields with values identical to the previous read segment ID) are marked. The parser then checks whether all read segment IDs have been processed (step 354). If all read segment IDs have been processed, process 340 ends; otherwise, process 340 returns to step 348 to parse another read segment ID.
[0155] At step 352, non-numerical fields (i.e., non-numerical fields with different values) can be efficiently stored directly or stored as an offset from the previous read segment ID. Non-numerical fields (i.e., non-numerical fields with different values) can be encoded by dividing the non-numerical field into a common part such as a common prefix (i.e., the part that is the same as the corresponding field of the previous read segment ID) and a non-dissimilar part such as a non-dissimilar suffix, and the non-dissimilar suffix is stored.
[0156] By using incremental and arithmetic coding methods, fields that remain the same from read to read (e.g., instrument name) can be compressed to a small or even negligible amount of space (e.g., for such fields, codes less than one (1) bit). Therefore, read segment IDs, typically 50 bytes or longer, are usually stored in two (2) to four (4) bytes after compression. It is worth noting that, from... The instrument (Illumina is a registered trademark of Illumina, Inc., San Diego, California, USA) generates reads in which most of the read ID can be compressed to consume almost no space; the remaining few bytes are occupied by tile coordinates that are almost never needed in downstream analysis, and removing them as a preprocessing step can further reduce the size of the compressed sequence data.
[0157] Statistical modeling of nucleotide sequences
[0158] In some implementations, statistical models based on higher-order Markov chains (such as 12-order Markov chains) are used to compress the nucleotide sequences of differentially read sequences, where, for example, the first 12 positions can be used to predict the nucleotides at a given position in the read. While the statistical models used in these implementations may require more memory than conventional general compression methods (e.g., possibly 4), the underlying technologies are not universally applicable. 13=67,108.864 parameters, each represented by 32 bits), but the statistical model is simple and extremely efficient (requiring very little computation, and the runtime is mainly limited by memory latency, as lookups in such a large table lead to frequent cache misses).
[0159] While a 12th-order Markov chain may not be well-suited for compressing extremely short files, after compressing multiple reads (such as millions of reads), the parameters of the 12th-order Markov chain can closely fit the nucleotide composition of the sequence data, allowing the remaining reads to be highly compressed. Therefore, compressing large files can achieve both close fit and high compression ratios.
[0160] Statistical modeling for quality scores
[0161] The quality score at a given position is highly correlated with the score at the previous position. Therefore, in some implementations, Markov chains are used as statistical models for initial compression of quality scores. However, unlike nucleotides, quality scores are on a much larger alphabet (typically 41 to 46 distinct scores), which limits the order of Markov chains because long chains require a large amount of space and impractical amounts of data for training.
[0162] In some implementations, to reduce the number of parameters, a 3rd-order Markov chain is used and the two far-end positions are coarsely binned or stored. In addition to the first three positions, the blockchain system 100 also uses large jumps in quality scores between positions within a read segment and between adjacent positions (where a "large jump" is defined as |q...). i -q i-1 The number of runs (>1) is conditional, which allows for the use of a separate model to encode reads with highly variable quality scores. Both variables are binned or stored to control the number of parameters.
[0163] Compression of assembled genomic data
[0164] In some implementations, the blockchain system 100 can compress genomic data using a de novo assembly method (i.e., an assembly method without reference data). This method uses probabilistic data structures to significantly reduce the memory required by traditional de Bruin graph assemblers, thereby allowing for the highly efficient assembly of millions of reads. The read sequences are then stored as positions within the contiguous group of the assembly. This, combined with statistical compression of read IDs, quality scores, alignment information, sequences, etc., effectively compresses very large datasets to less than 15% of their original size without loss of information. Compared to reference-based compression methods, the assembly-based compression method disclosed herein does not require an external sequence database and generates fully self-contained files.
[0165] Roughly speaking, the assembly-based compression method disclosed in this paper can be considered similar to the Lempel-Ziv algorithm, where, when reading a sequence of segments, they are matched with previously observed data, which can be contigs assembled from the previously observed data (these contigs are not explicitly stored but reassembled during decompression).
[0166] Figure 9 This is a flowchart illustrating an assembly-based lossless genomic data compression process 400 performed by an assembler (which may be an application or program module 164) of a blockchain system 100 (or more specifically, one or more processors 122 of the blockchain system) according to some implementations, for compressing genomic or sequence data while retaining all information from the original file. Sequence data can be stored in any suitable format, such as in FASTQ and SAM / BAM formats (e.g., as NGS data in FASTQ and SAM / BAM formats).
[0167] After the process begins (step 402), the blockchain system 100 uses multiple reads (e.g., the first 2.5 million reads) to assemble an overlap group, and then uses the overlap group as a reference sequence to compactly encode the aligned reads as positions (step 404).
[0168] Once the contiguous group is assembled, the read sequence is aligned using a simple “seed-expansion” method (step 406). At this step, a hash table is used to match the 12-mer seed, and candidate alignments are evaluated using Hamming distance. The best alignment is selected (e.g., the one with the lowest Hamming distance), which may be below a given cutoff. The read is then encoded as its position within the contiguous group set.
[0169] This alignment method is simple and fast, and works on platforms where erroneous insertions or deletions (indels) do not occur frequently (such as Illumina). After alignment, the remainder of process 400 can be similar to... Figure 6 The process shown (including steps 306 and 310) or Figure 7 The process shown includes steps 306, 308, and 310.
[0170] Traditionally, de novo assembly is computationally intensive. The most common technique involves constructing a de Bruin graph, a directed graph where each vertex represents a k-mer of nucleotides present in the data for a given k (e.g., k = 25). In this paper, a k-mer is a substring of length k in a given string (such as DNA, RNA, protein, or any string sequence).
[0171] A directed edge from a k-cluster u to v occurs if and only if the (k-1)-cluster suffix of u is also a prefix of v. In principle, given such a graph, an assembly can be generated by finding Eulerian paths (i.e., paths that run exactly once along every edge in the graph). In practice, due to the non-negligible error rate of NGS data, the assembler increments each vertex by the number of observed k-cluster occurrences and uses various heuristics to leverage these counts to filter out spurious paths.
[0172] A significant bottleneck of the de Bruin graph method is constructing the implicit representation of the graph by counting and storing the occurrences of k-clusters in a hash table. The assembler implemented in Quip largely overcomes this bottleneck by counting k-clusters using a Bloom filter-based data structure (a probabilistic data structure based on hashing). A Bloom filter is a probabilistic data structure that represents a set of elements in a very compact way, at the cost of elements occasionally colliding and being incorrectly reported as existing in the set. These collisions occur pseudo-randomly in a probabilistic sense, determined by the size of the table and the chosen hash function, but are typically low-probability.
[0173] Bloom filters are generalized to count Bloom filters, where any count can be associated with each element. The d-left count Bloom filter (dlCBF) is an improvement on the count Bloom filter, requiring significantly less space to achieve the same false positive rate.
[0174] In some implementations, the assembler of the blockchain system 100 is based on an implementation of dlCBF. Because the assembler uses a probabilistic data structure, k-clusters may occasionally be reported with incorrect (overstated) counts. While these incorrect counts may make the assembly less accurate, poor assembly will only result in a slightly reduced compression ratio. Regardless of the assembly quality, compression remains lossless, and in practice, collisions in dlCBF occur at a very low rate.
[0175] Given a probabilistic de Brouin graph, the assembler of Blockchain System 100 uses a simple and efficient greedy method to assemble contigs. A read sequence is used as a seed, and the system repeatedly searches for the richest k-mer that overlaps the ends of the contig by k-1 bases, expanding by one nucleotide at each end. More complex heuristics can also be used.
[0176] Memory efficiency
[0177] Efficient memory assembly is a target of particular interest and a subject of ongoing research. In the prior art, efficient means have been developed for representing de Bruin graphs using sparse bitmaps compressed with Elias-Fano encoding (an encoding method developed by Peter Elias and Robert Mario Fano in the 1970s). String graph assemblers rely on FM indexes (compressed full-text substring indexes based on the Burrows-Wheeler transform) to construct a compact representation of a set of short reads, generating overlap groups from this set by searching for overlaps. Both of these approaches improve space efficiency at the expense of time efficiency (significantly longer runtime than traditional assemblers).
[0178] The assembly-based compression method disclosed in this paper does not use precise representations, but instead relies on probabilistic data structures. This method uses Bloom filters to store k-clusters that appear only once, thereby reducing the memory required for hash tables.
[0179] Metadata
[0180] In some implementations, the sequence data compression method disclosed herein, used by the blockchain system 100, provides several useful features for protecting data integrity. First, the output of the sequence data compression method (also referred to as an "archive") is divided into blocks, each several megabytes in size. Within each block, a separate 64-bit checksum is calculated for the read ID, nucleotide sequence, and quality score. When the archive is decompressed, these checksums are recalculated on the decompressed data and compared with the stored checksums to verify the correctness or integrity of the archive.
[0181] Besides data corruption, reference-based compression methods are susceptible to data loss if the reference used for compression is lost or an incorrect reference is used. To prevent reference loss, the archive file of blockchain system 100 (i.e., the compressed sequence data file) stores a 64-bit hash of the reference sequence, ensuring that the same sequence is used for decompression. To aid in locating the correct reference, the filename, as well as the length and name of the sequence used in compression, are also stored without compression, making them accessible without decompression.
[0182] In addition, the block header stores the number of compressed reads and bases in the block without compression, allowing summary statistics of the sequence dataset to be listed without decompression.
[0183] FASTQ compression mode
[0184] In some implementations, blockchain system 100 supports the following recommended FASTQ compression modes:
[0185] 1. Lossless mode (default): In this mode, the FASTQ file is compressed so that it can be accurately reconstructed, that is, the read segments, quality values, read segment identifiers and read segment order information can all be perfectly recovered.
[0186] 2. Recommended Lossy Mode: In this mode, information relevant to most genome applications (such as alignment, assembly, variant detection, etc.) is preserved. This includes reads, pairing information, and binning quality values. Quality values undergo Illumina's normalized 8-level binning before compression (https: / / www.illumina.com / documents / products / whitepapers / whitepaper_com / compression.pdf) (Novaseq quality remains unchanged). Read identifiers and pair order are discarded (i.e., decompressed FASTQ files contain read pairs in any order). The relative order of the first and second reads within each pair is still preserved. Sequence data compression can be highly customized based on user needs and offers additional capabilities such as customized binning using QVZ (a lossy compression method) and binary thresholding of quality values. For short reads (up to 511 bp), read compression can be based on the Hash Read Compressor (HARC), which offers significant improvements and increased support for variable-length reads. Sequence data compression also supports long read compression, with the Block Sorting Compressor (BSC; https: / / github.com / IlyaGrebnov / libbsc / ) used as the read compressor. Furthermore, the sequence data compression method disclosed in this paper compresses the stream in blocks, allowing for fast decompression of subsets of reads (random access).
[0187] (1) Preprocessing
[0188] In long read mode, read segments, quality values, and read segment identifiers are separated and compressed into blocks (read segments and quality values use BSC, and identifiers use the dedicated identifier compressor mentioned above). By default, the block length for long reads is set to 10,000 read segments. The read segment length is also stored as a 32-bit integer in a separate stream, which is compressed using BSC. Preprocessing is followed directly by the Tar stage for long reads. In order-preserving mode, quality values and read segment IDs are compressed into blocks (quality values use BSC, and identifiers use the dedicated identifier compressor). If the corresponding flag is specified, QVZ quantization is applied before quality compression. By default, the block length for short reads is set to 256,000 read segments. After separating read segments containing the character "N", the read segments are written to a temporary file. In the "Encode Read Segment" stage, read segments containing "N" are considered directly. In non-order-preserving mode, quality values and read segment IDs are written to a temporary file. Read segments are processed in exactly the same way as in order-preserving mode. If the Illumina binning or binary threshold flag is activated, the quality is binned before compression / writing to temporary files.
[0189] (2) Reorder the reading segments
[0190] This step in read compression is based on HARC and has been extended and improved in several ways. In this step, SPRING reorders the reads so that they are approximately ordered according to their position in the genome. The reordering is performed iteratively: given a current read, the sequence data compression method disclosed in this paper attempts to find reads that match the prefix or suffix of the current read with a small Hamming distance. To do this efficiently, a hash table is used, which indexes the reads based on certain substrings of the reads. The sequence data compression method disclosed in this paper makes the following improvements to this stage:
[0191] HARC searches for matching reads in only one direction (matching the suffix of the current read). In contrast, the sequence data compression method disclosed in this paper searches for matches in both directions. This improves read compression efficiency by 5% to 10% on most datasets.
[0192] While HARC only supports fixed-length reads with a maximum length of 255, the sequence data compression method disclosed in this paper adds support for variable-length short reads with a maximum length of 10¹¹. To this end, the sequence data compression method disclosed in this paper stores an array containing read lengths, which is used to ensure the correct calculation of the Hamming distance between reads of different lengths.
[0193] As those skilled in the art will understand, a significant portion of the time in the reordering phase can be spent on a small fraction of the remaining reads, and attempts to find matches for these reads typically fail. To save time in this step, the sequence data compression method disclosed herein introduces early stopping at this stage. Each thread maintains a proportion of unmatched reads in the last (1) million reads and stops searching for matches once this proportion exceeds a certain threshold (e.g., 50%). Since this stage is the most time-consuming step in the sequence data compression method disclosed herein, early termination can reduce compression time by up to 20% without affecting the compression ratio.
[0194] (3) Encode the reading segment
[0195] In this step, a majority-based reference sequence is obtained using the sequence of reordered reads. The reordered reads are then encoded using the reference sequence. The final encoding includes the reference sequence, the position of the read in the reference sequence, and any mismatches of the reads relative to the reference sequence. An index mapping of the reordered reads to their positions in the original FASTQ file is also stored. The sequence data compression method disclosed herein provides support for variable-length reads up to 1055. This stage produces an encoded stream of the majority-based reference sequence and reads aligned to the reference. A small subset of reads typically remains unaligned with the reference and is stored separately. However, the encoded stream does not correspond to the original order of the reads in the FASTQ file. Furthermore, the reordering and encoding stages of the sequence data compression method disclosed herein treat paired-end FASTQ files as a single-end FASTQ file obtained by concatenating two files. Therefore, for both order-preserving and non-order-preserving modes, these streams can be transformed using information from the index mapping of the reordered reads to their positions in the original file. This is accomplished in the next two steps.
[0196] (4) Paired-end sequential encoding
[0197] This step is only used in non-order-preserving mode. A new read segment order is generated that preserves pairing information while achieving optimal compression. This step generates an index mapping between the reordered read segments and their positions in the new order. Read segments in file 1 (see...) Figure 5 The reads in file 1 are ordered in the same order as those obtained in the previous stage (encoding the reads), i.e., they are ordered according to their positions in the majority-based reference. This allows incremental coding to store the positions of these reads in the majority-based reference, resulting in improved compression. The read order in file 2 (see...) Figure 5The order of read segments is automatically determined by the order in file 1 (because pairing information is preserved). For single-ended files (in non-order-preserving mode), the read segments maintain the same order as obtained after the encoding stage (i.e., sorted according to their position in the majority-based reference).
[0198] (5) Reorder and compress the read stream
[0199] In this step, the final encoded stream is generated using BSC and compressed in blocks. To do this, the stream generated by the encoding stage is first loaded into memory, and then reordered according to the mode. In order-preserving mode, the stream is ordered according to the original order of the read segments in the FASTQ file. In non-order-preserving mode, the stream is ordered according to the new order generated in the paired-end sequence encoding step. The final stream is described as follows:
[0200] • seq: The seq storage is based on a majority reference sequence. This is packed into a two (2) bit / basic representation before compression.
[0201] • Flag: Indicates whether the reading segments have been compared and the distance between them on the reference.
[0202] ◆ Flag set to zero (0): Both reads have been aligned and the gap between the alignment positions is less than 32,767 (for single-end datasets, flag 0 means that the read has been aligned).
[0203] ◆ The flag is set to one (1): Both reading segments have been compared and the gap between the comparison positions is greater than or equal to 32,767.
[0204] ◆ The flag is set to two (2): neither read segment is matched (for a single-end dataset, flag 2 means that the read segment is not matched).
[0205] ◆ The flag is set to three (3): Segment 1 in a pair has been matched, while segment 2 has not been matched.
[0206] ◆ The flag is set to four (4): Segment 1 in a pair has not been matched, while segment 2 has been matched.
[0207] • pos: In order-preserving mode, pos uses eight (8) bytes to store the position of the first read segment (and possibly the second read segment) in the pair on the reference. If the flag is zero (0) or three (3), only the position of the first read segment is stored. If the flag is one (1), the positions of both the first and second read segments are stored. If the flag is two (2), nothing is stored. In non-order-preserving mode, the position of the first read segment in the pair is stored as the difference between the first read segment of the previous pair (except the first pair in the block). Note that this difference is always positive due to the way the new order is defined in the pairwise sequential encoding step. As long as this difference is less than 65,535, it is stored as a two (2)-byte unsigned integer; otherwise, 65,535 is first stored in two (2) bytes, followed by the actual difference in eight (8) bytes. Storing the difference instead of the absolute position allows SPRING to achieve better compression in non-order-preserving mode.
[0208] • pos pairs: For paired-end datasets, when the flag is zero (0), pos pairs store the gap between paired reads on the reference using a 16-bit signed integer. Since paired reads are sequenced from adjacent parts of the genome (paired reads are typically 50 to 250 bases apart), they are likely to be located close to each other in the reference. Processing the gaps between paired reads using separate streams allows us to take advantage of this fact.
[0209] Noise: Noise stores bases in the aligned reads relative to the reference. Encoding depends on both the reference and the bases in the read, allowing the use of the fact that certain errors are more likely to occur in Illumina sequencing. For example, the most likely transition for each reference symbol is encoded as zero (0), the next most likely transition is encoded as one (1), and so on. This results in more 0s in the encoded stream, leading to better compression. Newline characters separate noise from consecutive reads.
[0210] • noisepos: noisepos stores the positions of encoded noise bases in the noise stream. These positions are incrementally encoded to take advantage of the fact that sequencing errors mostly occur at the ends of reads. The incrementally encoded noise positions are stored as 16-bit unsigned integers.
[0211] • RC: RC stores the direction (forward / backward) of the paired read segment relative to a reference. If the flag is zero (0), RC does not store the direction of the second read segment in the pair (see RC convection).
[0212] • RC Pairs: For paired end datasets, when the flag is zero (0), RC pairs store the relative orientation of the second read with respect to the first read. If the paired reads have opposite orientations, zero (0) is stored; otherwise, one (1) is stored. Since the paired end reads have opposite genomic orientations, we expect to obtain mostly 0s in this stream, and therefore the stream is highly compressible.
[0213] • Unaligned: Unaligned stores unaligned read segments that have not been encoded.
[0214] • Length: The length is stored as a 16-bit unsigned integer.
[0215] Blockchain-based panomics
[0216] In some implementations, blockchain system 100 is a panomics system, such as the Anryton panomics system from BioAro in Calgary, Alberta, Canada, used for panomics analysis (such as analyzing genomic data) using the computing devices of blockchain system 100 (such as multiple server computers 102 and / or multiple client computing devices 104) to identify genome-related diseases. As those skilled in the art will understand, panomics (also known as “multi-omics”) is a class of bioanalytical techniques used to analyze various “groups” such as genomes, proteomes, transcriptomes, epigenomes, metabolomes, microbiomes, etc. Panomics typically requires large amounts of data (so-called “big data”) of the analyzed groups and may require significant computing power for analysis. In these implementations, blockchain system 100 may distribute the computational load among its various computing devices (such as multiple server computers 102 and / or multiple client computing devices 104) and use the aforementioned blockchain to link and / or combine analytical results from these computing devices.
[0217] For example, in these embodiments, the blockchain system 100 uses next-generation sequencing (NGS) to perform deep sequencing of DNA. Multiple sequencers in the blockchain system 100 first perform a preliminary analysis, in which bases are "called," i.e., read and reported in the output FASTQ file. Here, each sequencer can be a computing device 102 and / or 104 used by pan-omics professionals or general users.
[0218] Then, variant annotation and classification steps are performed (i.e., tertiary analysis), followed by a secondary analysis pipeline that combines the tertiary analyses. Analysis results from multiple sequencers are linked or combined via a blockchain to output a report in a suitable format (such as PDF) showing the patient's DNA sequencing results, their mutations, and their meaning. This PDF report is then stored on the blockchain and shared based on the user's authentication settings.
[0219] In some implementations, blockchain system 100 uses a variant call format (VCF) to retrieve genomic data files from the blockchain to a user. As those skilled in the art will understand, a VCF file is a text file with a VCF header that stores information about the VCF itself, such as metadata about annotations within the VCF file. Following the VCF header, the VCF file includes a table section with multiple columns (such as nine (9) columns), including CHROM, POS, ID, REF, ALT, QUAL, FILTER, INFO, and FORMAT columns for each sample listed in the VCF file. All this information can be encrypted and accessed by the user based on appropriate authentication.
[0220] Modern healthcare is a data-intensive field, representing the fusion of long-term electronic medical records, real-time patient monitoring data, and, more recently, sensor data from wearable computing. Genomics is a specialized field of biology that involves the structure, function, evolution, and editing of genomes. The human genome consists of approximately three (3) billion base pairs, equivalent to approximately 1.5 GB of data. Therefore, the blockchain system 100 disclosed in this paper is useful in genome analysis and, more broadly, in pan-omics technologies.
[0221] The decentralized nature of the blockchain system 100 disclosed herein allows for easy and secure data exchange between organizations. The blockchain system 100 also provides a platform for direct interaction between data providers (such as patients) and purchasers (such as pharmaceutical companies, research institutions, etc.). Furthermore, the blockchain system 100 also provides complete control over operation and management.
[0222] From a data security perspective, healthcare data is highly sensitive and important. The blockchain system 100 disclosed herein provides excellent data security and integrity. For example, in some implementations, the blockchain system 100 disclosed herein uses a special hashing technique to securely store data, which helps protect the data and also facilitates data sharing.
[0223] Figure 10This is a flowchart illustrating the steps of an initialization process 500 performed by a user, such as an administrator, on computing device 102 or 104 to initialize the blockchain system 100. As shown, the user installs pan-omics software on computing device 102 or 104 (step 502). The user then uploads reference genome data to computing device 102 or 104 (step 504) and downloads relevant datasets from remote data sources (such as public, research, and / or pharmaceutical databases) to the local reference library of computing device 102 or 104 (step 506). The user also uploads relevant report templates to computing device 102 or 104 (step 508) and sets default report templates and / or databases (step 510).
[0224] After initialization process 500, computing device 102 or 104 can execute as follows: Figure 11 The pan-omics analysis process shown is 550.
[0225] At step 552, a secondary genomic analysis is performed on the DNA sample (converting data into results such as alignment and expression) to align with a reference genome (or various types of microbiome) obtained at step 504 of the initialization process 500, and DNA fragments providing the complete sequence of the sample are assembled.
[0226] Then, pan-omics sequencing is performed to identify genomes associated with a specific disease (step 554). At this step, whole-exome sequencing (WES) (step 556) or whole-genome sequencing (WGS) (step 558) can be performed. At step 554 (or more specifically, steps 556 or 558), the aligned genomic data is compared with a dataset from a reference library. Since the reference library has been downloaded to computing device 102 or 104 from various sources such as public, research, and / or pharmaceutical databases at step 506 of the initialization process 500, the execution of step 558 can be performed by simply accessing the local reference library without downloading any data from a remote data source, thus accelerating the analysis.
[0227] At step 554, pan-omics sequencing can use templates for various sequencing tests. As mentioned above, such templates have been uploaded to the blockchain system 100 at step 508 of the initialization process 500. Furthermore, step 558 can also utilize one or more artificial intelligence (AI) models for pan-omics sequencing, which can detect errors, such as whether an incorrect DNA sample has been uploaded to the template (e.g., the skin microbiome and the gut microbiome have different bacteria). If such an error is detected, the pan-omics analysis process 550 can issue a warning and terminate the analysis because different regions have different microbiomes, and therefore such errors would invalidate the analysis results. By using one or more AI models, drug discovery and pharmacogenomics analysis can be accelerated.
[0228] Panomics sequencing can identify genomic fragments in a sample that are associated with a specific disease. Data from these disease-associated genomic fragments are then selected for subsequent steps. For ease of description, data from disease-associated genomic fragments in a sample are referred to as "clean data," and correspondingly, data from all genomic fragments in a sample are referred to as "raw data."
[0229] At step 560, the raw data is stored in FASTQ format (e.g., as a FASTQ file). At step 562, clean data is also stored in FASTQ format (e.g., as a FASTQ file). At step 564, the alignment data is stored in BAM format (e.g., as a BAM file). Annotations (such as difference combining) can then be added (step 566), and variants are stored in VCF format (step 568). At step 570, the patient data obtained in steps 560 through 568 is uploaded to a blockchain or stored in local storage.
[0230] Figure 12 It shows Figure 11 The details of step 558 are shown below. As shown in the figure, panomics sequencing can be used to analyze WES, WGS, or microbiome (box 602), where WES and WGS data are in VCF format, and microbiome data can be in comma-separated value (CSV) format (box 604).
[0231] At step 606, the WES, WGS, or microbiome data are filtered or otherwise selected based on the specific disease. At this step, a pre-upload script is also selected for the microbiome data based on the disease to be analyzed.
[0232] At step 608, the AI model and AI-based methods are used to process WES, WGS, or microbiome data using public / research databases (uploaded at step 506 of the initialization process 500). When processing microbiome data, the script selected at step 606 (box 610) is used.
[0233] In step 612, analysis results are generated for the disease to be analyzed.
[0234] In step 614, the generated analysis results can be verified by users such as technical experts who can modify the report as needed.
[0235] One or more reports can be generated using the analysis results (step 616), which can be sent to the patient for review (step 618).
[0236] The pan-omics analysis process 550 is applicable to all types of sequencing, where users can simply upload DNA sample data and select one or more analysis templates. Analysis reports are then automatically generated.
[0237] The pan-omics analysis process 550 is applicable to all comparative analyses. By using AI models and related AI algorithms, the pan-omics analysis process 550 can understand key differentiating factors between WGS / WES and microbiome samples, learn multiple patterns, and select appropriate patterns in the analysis, thereby simplifying the discovery of new diseases and / or treatments. In some implementations, the pan-omics analysis process 550 can be overlaid with or otherwise integrated with drug databases to enable rapid drug discovery based on genomes or to provide timely personalized medication recommendations. Furthermore, the pan-omics analysis process 550 is suitable for analyzing the skin microbiome, thereby allowing for skin microbiome-based cosmetic recommendations.
[0238] In some implementations, the local reference library can be updated periodically or periodically. More specifically, users can periodically or periodically download datasets from various sources, then compare the downloaded datasets with the local reference library to identify updated data in the various sources, and then use the identified data to update the local reference library. Comments can be added to notify users of updates.
[0239] In the above embodiments, the reference genome data, downloaded dataset, and report template obtained in steps 504 to 508 are stored in computing device 102 or 104. In some embodiments, such data may be stored in one or more blockchains.
[0240] In the above embodiments, the computer network system 100 is a blockchain system. In some embodiments, the computer network system 100 does not use blockchain (i.e., system 100 is not a blockchain system) and the pan-omics analysis process 550 does not use blockchain.
[0241] Other applications
[0242] Those skilled in the art will understand that the blockchain system 100 disclosed herein can be used in other applications.
[0243] For example, in some implementations, blockchain system 100 can be used in the aviation sector to automate various repetitive processes with sufficient security and efficiency, such as purchasing travel insurance, loyalty settlements, and paying taxes to authorities. Blockchain system 100 can also be used in airports to improve customer satisfaction, reduce costs, and increase revenue among network members. Blockchain system 100 can also be used in airports to improve ground operations through enhanced transparency in tracking, tracing, and operations, transactions, costs, and revenues, and to increase efficiency by reducing complexity and streamlining processes.
[0244] Blockchain System 100 can unify systems across various industries, such as the aviation and other travel industries (e.g., ticketing, loyalty programs) and non-aviation logistics industries (e.g., transportation, hotels), while enabling streamlined and seamless experiences and improving customer satisfaction.
[0245] In some implementations, the blockchain system 100 can be used in the legal field to allow lawyers and legal professionals to verify the authenticity of wills and trusts, effectively track chains of custody, assess intellectual property rights, and so on.
[0246] Blockchain System 100 allows lawyers and legal professionals to spend more time on less regulated aspects of the law. For example, compared to traditional technologies that can take longer to compile records and may require additional legwork to actually support reliable evidence from each previous custodian, lawyers with property ownership cases can use Blockchain System 100 to quickly verify the custodian chain of a client's property in seconds.
[0247] Blockchain system 100 can also facilitate lawyers in creating verifiable digital contracts, which are then digitally signed by the relevant parties using appropriate signing authorities. These digital contracts are stored on the blockchain, and content consumers can request to consume digital content from the blockchain.
[0248] In Blockchain System 100, the custodian chain document provides information on the collection, storage, transportation, and processing of electronic evidence. Blockchain System 100 provides evidence custodians to facilitate litigation by plaintiffs.
[0249] In some implementations, blockchain system 100 can be used in the education sector to simplify record keeping and sharing, enhance security and trust, streamline the recruitment process, and empower students with lifelong ownership of their academic records. Blockchain system 100 can store diplomas on the blockchain to allow students to own and manage their academic achievements, providing them with the ability to share and manage their academic achievements as needed and / or as desired. Historically, universities have owned and controlled student records, thus making students dependent on the institution to access and share their academic history and achievements.
[0250] A recent study by the University of Rome indicates that the university spends nearly €19,000 (over $20,000) annually on diploma verification, equivalent to approximately 36 weeks of work. In some implementations, a blockchain system 100 can be used to issue diplomas, streamlining the verification process at higher education institutions and saving them time and money. Because diplomas issued from the blockchain system 100 are virtually tamper-proof, blockchain-issued diplomas also make verifying students' academic records much easier for graduate schools.
[0251] Blockchain System 100 can be used to provide new and more affordable learning pathways and break down existing relationships between schools and students. For example, managing student tuition payments is a labor-intensive process involving multiple stakeholders such as students, parents, scholarship funds, private lending companies, federal and state agencies, and university finance departments. Blockchain System 100 can streamline this process, reducing administrative overhead and lowering tuition costs.
[0252] Blockchain System 100 can be useful for lifelong learning. More specifically, Blockchain System 100 can be used to provide broad access to open educational resources, defined as teaching and learning materials, such as books, podcasts, and videos, that exist in the public domain and can be freely used and redistributed. Blockchain System 100 allows these resources to be shared cheaply and securely on public networks.
[0253] Although human genome sequencing data has been described as an example in some of the above embodiments, those skilled in the art will understand that in other embodiments, systems, devices, methods, and computer-readable storage devices / media may also be used to store genome sequencing data of other species or life forms.
[0254] Those skilled in the art will understand that, in some embodiments, the methods disclosed herein can be implemented as one or more circuits, such as modules, apparatuses, devices, systems, etc. In some embodiments, the methods disclosed herein can be implemented as computer-executable instructions stored in one or more non-transitory computer-readable storage devices, such that, when executed, the instructions can cause one or more circuits (such as one or more processors) to perform the methods disclosed herein.
[0255] Those skilled in the art will understand that the various embodiments and / or features disclosed herein can be customized and / or combined as needed or desired. Furthermore, although embodiments have been described above with reference to the accompanying drawings, those skilled in the art will understand that variations and modifications can be made without departing from the scope of this disclosure as defined by the appended claims.
Claims
1. A computerized method for compressing genome sequencing data, the method comprising: The genome sequencing data was compared with reference sequencing data; Obtain one or more differential read sequences, each of which is a read sequence of the genome sequencing data that is different from the corresponding read sequence of the reference sequencing data; as well as Compressed genome sequencing data is obtained by compressing the one or more differentially read sequences using statistical compression methods or by using assembly methods with probabilistic data structures.
2. The method according to claim 1, wherein, The statistical compression method is an arithmetic coding method.
3. The method according to claim 1 or 2, further comprising: The one or more differential read sequences are processed using one or more statistical modeling methods; Wherein, compressing the one or more differential read sequences using the statistical compression method includes: The statistical compression method described above is used to compress one or more differential read sequences after processing.
4. The method according to any one of claims 1 to 3, further comprising: Multiple reads are assembled to form the reference sequencing data.
5. The method according to any one of claims 1 to 4, wherein, The processing of the one or more differential read sequences includes: The first read identifier (ID) of the genome sequencing data is parsed into one or more first fields; Store one or more fields of the first read segment ID; The second read ID of the genome sequencing data is parsed into one or more second fields; and Obtain one or more difference fields, each difference field representing a field in the one or more second fields that is different from the corresponding field in the one or more first fields.
6. The method according to claim 5, wherein, The one or more difference fields include a first difference field, which corresponds to a numerical field in the one or more second fields that is different from the corresponding field in the one or more first fields; and wherein the first difference field includes the numerical field, or includes an offset between the numerical field and the corresponding field in the one or more first fields.
7. The method according to claim 5, wherein, The one or more difference fields include a second difference field, which corresponds to a non-numeric field in the one or more second fields that is different from the corresponding field in the one or more first fields; and wherein the second difference field includes a portion of the non-numeric field that is different from the corresponding portion of the corresponding field in the one or more first fields.
8. The method according to any one of claims 1 to 7, wherein, The processing of the one or more differential read sequences includes: One or more nucleotide sequences in the one or more differential read sequences are compressed using a statistical model based on Markov chains.
9. The method according to claim 8, wherein, The statistical model is based on a 12th-order Markov chain.
10. The method according to any one of claims 1 to 9, further comprising: Multiple contigs are assembled into the reference sequencing data using multiple reads from the genome sequencing data.
11. The method according to claim 10, wherein, The assembly of multiple contigs into the reference sequencing data using multiple reads from the genome sequencing data includes: The de Bruin graph is constructed using a Bloom filter-based data structure.
12. The method according to claim 10, wherein, The assembly of multiple contigs into the reference sequencing data using multiple reads from the genome sequencing data includes: The de Bruin graph is constructed using a data structure based on the d-left count Bloom filter (dlCBF).
13. The method according to claim 11 or 12, wherein, The step of comparing the genome sequencing data with the reference sequencing data includes: The de Bruin graph is represented using a sparse bitmap, which is compressed using Elias-Fano encoding.
14. The method according to any one of claims 1 to 4, further comprising: Compressed genome sequencing data is stored in a blockchain.
15. A computerized method, comprising: At least the first part of the entity's genome sequencing data is stored at the network nodes of the blockchain system as transaction data.
16. The method of claim 15, further comprising: At least a first portion of the entity’s genome sequencing data is converted into a unique first hash string.
17. The method of claim 16, further comprising: The first hash string is combined with one or more first hash strings from at least a first portion of the genome sequencing data of the entity from one or more other nodes of the blockchain system to obtain a second hash string.
18. The method according to any one of claims 15 to 17, wherein, At least a first portion of the genome sequencing data of the entity is stored as one or more binary alignment and mapping (BAM) files.
19. The method according to claim 18, wherein, The one or more Binary Comparison and Mapping (BAM) files are associated with the first hash string.
20. The method according to any one of claims 15 to 19, further comprising: In the network nodes of the blockchain system, at least a first portion of the entity's genome sequencing data is stored as one or more non-fungible tokens (NFTs).
21. The method according to any one of claims 15 to 20, further comprising: At least a first portion of the entity’s genome sequencing data is protected using pseudo-anonymity and public key infrastructure (PKI).
22. The method according to any one of claims 15 to 21, further comprising: Unauthorized users are denied access to at least a first portion of the entity's genome sequencing data.
23. The method of claim 22, further comprising: Metadata of at least a first portion of the entity's genome sequencing data is presented to the unauthorized user.
24. The method according to any one of claims 15 to 23, wherein, At least a first portion of the entity's genome sequencing data is stored in the blockchain system, while a second portion of the entity's genome sequencing data is not present.
25. A computerized method, comprising: The genomic data of the DNA sample is compared with the data of a reference genome through secondary genomic analysis; as well as Panomics sequencing is performed by comparing the aligned genomic data with a local reference dataset to identify disease-related genomic fragments.
26. The method of claim 25, further comprising: The genomic data and / or identified genomic fragments of the DNA sample are stored in FASTQ format.
27. The method of claim 26, further comprising: The alignment data obtained by aligning the genomic data of the DNA sample with the data of the reference genome is stored in BAM format.
28. The method according to any one of claims 25 to 27, wherein, Performing the pan-omics sequencing includes: Perform whole exome sequencing (WES) or whole genome sequencing (WGS).
29. One or more non-transitory computer-readable storage devices, including computer-executable instructions for compressing genome sequencing data, wherein, When the instruction is executed, it causes the processing structure to perform the method according to any one of claims 1 to 28.
30. One or more processors configured to perform the method according to any one of claims 1 to 28.
31. A blockchain system, comprising: Multiple nodes are used to distribute and store the genomic sequencing data of entities as transaction data.
32. The blockchain system according to claim 31, wherein, The plurality of nodes are configured to compress the genome sequencing data using the method according to any one of claims 1 to 14.
33. The blockchain system according to claim 31 or 32, wherein, The plurality of nodes are configured to perform the method according to any one of claims 15 to 24 to distribute the genome sequencing data of the entity in the nodes.
34. The blockchain system according to any one of claims 31 to 33, wherein, The plurality of nodes are configured to perform the method according to any one of claims 25 to 28.
Citation Information
Cited By
Gene data storage method and device based on mixed medium and storage medium
CN121191603A