A method and apparatus for parsing and processing large-scale sequence data

Through multi-level analytical queues and heterogeneous storage solutions in a distributed environment, the problem of time-consuming and unreliable large-scale sequence data on a single server is solved, fast and reliable analysis and efficient storage is achieved, and data processing efficiency and reliability of bioinformatics research are improved.

CN116431700BActive Publication Date: 2025-07-29COMP NETWORK INFORMATION CENT CHINESE ACADEMY OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310267230.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-15
Publication Date
2025-07-29
Estimated Expiration
2043-03-15

AI Technical Summary

Technical Problem

The analysis and processing of large-scale sequence data is time-consuming and unreliable on a single server, making it difficult to meet the needs of fast analysis and efficient storage of bioinformatics research, especially in a distributed environment, the reliability of network communication and task scheduling across nodes is difficult to ensure.

Method used

Using a multi-level analytical queue and heterogeneous storage scheme in a distributed environment, sequence files are allocated through the main queue and work queue, combined with message middleware and heterogeneous fusion repository, the parsing process and incoming process are realized independently, the parsing expiration time and heartbeat detection are set, fault tolerance is ensured, and data storage efficiency is improved through compression and merging storage.

Benefits of technology

It realizes reliable and fast analysis and efficient storage of large-scale sequence data, supports dynamic server addition and decrease, improves the efficiency of fusion storage and retrieval and analysis of heterogeneous sequence data, prevents interruptions in the parsing process, and improves the reliability and efficiency of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116431700B_ABST
    Figure CN116431700B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and apparatus for parsing and processing large-scale sequence data. The method is as follows: 1) Traverse all sequence files to be parsed and processed, and record the file path, size, and type in the main queue; 2) Select multiple parsing servers and deploy a parsing process on each of the parsing servers for sending parsing requests to the main queue; after receiving the parsing request, the main queue creates a work queue for the corresponding parsing server and migrates the set of sequence files Fi assigned to the parsing server to the work queue; 3) When the ith parsing server starts to parse the jth sequence file Fij in the set of sequence files Fi, record the start parsing time and the process identifier for parsing Fij in Fij; 4) The ith parsing server stores the parsing result of Fij in an intermediate file; 5) Write the paths of different intermediate files into different queues and monitor, and store the corresponding intermediate files in the database.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of applied bioinformatics, and mainly relates to a method and device for parsing and processing large-scale sequence data, which can be applied to related fields such as biological phylogeny, epidemiological analysis research, and biogeographical flora occurrence analysis, and provides a method for processing and applying data mining of internationally public biological sequences. Background Art

[0002] Biological big data resources are important strategic resources. Among them, massive biological sequence data such as nucleic acids and proteins generated by high-throughput sequencing and other technologies are important sources of biological big data resources. In the past 30 years, a vast amount of sequence data resources have been accumulated in the field of bioinformatics. For example, the nucleic acid sequence data from the 252nd version (2022 / 10 / 19) of GenBank (a member of the International Nucleotide Sequence Database Collaboration (INSDC)) released 240 million nucleic acid data, including 1.5 trillion key-value pairs. Currently, the total number of data entries in GenBank has reached 3.1 billion, containing approximately 20 trillion base pairs in total. Additionally, there are 93 billion nucleotides located in the WGS sequence regions of 787 registered genome projects (Eric W Sayers et al., Nucleic Acids Research, 2021; Benson et al., Nucleic Acids Research, 1999; Rindone, Trends in Pharmacological Sciences, 1983). To improve the utilization efficiency of biological big data resources, there is an urgent need for methods to process large-scale sequences to meet the needs of refined retrieval and knowledge extraction in subsequent mining application fields such as systematic evolution research, epidemiological analysis research, and biogeographical flora occurrence analysis research.

[0003] Large-scale sequence data has characteristics such as a large number of data entries, wide sources, strong heterogeneity, and clear structure. The same piece of data corresponds to multiple pieces of information such as a brief description, metadata, and sequence data. Parsing one by one based on a single server takes several months, and the time is unacceptable. Based on an elastic and scalable distributed environment, the parsing speed can be greatly accelerated. However, due to operations such as cross-node network communication and task scheduling, the reliability is difficult to guarantee. Moreover, the parsed data contains metadata structures with large structural differences and detail data and sequence data of different sizes. A single type of storage scheme is difficult to meet the storage service requirements of complex parsed data. Therefore, the present invention provides a method and device for parsing and processing large-scale sequence data, which can reliably and quickly complete the parsing and processing of large-scale sequence files in a distributed environment, and realize the efficient storage and service of massive heterogeneous metadata and sequence data based on a fusion storage scheme. Summary of the Invention

[0004] In view of the efficiency and reliability problems faced by large-scale sequence file parsing applications, the present invention provides a method and apparatus for parsing and processing large-scale sequence data. By implementing this solution, reliable and fast parsing and processing of large-scale sequence files can be achieved, supporting dynamic addition and deletion of parsing servers, parsing exception detection and recovery, etc., and significantly improving the efficiency of application services such as heterogeneous sequence data fusion storage, retrieval, and analysis.

[0005] The technical solution of the present invention is as follows:

[0006] A method for parsing and processing large-scale sequence data, the steps of which include:

[0007] 1) Traverse all sequence files to be parsed and processed, and record relevant information such as file paths, sizes, and types into a thread-safe main queue; only file information is saved in the main queue, and the sequence files are not saved. The sequence files are stored in a shared file system; the main queue and the parsing process are located on different servers;

[0008] 2) Multiple parsing processes deployed on different parsing servers actively request sequence files to be parsed from the main queue. After receiving the parsing request, the main queue creates a work queue for each parsing server and migrates all the set of sequence files F i assigned to this parsing server to its bound work queue;

[0009] 3) When starting to parse the j-th sequence file F i in the set of sequence files F ij , record the start parsing time and the process identifier for parsing F ij into F ij . At the same time, the work queue sets a parsing expiration time for each file to be parsed according to the type and size of the sequence file. If the time is exceeded, the current parsed file is considered to have failed in parsing;

[0010] 4) According to the different types of sequence files, set different types of parsing processes. The parsing result data includes three types: semi-structured metadata, file detail data, and text sequence data. To improve the efficiency and reliability of data parsing and storage into the database, the above three types of parsing result data are all temporarily stored in intermediate files first, and then deposited into the heterogeneous fusion storage repository by a batch storage program;

[0011] 5) The semi-structured metadata is stored line by line in the intermediate file, and the maximum number of lines per file is limited to Mx. To further improve the data storage efficiency, first compress the text type of detail data and sequence data. To prevent data corruption caused by network or storage device exceptions, calculate the message digest of each detail data or sequence data before compression and splice it to the end of the data, and then compress. The number of bytes after compression is less than Smin The files are further merged into large files not smaller than S max and the offsets of the single compressed detail data or sequence data in the merged files are recorded in the index database;

[0012] 6) After all the above intermediate files are generated, they are written into the message queues in the message middleware. The paths of the intermediate files storing semi-structured metadata are written into the META.MQ queue, the paths of the detail data files of the text type are written into the DETAIL.MQ queue, and the paths of the sequence data files of the text type are written into the SEQ.MQ queue. The semi-structured metadata storage program listens to the META.MQ queue in a loop, the detail data storage program of the text type listens to the DETAIL.MQ queue in a loop, and the sequence data storage program of the text type listens to the SEQ.MQ queue in a loop. When an intermediate file is put into one of the above message queues by the parsing process, the corresponding storage program is started, the heterogeneous fusion storage interface is called, the corresponding data is batch stored in the database, and then the intermediate file after successful storage is marked by renaming;

[0013] 7) Each parsing server can start multiple parsing and storage processes according to its own resource configuration. The parsing process and the storage process of the same sequence file do not have to run on the same server. The three message queues in the above message middleware are shared by the parsing and storage processes on all servers. Each message queue can receive messages generated by multiple parsing processes of the same type at the same time and send messages to multiple storage processes of the same type; The parsing process and the storage process communicate through the message middleware. These two types of processes can run on different servers. The purpose of this design is to prevent the parsing and storage from affecting each other due to factors such as speed mismatch or abnormal operation, and improve the fault tolerance of the entire parsing process;

[0014] 8) A management process is set on each parsing server to monitor the status of the parsing and storage processes. When a process exits abnormally, it automatically tries to restart the process. If the attempt fails after multiple times, it notifies the administrator for manual inspection and debugging; The management process is also responsible for periodically responding to the heartbeat detection information sent by the work queue. If the heartbeat information of the work queue is not responded to within a certain period, the corresponding server is marked as an abnormal state, and then the main queue reclaims the sequence files in the work queue and destroys the work queue;

[0015] 9) If the work queue detects that the parsing process corresponding to a certain parsing file has not completed the parsing within the set expiration time, it notifies the task manager process to terminate the process, reset the corresponding parsing file, and then move it back to the main queue to wait for reallocation;

[0016] 10) The heterogeneous fusion repository responsible for storing the above three types of parsed result data consists of an object repository, a key-value index repository, and a full-text retrieval repository. The object repository is responsible for storing the compressed detailed data and sequence data. The key-value index repository is responsible for storing the offset positions of individual compressed data in the merged file. The file retrieval repository is responsible for storing semi-structured metadata;

[0017] 11) To improve the management efficiency of large-scale data, a certain type of database can be split into multiple sub-repositories for separate storage according to data characteristics, etc. internally. When performing full-text retrieval, each sub-repository is retrieved separately, and the top N records of each sub-repository are determined by heap sorting according to the specified fields, and then all the top N records of all sub-repositories are obtained through merge sorting.

[0018] 12) Encapsulate the read and write interfaces of the internal database, and modify the read and write interfaces for the data storage repository to operation interfaces for specific biological sequence data; all data write interfaces are idempotent interfaces, that is, writing the same data multiple times automatically overwrites the existing data, ensuring that even if the parsing program fails and restarts multiple times to generate the same data, there will be no duplicate data inside the database.

[0019] This device mainly includes three modules: a multi-level parsing queue for sequence files, a distributed parsing workflow for sequence files, and heterogeneous sequence data fusion storage and services. The multi-level parsing queue for sequence files caches parsed files in different parsing states through different queues, allocates sequence files to be parsed for different parsing tasks, and communicates with the task management program to prevent data loss caused by abnormal exits of parsing tasks. The distributed parsing workflow for sequence files splits the complex and time-consuming sequence file parsing process into multiple relatively independent sub-tasks, and constructs a workflow for parsing and storing sequence files into the database based on a message middleware to achieve reliable and flexible parallel acceleration parsing in a distributed environment. The heterogeneous sequence data fusion storage and service module stores billions of heterogeneous sequence data in multiple sub-repositories separately, constructs a fusion storage system that is compatible with semi-structured metadata, text sequence data, and text detailed data based on object storage, key-value index storage, full-text retrieval storage, etc., and provides cross-library parallel retrieval and analysis application services externally. The multi-level parsing queue allocates files for the parsing workflow and is responsible for recording the status information during the file parsing process. The comprehensive management and fault tolerance of the cluster status are mainly achieved through the multi-level queue. The parsing and storage workflow is constructed according to the type of input sequence files and the type of output parsed results, including multiple parsing processes and storage processes of different types, and forms a parsing and storage workflow through the message middleware. These processes run independently on the parsing server and are resident processes, not dynamically created. The fusion storage service is built based on the parsed result data, which includes optimized storage modules for different types of data and service interface encapsulation for applications.

[0020] 1) In the multi-level parsing queue module of the sequence file, the list of sequence files to be parsed and the list of sequence files being parsed are cached through the main queue and the working queue respectively. The main queue is a thread-safe queue that supports allocating different sequence files to multiple parsing programs accessing the queue simultaneously. The sequence file information in the main queue includes information such as the file storage path, file type, and file size.

[0021] 2) After a sequence file is allocated to a parsing process on a certain server, the main queue will automatically remove it. To prevent the loss of parsing data due to the abnormal exit of the parsing process, all sequence files allocated to a certain parsing server are recorded in the working queue bound to that server, and information such as the start parsing time and the parsing process identifier is added based on the existing file information. When a certain sequence file is parsed, it is removed from the associated working queue.

[0022] 3) For the sequence files in the working queue, parsing time thresholds are set according to factors such as size and type. If a certain parsing process fails to complete parsing within the set threshold time, it is considered that the parsing process is abnormal. The corresponding parsing file in the working queue is moved back to the main queue, and the management and monitoring program of the parsing process is notified.

[0023] 4) Each working queue periodically sends heartbeat information to the monitoring and management program on the associated server. If a certain working queue does not receive a response from the corresponding monitoring and management program for multiple cycles, it is considered that the server has stopped working. The working queue automatically migrates the sequence files that have not been parsed to the main queue. After the main queue receives the parsing files returned by the working queue, it resets their status and then destroys the working queue.

[0024] 5) The parsing progress information is summarized from the queue information: the number of files to be parsed is the length of the main queue, and the number of files being parsed is the sum of the elements in all sub-queues.

[0025] 6) In the distributed sequence data parsing workflow module, the parsing process of the parsing files is divided into two categories: parsing programs and warehousing programs. Among them, the parsing programs are classified according to the types of parsing files to be parsed, and each type of parsing file to be parsed corresponds to a type of parsing program; the warehousing programs are divided into three categories: bulk warehousing of detailed data, bulk warehousing of sequence data, and bulk warehousing of metadata. The parsing programs and warehousing programs form a parsing workflow through a message middleware.

[0026] 7) Among them, the parsed metadata is stored line by line in the intermediate temporary text file in JSON format, and the parsed detailed data and sequence data are stored in the form of individual files.

[0027] 8) The parsed metadata fields are divided into three categories: basic attribute fields, foreign key fields, and detailed attribute fields. Among them, the basic attribute fields contain brief description information of the current organism, allowing users to perform combined queries based on the above fields through Bool conjunctions. The foreign key fields are responsible for establishing association relationships with other databases. The detailed attribute fields are generally only used for applications such as data description and data visualization. To improve data retrieval efficiency, indexes are built on the basic attribute fields and foreign key fields.

[0028] 9) Multiple pieces of data can be obtained after parsing each sequence file. Among them, the metadata is stored line by line. To facilitate subsequent reliable parallel batch data import, the maximum number of lines threshold for each file is set to Mx, and the file name rule is: {bakfile}.{n}.meta, where bakfile is the parsed file name of the organism to be parsed, and n is the corresponding intermediate file number. The sequence data and detail data are both stored in the form of separate files. The detail data file with the info suffix, and the sequence data file with the seq suffix. The naming rules of the files are {bakfile}.{accession}.{version}.info and {bakfile}.{accession}.{version}.seq respectively, where bakfile is the parsed file name of the organism to be parsed, accession is the identifier of a piece of organism data, and version is the version of a piece of organism data. The two jointly uniquely identify a piece of organism data.

[0029] 10) The FASTA parsing program can parse out multiple nucleic acid or protein sequence data, and each sequence data is stored in the form of a separate file. The file naming rule is {bakfile}.{accession}.{version}.seq, where bakfile is the file name of the backup data of the currently parsed organism, accession is the identification number of the current organism data, and version is the version number of the current organism data.

[0030] 11) The path information of the intermediate files of the parsed detail data, sequence data, and metadata is respectively placed in three message queues INFO.MQ, SEQ.MQ, and META.MQ in the message middleware.

[0031] 12) Three subtasks, namely batch storage of detail data, batch storage of sequence data, and batch storage of metadata, respectively loop and listen to the three message queues INFO.MQ, SEQ.MQ, and META.MQ, and call the batch import interface of the heterogeneous fusion storage service to write the detail data, sequence data, and metadata into the corresponding repositories respectively. After successful writing, the intermediate file is named {filename}.done, where filename is the original intermediate file name.

[0032] 13) A task manager is set on each server to monitor the server load and the running status of the task processes running on the service. When a task process exits abnormally, the task manager attempts to restart the abnormally exited process. If the retry fails after multiple attempts, it notifies the administrator for manual debugging and inspection. When the server load is too high, on the premise of ensuring that at least one process of each type of subtask is running normally, the inbound subtask process that occupies more computing resources is preferentially selected to be terminated and the relevant resources it occupies are released.

[0033] 14) The task manager is also responsible for communicating with the work queue, replying in real time to the heartbeat detection information sent by the work queue, and after receiving the timeout message sent by the work queue, searching and terminating the corresponding parsing process according to the task ID, and at the same time cleaning the intermediate files generated by the parsing task according to the sequence file name.

[0034] 15) In the heterogeneous data fusion storage module, the parsed metadata and sequence data are respectively stored in three types of data repositories, namely, the distributed object storage repository, the key-value index storage repository, and the full-text retrieval database. Among them, the distributed object storage repository is responsible for storing the detailed data and sequence data, the full-text retrieval database is responsible for storing the JSON format metadata, and the key-value index storage repository is responsible for storing the index information of the detailed objects and sequence object data.

[0035] 16) Considering that the data scale of the parsed sequence files may be very large, the single-library storage has low read and write efficiency and is difficult to manage and maintain. Therefore, the biological data is split into sub-libraries by type internally.

[0036] 17) In order to shield the internal storage details, the present invention combines and encapsulates the internal database interfaces to provide service interfaces for accessing, retrieving, and storing various types of sequence data, including full-library keyword retrieval, multi-field combination retrieval, metadata reading and writing and batch import and export, sequence reading and writing and batch import and export, partial sequence export based on keywords, etc.

[0037] 18) Considering that storing a large number of small detailed and sequence files will seriously reduce the efficiency of object storage, in the sequence data and detailed data writing interfaces, a compressed and merged storage method is adopted to improve the storage efficiency: before the small sequence and detailed files are written into the object storage, the message digest of the data content is first calculated and spliced to the end of the data, and then it is compressed. Considering that there are many repeated letters in both the detailed and sequence data, compression will significantly reduce the data size, and after further combined compression, large files are formed. The offset position of a single detailed or sequence data in the merged file is recorded in the index storage repository, where the primary key in the index library is the unique identifier of the data, and the value is the starting storage position of the sequence in the large file and the sequence size.

[0038] 19) When reading detailed information or sequence data, first determine the object storage location by the index database, then read the corresponding detailed information or sequence data according to the specified offset and size. After the detailed information or sequence data is read, it is decompressed and then the message digest of the data part is calculated and compared with the message digest at the end of the data. If the two are consistent, the data is returned; if not, it means the data is corrupted and the data is discarded.

[0039] 20) The full-text retrieval database and the key-value index repository use the sequence data identifier and version number as the only primary key, and the data with the same primary key automatically overwrites the existing data. Based on this method, when a certain data file is re-parsed and stored in the database due to the failure of the parsing task halfway, it can be ensured that there will be no duplicate data in the database.

[0040] 21) When performing full-text retrieval, each sub-database is retrieved separately, and the top n of each sub-database are determined by heap sorting according to the specified field, and then the top n data of all sub-data are obtained by the way of merge sorting.

[0041] The present invention also provides a server, which is characterized in that it includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing the steps in the above method.

[0042] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and is characterized in that the steps of the above method are realized when the computer program is executed by a processor.

[0043] The advantages of the present invention are as follows:

[0044] The parsing of large-scale sequence files is very time-consuming and usually lasts for several months. During the parsing process, the parsing process often terminates due to abnormal operating environments (such as server failures, high loads, network problems, etc.). If it is necessary to completely re-parse from the beginning every time, the cost is very high. For this reason, this patent proposes a method for quickly and reliably completing the parsing and storage in a distributed cluster. Based on the method described in this patent, it is allowed to dynamically add or subtract parsing servers during the parsing process without causing the parsing and storage process to interrupt, and server failures can be detected and abnormal parsing servers can be automatically removed from the cluster. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 is the parsing processing method and device of the present invention DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] The present invention will be further described in detail below with reference to the drawings. The examples given are only used to explain the present invention and are not intended to limit the scope of the present invention.

[0047] A parsing and processing device for large-scale sequence data described in this embodiment mainly includes three core modules: a multi-level parsing queue for sequence files, a distributed sequence file parsing workflow, and heterogeneous sequence data fusion storage and service. Its system module composition and main data flow are as Figure 1 shown.

[0048] 1) Traverse all nucleic acid parsing files, and store information such as file path, size, and type into the main queue. Common nucleic acid parsing files include three types of files: gbff.gz, seq.gz, and fasta.gz;

[0049] 2) Start processes such as the task management program, GBFF parsing program, FASTA parsing program, detailed data storage program, sequence data storage program, and metadata storage program on the parsing server one by one. After the above processes are started, the task management process sends a registration request containing information such as the current parsing server IP address to the main queue. After receiving the registration request, the main queue creates a work queue for it;

[0050] 3) The GBFF parsing process requests files of gbff.gz and seq.gz types from the main queue, and the FASTA parsing program requests files of fasta.gz type from the main queue. After the parsing files are allocated to the parsing process, they are migrated from the main queue to the work queue of the parsing server where the parsing process is located. The work queue sets the parsing expiration time according to the file size and type;

[0051] 4) To prevent the loss of uncompleted parsing file data due to the downtime of the parsing server, the main queue and the work team are deployed on a separate server, and messages are received and sent through the network;

[0052] 5) After the parsing process starts working, it records information such as the start time and process ID into the work queue, and periodically sends heartbeat information to the task management program;

[0053] 6) If the parsing process corresponding to a certain parsing file fails to complete the parsing within the set expiration time, the work queue sends a parsing file recovery message to the task management program. The task management program terminates the corresponding process, and the work queue migrates the expired file back to the main queue for reallocation;

[0054] 7) If multiple heartbeat messages continuously sent by the work queue to the task management program are not responded to, it is considered that the parsing server has crashed. The work queue sends a request to the main queue to cancel the server. After receiving the request, the main queue reclaims all parsing files allocated to the work queue and destroys the queue;

[0055] 8) The GBFF parser parses to obtain three types of data: semi-structured metadata, text detail data, and text sequence data. The FASTA file parser parses to obtain text sequence data. Multiple semi-structured metadata are merged and stored in an intermediate file by row. The text type detail data and text type sequence data are stored separately in a file. The intermediate file is stored in a distributed file system shared by all servers.

[0056] 9) Nucleic acid metadata fields are divided into three categories: basic attribute fields (such as Accession, Version, Title, Length, etc.), foreign key fields (such as Pubmed, Taxonomy, BioProject, BioSample, etc.), and detailed attribute fields (such as Gene, CDS, etc.). Basic attribute fields contain a brief description of the current organism, allowing users to perform combined queries based on these fields using Boolean conjunctions. Foreign key fields are responsible for establishing associations with other databases, allowing for associated queries and statistical analysis. Detailed attribute fields are generally only used for applications such as data description and data visualization.

[0057] 10) To improve data retrieval efficiency, hash indexes and word-based inverted indexes are built for basic attribute fields, and hash indexes are built for foreign key fields to facilitate quick location of related data.

[0058] 11) The maximum number of lines in the metadata intermediate file is set to 10,000 lines, and the file name rule is: {bakfile}.{n}.meta, where bakfile is the name of the biological analysis file to be parsed, and n is the corresponding intermediate file number. Sequence data and detailed data are stored in the form of separate files, with detailed data files ending in info and sequence data files ending in seq. The file naming rules are {bakfile}.{accession}.{version}.info and {bakfile}.{accession}.{version}.seq, respectively, where bakfile is the name of the biological analysis file to be parsed, accession is the identifier of a biological data item, and version is the version of a biological data item. The two together uniquely identify a biological data item. After the storage is completed, the intermediate file is renamed to {filename}.done, and filename is the original file name;

[0059] 12) The paths of the intermediate result files obtained by the GBFF parsing process and the FASTA parsing process are passed to the detailed data storage process, the sequence data storage process, and the metadata storage process through the message middleware. Three types of message queues are set in the message middleware to store different types of intermediate result files: META.MQ (metadata), DETAIL.MQ (detailed data), SEQ.MQ (sequence data). The detailed data storage process, the sequence data storage process, and the metadata storage process continuously monitor the above three queues;

[0060] 13) Multiple parsing processes and storage processes can be run respectively according to the amount of data to be parsed and the server configuration. The storage subtasks can run on different servers, without being restricted by the original nucleic acid parsing file. The message middleware serves as an intermediate buffer storage, which can coordinate each subtask with mismatched speeds so that they do not affect each other;

[0061] 14) The metadata intermediate file is directly imported into the heterogeneous fusion database by the corresponding storage process. Considering that there are a large number of relatively small detailed data and sequence data, directly storing such data will affect the storage efficiency. Therefore, before writing the detailed data and sequence data into the heterogeneous fusion storage repository, the storage efficiency is improved and the data security is ensured by merging, compressing, and adding a message digest. Before writing all nucleic acid detailed data and sequence data into the object storage, first calculate their corresponding message digests and splice them to the end of the message, and then compress. Data files larger than 4MB after compression are directly stored in the object storage repository, and the access path of the object file is recorded in the key-value index library; data files smaller than 4MB are further merged into large files not less than 4MB, and then written into the object storage repository, and the access path of the object file and the offset of each small file in it are recorded in the key-value index library; the data record structure in the key-value index library is: (accession,version)→(object,offset,size), where accession and version are the nucleic acid data identifier and version number respectively, object is the access path of the object file, offset is the offset of the current nucleic acid detailed data or sequence data in the object file, and size is the size of the nucleic acid detailed data or sequence data;

[0062] 15) All parsing processes and storage processes located on the same parsing server are managed by the same task monitoring process. When it is detected that a certain parsing process or storage process fails, the process is automatically attempted to be restarted. If it still fails after multiple retries, the administrator is notified for inspection and debugging;

[0063] 16) The heterogeneous fusion repository responsible for storing the parsed result data consists of an object storage, a key-value index library, and a metadata database. The object storage repository is responsible for storing the detail data and sequence data in the form of single files. The full-text retrieval library is responsible for storing the metadata of the JSON type. The key-value index library is responsible for storing the index information of the detail data and sequence data. To ensure the efficiency of massive data storage management, CEPH is recommended for the object storage, TiDB is recommended for the key-value index library, and ElasticSearch is recommended for the metadata database. Further encapsulation is performed on the read-write, retrieval, and other interfaces of CEPH, TiDB, and ElasticSearch internally to implement service interfaces such as full-library keyword retrieval, multi-field combination retrieval, nucleic acid detail data read-write, and nucleic acid sequence data read-write for nucleic acid data.

[0064] 17) All data operation interfaces are idempotent operation interfaces. Repeatedly writing the same piece of data automatically overwrites the existing data to prevent duplicate data in the internal database. When performing full-text retrieval, each sub-library is retrieved separately, and the top n records of each sub-library are determined by heap sorting according to the specified fields. Then, through the method of merge sorting, the top n records of all sub-data are obtained.

[0065] The present invention also provides a server, which is characterized by including a memory and a processor. The memory stores a computer program, and the computer program is configured to be executed by the processor. The computer program includes instructions for executing the steps in the above method.

[0066] The present invention also provides a computer-readable storage medium, on which a computer program is stored. The computer program is characterized in that when it is executed by a processor, the steps of the above method are implemented.

[0067] Although specific embodiments of the present invention are disclosed for illustrative purposes, the purpose is to help understand the content of the present invention and implement it accordingly. Those skilled in the art can understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the content disclosed in the best embodiments, and the scope of protection required by the present invention is defined by the scope defined in the claims.

Claims

1. A method for parsing and processing large-scale sequence data, the steps of which include: 1) Traverse all sequence files to be parsed and processed, and record the file path, size, and type in the main queue; 2) Select multiple parsing servers and deploy a parsing process on each of the parsing servers; The parsing process on the i-th parsing server sends a parsing request to the main queue. After receiving the parsing request, the main queue creates a work queue for the corresponding parsing server and migrates the set of sequence files F i assigned to the parsing server into the work queue. The work queue sets a parsing expiration time for each sequence file to be parsed according to the type and size of the sequence file. If the parsing expiration time is exceeded, it is considered that the corresponding sequence file parsing fails. 3) The i-th parsing server starts to parse the j-th sequence file F in the sequence file set F i and records the start parsing time and the process identifier for parsing F ij into F ij ; ij ​ 4) The i-th parsing server stores the parsing result of F ij in an intermediate file; the parsing result includes semi-structured metadata, file detail data, and text sequence data; 5) Write the intermediate file path storing semi-structured metadata in step 4) into the META.MQ queue, the intermediate file path storing text detail data into the DETAIL.MQ queue, and the intermediate file path storing text sequence data into the SEQ.MQ queue; The semi-structured metadata warehousing process circularly monitors the META.MQ queue. When there is an intermediate file path written into the META.MQ queue, store the corresponding intermediate file into the database; The text type detail data warehousing process circularly monitors the DETAIL.MQ queue. When there is an intermediate file path written into the DETAIL.MQ queue, store the corresponding intermediate file into the database; The text type sequence data warehousing process circularly monitors the SEQ.MQ queue. When there is an intermediate file path written into the SEQ.MQ queue, store the corresponding intermediate file into the database.

2. The method according to claim 1, wherein The method for storing the parsing result in an intermediate file is as follows: store the parsed semi-structured metadata line by line in the first type of intermediate file, and set the maximum number of lines of the first type of intermediate file to Mx. If the number of lines in the first type of intermediate file exceeds Mx, generate a new first type of intermediate file to store the subsequent parsed semi-structured metadata; compress the parsed text detail data and store it in the second type of intermediate file; the compression method for the text detail data is: first calculate the message digest of each text detail data and splice it to the end of the text detail data, and then compress the spliced data. If the number of bytes of the compressed file is less than the set value S min , then the compressed file smaller than the set value S min is merged into a file not less than the set value S max , and the offset of each compressed text detail data in the merged file is recorded in the index database; The parsed text sequence data is compressed and stored in a third type of intermediate file. The compression method for the text sequence data is as follows: First, calculate the message digest of each text sequence data and splice it to the end of the text sequence data. Then, compress the spliced data. If the number of bytes of the compressed file is less than the set value S min , then the compressed files smaller than the set value S min are merged into a file not less than the set value S max , and the offset of each compressed text detail data in the merged file is recorded in the index database.

3. The method according to claim 1, characterized in that, The parsing process and the warehousing process of the same sequence file can run on different parsing servers; The META.MQ queue, the DETAIL.MQ queue, and the SEQ.MQ queue are all shared queues shared by the parsing processes and the warehousing processes.

4. The method according to claim 1 or 2 or 3, characterized in that The interface for data writing is an idempotent interface, including the interface for writing the parsing result into the intermediate file in step 4) and the interface for writing the intermediate file into the database in step 5).

5. The method according to claim 1 or 2 or 3, characterized in that, A management process is set on each of the parsing servers, which is responsible for monitoring the status of the parsing and warehousing processes. When a process exits abnormally, automatically restart the process. If the restart fails multiple times, notify the administrator for manual inspection and debugging; The management process is also responsible for periodically responding to the heartbeat detection information sent by the work queue. If the heartbeat information of work queue i has not been responded to for more than a certain period, mark the corresponding parsing server i as an abnormal state, and then the main queue reclaims the sequence file in the work queue i and destroys the work queue i; If work queue a detects that the corresponding parsing process has not completed the parsing of the sequence file within the set expiration time, notify the task manager process to terminate the parsing process, and reset the corresponding sequence file and then move it back to the main queue to wait for reallocation.

6. The method according to claim 1 or 2 or 3, characterized in that, The database consists of an object repository, a key-value index library, and a full-text retrieval library. Among them, the object repository is responsible for storing the compressed detail data and sequence data, the key-value index library is responsible for storing the offset position of a single compressed data in the merged file, and the full-text retrieval library is responsible for storing semi-structured metadata.

7. The method according to claim 6, wherein The full-text retrieval library includes multiple sub-libraries, and the same sub-library stores data with the same characteristics; When performing retrieval, retrieve each sub-library separately, determine the top N matching data of each sub-library by heap sorting according to the specified field, and then obtain the top N data of each sub-library through merge sorting.

8. An analysis and processing device for large-scale sequence data, characterized in that, It includes a multi-level parsing queue module for sequence files, a distributed sequence file parsing workflow module, and a heterogeneous sequence data fusion storage and service module; The multi-level parsing queue module for sequence files is used to traverse all sequence files to be parsed and record the file paths, sizes, and types into the main queue; And create a work queue for the parsing server, and migrate the set of sequence files assigned to the parsing server to the work queue; the work queue sets a parsing expiration time for each sequence file to be parsed according to the type and size of the sequence file, and if the parsing expiration time is exceeded, the corresponding sequence file is considered to have failed to be parsed; The distributed sequence file parsing workflow module is used to parse each sequence file in the set of sequence files, and record the start parsing time and the process identifier for parsing the sequence file into the corresponding sequence file; And store the parsing result in an intermediate file; the parsing result includes semi-structured metadata, file detail data, and text sequence data; The heterogeneous sequence data fusion storage and service module is used to write the intermediate file path storing the semi-structured metadata into the META.MQ queue, the intermediate file path storing the text detail data into the DETAIL.MQ queue, and the intermediate file path storing the text sequence data into the SEQ.MQ queue; the semi-structured metadata warehousing process listens to the META.MQ queue in a loop, and when there is an intermediate file path written into the META.MQ queue, stores the corresponding intermediate file into the database; the text type detail data warehousing process listens to the DETAIL.MQ queue in a loop, and when there is an intermediate file path written into the DETAIL.MQ queue, stores the corresponding intermediate file into the database; the text type sequence data warehousing process listens to the SEQ.MQ queue in a loop, and when there is an intermediate file path written into the SEQ.MQ queue, stores the corresponding intermediate file into the database.

9. A server, characterized in that, It includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing the steps in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Method and system for quickly loading data to database based on multi-process concurrency and plug-ins

    CN110347440A

  • File analysis and storage method and device and file generation method and device

    CN111339041A