Data splitting method, computer device, and computer readable storage medium
By splitting and matching tag sequences in second-generation nucleic acid sequencing and utilizing hash table encoding and fault-tolerant tag sequences, the efficiency and accuracy issues of mixed sample data splitting are solved, achieving efficient data splitting and sample source identification.
Patent Information
- Application Number
- PCT/CN2024/104643
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-15
- Filing Date
- 2024-07-10
- Publication Date
- 2025-10-23
AI Technical Summary
During the second-generation nucleic acid sequencing process, how to effectively split the sequencing data of mixed samples to identify the sample sources of each sequencing object, especially when the tag sequence lengths are inconsistent, existing technologies make it difficult to efficiently split the data.
By splitting the tag sequence in the sequencing data and matching it with the reference tag sequence, hash table encoding and fault-tolerant tag sequence are used to determine the sample source of the sequencing data, and then the sequencing data of the same biological sample can be split.
It improves the efficiency and convenience of data splitting, can accurately identify the label sequence part in mixed sample data, and improves the accuracy and speed of data splitting.
Smart Images

Figure CN2024104643_23102025_PF_FP_ABST
Abstract
Description
Data splitting method, computer device and computer readable storage medium
[0001] The present disclosure claims priority to the Chinese patent application No. 2024104518738, filed on April 15, 2024, entitled "Data splitting method, computer device and computer storage medium", the entire content of which is incorporated herein by reference. TECHNICAL FIELD
[0002] The present disclosure relates to the technical field of computer, in particular, to a data splitting method, a computer device and a computer readable storage medium. BACKGROUND
[0003] In the process of second-generation nucleic acid sequencing, in order to improve the sequencing speed and reduce the sequencing cost, multiple biological samples are mixed and sequenced together. Here, the multiple biological samples become nucleic acid samples (e.g., nucleic acid fragments) after library preparation, which can also be referred to as sequencing objects. The sequencing objects not only contain nucleic acid sequences to be sequenced, but also contain tag sequences, which are used to identify the sample source of the sequencing object. Since the lengths and quantities of the tag sequences contained in different sequencing objects may not be the same during library preparation, it is particularly important to split the sequencing data from the same sequencing object from all mixed sequencing data during the sequencing process of mixed samples.
[0004] SUMMARY
[0005] The present disclosure provides at least a data splitting method, a computer device and a computer readable storage medium.
[0006] In a first aspect, the present disclosure provides a computer device, comprising: a processor, a memory and a bus, the memory storing machine readable instructions executable by the processor, when the computer device is running, the processor and the memory communicate through the bus, the machine readable instructions are executed by the processor to perform a splitting process of mixed sample data, the mixed sample data comprises a plurality of sequencing data from a plurality of sequencing objects, any sequencing object in the plurality of sequencing objects comprises a tag sequence and a nucleic acid sequence, the tag sequence is used to indicate the sample source of the sequencing object, the sequence lengths of the tag sequences in the plurality of sequencing objects are not completely the same, for any sequencing data in the plurality of sequencing data, the splitting process of the mixed sample data comprises:
[0007] obtaining the sequencing data, the sequencing data comprising sequencing data of the tag sequence and sequencing data of the nucleic acid sequence;
[0008] The sequencing data of the tag sequence is cut out from the sequencing data according to at least one sequence length, to obtain at least one sequencing data of the tag sequence; wherein the at least one sequence length covers the sequence length of at least part of the tag sequences in the plurality of sequencing objects;
[0009] The at least one sequencing data of the tag sequence is matched with a reference tag sequence respectively, and the sample source of the sequencing data is determined according to the matching result;
[0010] At least one sequencing data derived from the same biological sample is split out from the plurality of sequencing data according to the identified sample source.
[0011] In a possible implementation, in the splitting process of the mixed sample data performed by the processor, the at least one sequencing data of the tag sequence is matched with a reference tag sequence respectively, and the sample source of the sequencing data is determined according to the matching result, comprising:
[0012] The at least one sequencing data of the tag sequence is encoded respectively, and is converted into a key in a hash table through the encoding, to obtain at least one to-be-processed key;
[0013] The at least one to-be-processed key is queried in the hash table;
[0014] If a key matching any one of the at least one to-be-processed key is queried from the hash table, a mapping value is obtained according to the mapping relationship of the hash table, and the sample source of the sequencing data is determined according to the mapping value;
[0015] The key in the hash table is the encoding data of the reference tag sequence, and the value is a number indicating the sample source.
[0016] In a possible implementation, the processor is further configured to perform:
[0017] If a key matching any one of the at least one to-be-processed key is not queried from the hash table, it is determined that the sample source of the sequencing data is an unknown sample.
[0018] In a possible implementation, the key in the hash table further comprises the encoding data of a fault-tolerant tag sequence of the reference tag sequence;
[0019] The fault-tolerant tag sequence is obtained by replacing the bases at at most a preset number of base positions in the reference tag sequence with base types other than the base type currently corresponding to the base position respectively, to obtain the fault-tolerant tag sequence of the reference tag sequence; wherein the preset number is a base fault-tolerant value.
[0020] In a possible implementation, in the splitting process of the mixed sample data performed by the processor, the at least one sequencing data of the tag sequence is encoded respectively, and the encoding is converted into a key in a hash table to obtain at least one to-be-processed key, including:
[0021] The occurrence frequencies of the bases are determined, and the encoding values corresponding to the bases are determined based on the occurrence frequencies;
[0022] The at least one sequencing data of the tag sequence is encoded respectively based on the encoding values corresponding to the bases to obtain at least one to-be-processed key.
[0023] In a possible implementation, in the splitting process of the mixed sample data performed by the processor, the at least one to-be-processed key is queried in the hash table, including:
[0024] The at least one to-be-processed key is queried in the hash table corresponding to the sequence length respectively; wherein the reference tag sequences corresponding to the keys in the same hash table have the same sequence length.
[0025] In a possible implementation, in the splitting process of the mixed sample data performed by the processor, the at least one to-be-processed key is queried in the hash table corresponding to the sequence length respectively, including:
[0026] K threads are established, and the corresponding hash table is queried in parallel based on each thread and the at least one to-be-processed key; wherein K is the number of the hash tables; or,
[0027] The number of keys contained in each hash table is determined, and the searching order of each hash table is determined based on the number of the keys, and each hash table is queried based on the at least one to-be-processed key in the searching order.
[0028] In a possible implementation, the computer device is a computer device on-board a nucleic acid sequencer, or a computer device in wireless communication with the nucleic acid sequencer.
[0029] In a second aspect, the embodiments of the present disclosure provide a data splitting method, applied to mixed sample data, the mixed sample data including a plurality of sequencing data from a plurality of sequencing objects, any sequencing object in the plurality of sequencing objects including a tag sequence and a nucleic acid sequence, the tag sequence being used to indicate the sample source of the sequencing object, the sequence lengths of the tag sequences in the plurality of sequencing objects being not completely same, and the splitting process of the mixed sample data including:
[0030] The sequencing data is obtained, the sequencing data including sequencing data of a tag sequence and sequencing data of a nucleic acid sequence;
[0031] cutting the sequencing data of the tag sequences from the sequencing data according to at least one sequence length, to obtain at least one sequencing data of the tag sequences; wherein the at least one sequence length covers sequence lengths of at least part of the tag sequences in the plurality of sequencing objects;
[0032] matching the at least one sequencing data of the tag sequences with reference tag sequences respectively, and determining sample sources of the sequencing data according to matching results;
[0033] splitting at least one sequencing data derived from the same biological sample from the plurality of sequencing data according to the identified sample sources.
[0034] In a possible implementation, in the splitting process of the mixed sample data performed by the processor, the at least one sequencing data of the tag sequences is matched with reference tag sequences respectively, and sample sources of the sequencing data are determined according to matching results, including:
[0035] encoding the at least one sequencing data of the tag sequences respectively, and converting into keys in a hash table through encoding, to obtain at least one to-be-processed key;
[0036] querying the at least one to-be-processed key in the hash table;
[0037] if a key matching any one of the at least one to-be-processed key is queried from the hash table, obtaining a mapping value according to a mapping relationship of the hash table, and determining a sample source of the sequencing data according to the mapping value;
[0038] wherein the keys in the hash table are encoding data of reference tag sequences, and the values are numbers indicating sample sources.
[0039] In a possible implementation, the method further includes:
[0040] if a key matching any one of the at least one to-be-processed key is not queried from the hash table, determining that the sample source of the sequencing data is an unknown sample.
[0041] In a possible implementation, the keys in the hash table further include encoding data of a fault-tolerant tag sequence of a reference tag sequence.
[0042] the fault-tolerant tag sequence is obtained by replacing bases at at most a preset number of base positions in the reference tag sequence with base types other than a current base type of the base positions, to obtain a fault-tolerant tag sequence of the reference tag sequence; wherein the preset number is a base fault-tolerance value.
[0043] In a possible implementation, in the splitting process of the mixed sample data performed by the processor, the at least one sequencing data of the label sequence is respectively encoded, and the encoding is converted into a key in a hash table to obtain at least one to-be-processed key, including:
[0044] The occurrence frequency of each base is determined, and based on the occurrence frequency, an encoding value corresponding to each base is determined;
[0045] The at least one sequencing data of the label sequence is respectively encoded based on the encoding value corresponding to each base to obtain the at least one to-be-processed key.
[0046] In a possible implementation, in the splitting process of the mixed sample data performed by the processor, the at least one to-be-processed key is queried in the hash table, including:
[0047] The at least one to-be-processed key is queried in the hash table corresponding to each sequence length, and the reference label sequence corresponding to the keys in the same hash table has the same sequence length.
[0048] In a possible implementation, the querying the at least one to-be-processed key in the hash table corresponding to each sequence length includes:
[0049] K threads are established, and the corresponding hash table is queried in parallel based on each thread and the at least one to-be-processed key, where K is the number of the hash tables; or
[0050] The number of keys contained in each hash table is determined, and the searching order of each hash table is determined based on the number of the keys, and each hash table is sequentially queried based on the at least one to-be-processed key in the searching order.
[0051] In a third aspect, the present disclosure also provides a computer readable storage medium, which stores a computer program. When the computer program is run by a processor, the steps of the second aspect or any possible implementation of the second aspect are performed.
[0052] In order to make the above objectives, characteristics and advantages of the present disclosure more apparent and understandable, the following preferred embodiments are specifically described with reference to the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0053] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings needed to be used in the embodiments will be briefly introduced hereinafter, and the drawings incorporated into the description and form a part of the description, which show the embodiments consistent with the present disclosure, and are used to explain the technical solutions of the present disclosure together with the description. It should be understood that the following drawings only show some of the embodiments of the present disclosure, and therefore should not be regarded as a limitation on the scope, and for those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.
[0054] Fig. 1 shows a flow chart of a data splitting method provided by an embodiment of the present disclosure;
[0055] Fig. 2 shows a whole flow chart of the data splitting method provided by an embodiment of the present disclosure;
[0056] Fig. 3 shows a structural schematic diagram of a computer device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0057] In order to make the purpose, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only some of the embodiments of the present disclosure, but not all the embodiments. The components of the embodiments of the present disclosure described and shown herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure is not intended to limit the scope of the claimed present disclosure, but only represents selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present disclosure.
[0058] In the process of second-generation nucleic acid sequencing, in order to improve the sequencing speed and reduce the sequencing cost, a plurality of biological samples are generally mixed and sequenced together (i.e. mixed sample sequencing). Here, a plurality of biological samples are each prepared into a respective nucleic acid sample (which can also be referred to as a library or sequencing object) through a library preparation process, and a plurality of sequencing objects derived from a plurality of biological samples are mixed and sequenced together. After sequencing by a nucleic acid sequencer, sequencing data of each sequencing object can be obtained simultaneously. Since the sequencing data of each sequencing object is mixed together, the sequencing data of each sequencing object derived from the same biological sample needs to be separated from all the mixed sequencing data before data analysis is performed on the sequencing data of the same biological sample. In order to achieve data separation, different sequencing objects in the library preparation process are added with a tag sequence (also referred to as a barcode sequence or an Index sequence) identifying the sample source thereof. Through sequencing by a nucleic acid sequencer, sequencing data of the nucleic acid sequence and the tag sequence can be obtained, and the tag sequence can be obtained by reading the sequencing data of the tag sequence, and the sample source of each sequencing object (e.g. derived from Zhang San or Li Si) can be identified through the tag sequence.
[0059] Since the types and lengths of sample tags in sequencing objects provided by different customers can be different, for example, some sequencing objects use single tag type, and some sequencing objects use double tag type. For another example, the length of sample tags in some sequencing objects is 6 bp, and the length of sample tags in some sequencing objects is 8 bp.
[0060] For this situation, one method in the industry is to pad the tag sequences of sequencing objects derived from different customers to a uniform length in the library preparation process. However, since the library preparation personnel and the data separation personnel can not be the same personnel of an institution, it is difficult to require the library preparation personnel to prepare the library according to the needs of the data separation personnel.
[0061] Based on the above research, the present disclosure provides a data splitting method, a computer device and a computer readable storage medium. After obtaining sequencing data of sequencing objects derived from multiple biological samples through on-machine sequencing, in the case that the length of the tag sequence is uncertain, the sequencing data of the tag sequence can be first cut out from the sequencing data according to at least one sequence length, and the cut-out sequencing data of the tag sequence is matched with a reference tag sequence. The reference tag sequence can be a standard tag sequence added to the sequencing objects derived from multiple biological samples during library preparation. The sample source of the sequencing data is determined based on the matching result, and the sequencing data derived from the same biological sample is split from the multiple sequencing data according to the identified sample source. Through this method, the tag sequence part of each sequencing data in the mixed sample data can be directly identified, and the efficiency and convenience of data splitting are improved.
[0062] It should be noted that similar reference numerals and letters refer to similar items throughout the accompanying drawings, and therefore, once an item is defined in one drawing, it need not be further defined and explained in subsequent drawings.
[0063] The term "and / or" herein is merely used to describe an associated relationship, that is, there can be three relationships, for example, A and / or B can represent three cases of A alone, A and B together, and B alone. In addition, the term "at least one" herein represents any one of a plurality or any combination of at least two of a plurality, for example, including at least one of A, B and C can represent including any one or more elements selected from the set consisting of A, B and C.
[0064] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the personal information (including but not limited to attribute information, face image, etc.) involved in the present disclosure is obtained with the authorization of the user. Specifically, the user can be prompted to authorize the request information through a pop-up window in the page, information push and the like, and the above-mentioned personal information can be obtained after the user agrees to authorize.
[0065] In order to facilitate the understanding of the present embodiment, first, a data splitting method disclosed by the embodiments of the present disclosure is introduced in detail. The data splitting method provided by the present disclosure is applied to mixed sample data, and the mixed sample data includes multiple sequencing data of multiple sequencing objects. Here, the multiple sequencing objects can refer to sequencing objects derived from multiple biological samples, for example, multiple species of samples, which can exemplarily include chicken, cow and horse samples; or can refer to samples of different individuals of the same species, for example, multiple human samples.
[0066] Any of the plurality of sequencing objects comprises a tag sequence and a nucleic acid sequence, the tag sequence (i.e. Index sequence) can be a sequence used to distinguish different biological samples, i.e. used to indicate the sample source; the sequence length of the tag sequence refers to the sequence length of the tag sequence after library preparation, and the mixed sample data comprises sequencing data after the plurality of sequencing objects are mixed for sequencing, and the sequence length of the tag sequence contained in the plurality of sequencing objects can not be completely the same, for example, the tag sequence can be a single tag sequence, or a double tag sequence, or in each single tag sequence or each double tag sequence, the number of bases contained in each tag sequence can also not be completely the same.
[0067] In a possible implementation, before starting sequencing, the Index sequence in the sequencing object can be padded to the same length to facilitate simultaneous sequencing.
[0068] Specifically, when padding the Index sequence, the length of the Index sequence of each sequencing object in the sequencing object can be determined first, and the Index sequence of other sequencing objects can be padded according to the longest length.
[0069] For example, if the longest length is a double Index sequence with a length of 8+8 (i.e. each end of the Index sequence contains 8 bases), the single Index sequence in other sequencing objects can be padded to a double Index sequence, i.e. the sequence length is padded to 8+8.
[0070] Alternatively, when padding the Index sequence of the sequencing object, the padding can be performed according to a preset fixed sequence length, and here, the preset fixed sequence length is greater than or equal to the length of the Index sequence of each sequencing object.
[0071] It should be noted that for any Index sequence, after the Index sequence is padded, the padded Index sequence is also not completely the same as the Index sequence of other sequencing objects, and here the Index sequence of other sequencing objects includes the unpadded Index sequence and other padded Index sequences.
[0072] When padding the Index sequence of the sequencing object, other bases connected to the Index sequence can be taken as part of the Index sequence to complete sequencing. For example, if the Index sequence with a sequence length of 6 is padded to an Index sequence with a sequence length of 8, two bases connected to the Index sequence from the adapter sequence connected to the Index sequence can be taken as padding bases.
[0073] Referring to FIG. 1, a flowchart of a data splitting method provided by an embodiment of the present disclosure is shown, and the method comprises steps 101-104, wherein:
[0074] In step 101, sequencing data is obtained, which comprises sequencing data of tag sequences and sequencing data of nucleic acid sequences.
[0075] In step 102, the sequencing data of the tag sequences is cut out from the sequencing data according to at least one sequence length, to obtain at least one sequencing data of the tag sequences; wherein the at least one sequence length covers the sequence length of at least part of the tag sequences in the plurality of sequencing objects.
[0076] In step 103, the at least one sequencing data of the tag sequences is matched with reference tag sequences respectively, and the sample source of the sequencing data is determined according to the matching result.
[0077] In step 104, at least one sequencing data derived from the same biological sample is split from the plurality of sequencing data according to the identified sample source.
[0078] The following is a detailed description of the above steps.
[0079] The sequencing data can comprise a plurality of bases arranged in sequence. Since the sequencing data is the sequencing data corresponding to a plurality of sequencing objects of a mixed biological sample, the sequence length of the tag sequences contained in the sequencing data of different sequencing objects is not completely the same. It is impossible to determine which part of the sequencing data is the sequencing data of the tag sequences and which part is the sequencing data of the nucleic acid sequences from the sequencing data alone.
[0080] In a specific implementation, the sequence length of the tag sequences can comprise the sequence length of a single tag sequence, the sequence length of a double tag sequence, etc., and the sequence length of different single tag sequences or different double tag sequences is also not completely the same. When cutting the sequencing data of the tag sequences, it can be cut according to a fixed position, for example, it can be cut from the tail end of the sequencing data.
[0081] Optionally, when cutting the sequencing data of the tag sequences from the sequencing data obtained by sequencing, it can be done in any one of the following two ways:
[0082] Method one: cutting first and then matching.
[0083] Specifically, the sequence length of the tag sequence can be multiple, for any sequencing data obtained by sequencing, the sequencing data can be split multiple times according to multiple sequence lengths to obtain multiple possible sequencing data of the tag sequence corresponding to the sequencing data, and then the multiple possible sequencing data of the tag sequence obtained by splitting is matched with the reference tag sequence to determine the matching result.
[0084] For example, if the sequence length of the tag sequence includes 8+8, 6+6, and 6, i.e., a double tag sequence containing 8 bases+8 bases, a double tag sequence containing 6 bases+6 bases, and a single tag sequence containing 6 bases, for sequencing data K, the sequencing data K can be split three times to obtain sequencing data of three possible tag sequences contained in the sequencing data K, and then the sequencing data of the three tag sequences is matched with the reference tag sequence one by one.
[0085] Method two, split and match at the same time.
[0086] Specifically, the splitting process of the multiple sequencing data of the multiple sequencing objects can be performed simultaneously, therefore, in order to improve the matching efficiency, the sequence length of the tag sequence can be sorted first, for example, from long to short, and then the sequence length corresponding to the sorting order is split in turn, after each splitting, the sequencing data of the split tag sequence is matched with the reference tag sequence, for the sequencing data of the tag sequence that is not matched, the next splitting is performed on the sequencing data to which it belongs, and for the sequencing data of the tag sequence that is matched, the subsequent splitting of the sequencing data to which it belongs is not required, so as to reduce the number of splitting.
[0087] For example, if the sequence length of the tag sequence includes 8+8, 6+6, and 6, the sequence length of 8+8 can be used to split each sequencing data first to obtain the first sequencing data of the tag sequence in each sequencing data, and then the sequencing data of the tag sequence is matched with the reference tag sequence, for the sequencing data of the tag sequence that is not matched, the sequence length of 6+6 is used to continue splitting the sequencing data to which it belongs to obtain the second sequencing data of the tag sequence in each sequencing data, and so on, until the splitting according to all sequence lengths is completed, or the sequencing data of the tag sequence in all sequencing data is successfully matched.
[0088] Here, the sorting of the sequence length of the tag sequence can refer to sorting according to the number of bases contained in the sequence length, or random sorting, and the disclosure is not limited to other sorting methods.
[0089] It should be noted that the above segmentation process obtains at least one sequencing data of the tag sequence, and if the tag sequence has multiple sequencing data, the multiple sequencing data of the tag sequence are all or part of possible sequencing data of the tag sequence, and the real sequencing data of the tag sequence is most likely one of the multiple sequencing data.
[0090] In one possible implementation, the reference tag sequence can refer to a standard (real) tag sequence added to a library (also referred to as a sequencing object) when the library is prepared. For example, if the tag sequence added to the sequencing object A is tag sequence 1, the tag sequence 1 is one of the reference tag sequences. Similarly, the tag sequence added to other sequencing objects when the library is prepared will also be one of the reference tag sequences.
[0091] In one possible implementation, when the at least one sequencing data of the tag sequence is matched with the reference tag sequence respectively, for the at least one sequencing data of any tag sequence, each sequencing data of the tag sequence can be matched with the reference tag sequence one by one.
[0092] Due to the sequencing accuracy of the sequencing object in the sequencing process, some bases may be detected incorrectly, and further, when the at least one sequencing data of the tag sequence is matched with the reference tag sequence, the matching may fail.
[0093] Therefore, in another possible implementation, when the at least one sequencing data of the tag sequence is matched with the reference tag sequence respectively, the at least one sequencing data of the tag sequence can be matched with the reference tag sequence and the fault-tolerant tag sequence of the reference tag sequence respectively, to expand the possibility of fault tolerance.
[0094] In one possible implementation, when the fault-tolerant tag sequence corresponding to each reference tag sequence is determined, for any reference tag sequence, at most a preset number of bases in the reference tag sequence can be replaced with other base types except the base type currently corresponding to the base position, to obtain the fault-tolerant tag sequence corresponding to the reference tag sequence, wherein the preset number is a base fault-tolerant value.
[0095] Here, the preset number can be related to the fault-tolerant accuracy, and generally, the preset number can be set to 1. When the base type is replaced, all possible replacement base types can be exhausted.
[0096] For example, if a reference tag sequence contains 8 bases and the base error tolerance value is 1, the reference tag sequence can have 8 base positions with error tolerance, and each position can have 4 possible error types, so the number of error-tolerant tag sequences corresponding to the reference tag sequence is 8*4=32. Therefore, the number of error-tolerant tag sequences corresponding to each reference tag sequence is N*4, where N is the number of bases contained in the reference tag sequence.
[0097] For example, if a reference tag sequence is ATACGA, the first base position of the error-tolerant tag sequence can be any one of the other three bases except A, or can be an unknown base N, for example, TTACGA, CTACGA, GTACGA, NTACGA, and the remaining base positions are similar.
[0098] Generally, the number of different bases between each reference tag sequence is greater than the above-mentioned base error tolerance value, so after determining the error-tolerant tag sequences of each reference tag sequence, there will also be differences between the error-tolerant tag sequences, avoiding the situation that the sequencing data of the same tag sequence matches multiple reference tag sequences or multiple error-tolerant tag sequences.
[0099] In a possible implementation, in order to improve the matching speed, the matching can be performed by searching a hash table.
[0100] Specifically, when matching the at least one sequencing data of the tag sequence with the reference tag sequence respectively, the at least one sequencing data of the tag sequence can be encoded first, converted into a key in the hash table through encoding, and at least one key to be processed is obtained; then the at least one key to be processed is queried in the hash table; wherein the key in the hash table is the encoding data of the reference tag sequence, and the value is a number indicating the sample source. Alternatively, the key in the hash table can also include the encoding data of the error-tolerant tag sequence corresponding to the reference tag sequence, and the value corresponding to the encoding data of the reference tag sequence in the hash table is the same as the value corresponding to the encoding data of the error-tolerant tag sequence in the hash table.
[0101] Here, when encoding the sequencing data of the tag sequence, the encoding can be performed by constructing a Huffman coding tree, for example.
[0102] Specifically, the frequency of occurrence of each base can be determined in advance, and the encoding value corresponding to each base is determined based on the frequency of occurrence of each base; then the sequencing data of the tag sequence is encoded based on the encoding value corresponding to each base to obtain the key to be processed.
[0103] The frequency of each base can refer to the frequency of the base in the historical sequencing process. The higher the frequency of the base, the shorter the corresponding encoding value can be. In this way, the first encoding data generated can be the shortest, thereby improving the matching speed.
[0104] For example, if the arrangement order of the frequency of each base is A > T > C > G > N, the corresponding encoding values can be A-11, T-01, G-00, C-101, and N-100, respectively. Since the frequencies of A, T, and G are higher, shorter encoding values can be used to represent them, while the frequencies of N and C are lower, and longer encoding values can be used to represent them.
[0105] Continuing the above example, if a certain tag sequence is ACATGA, the corresponding first encoding value is 1110111010011.
[0106] Correspondingly, the encoding data of the reference tag sequence and the fault-tolerant tag sequence corresponding to the reference tag sequence can also be encoded according to the respective encoding values of each base contained in the reference tag sequence and the fault-tolerant tag sequence.
[0107] Here, the process of encoding the reference tag sequence and the fault-tolerant tag sequence corresponding to the reference tag sequence can be performed in advance, i.e., before step 103 is performed.
[0108] Since the sequence lengths of each reference tag sequence are also not completely the same, after encoding the reference tag sequence and the fault-tolerant tag sequence corresponding to the reference tag sequence, the encoding data can be classified according to the sequence length. The encoding data of the reference tag sequence and the fault-tolerant tag sequence of the same sequence length are stored in the same hash table. In this way, when matching the at least one sequencing data of the tag sequence with the reference tag sequence, each hash table can be queried in sequence according to the sequence length.
[0109] For example, if the lengths of each reference tag sequence are 8+8, 6+6, and 6, respectively, three hash tables can be constructed. One is used to store the encoding data of the reference tag sequence and the fault-tolerant tag sequence corresponding to the reference tag sequence with a sequence length of 8+8. One is used to store the encoding data of the reference tag sequence and the fault-tolerant tag sequence corresponding to the reference tag sequence with a sequence length of 6+6. One is used to store the encoding data of the reference tag sequence and the fault-tolerant tag sequence corresponding to the reference tag sequence with a sequence length of 6.
[0110] When querying the at least one to-be-processed key in the hash table, each hash table can be searched in sequence according to a preset order, for example, the tag sequences of different sequence lengths can be searched according to the number of sequencing objects in which the tag sequences are located in the order of size. The sequence lengths of the reference tag sequences corresponding to the keys in the same hash table are the same.
[0111] For example, in a plurality of sequencing objects of a plurality of biological samples for mixed sequencing, if the sequence length of the tag sequence of 50 sequencing objects is 8+8, the sequence length of the tag sequence of 30 sequencing objects is 6+6, and the sequence length of the tag sequence of 10 sequencing objects is 6, the hash table corresponding to the tag sequence with the sequence length of 8+8 can be searched first. If the hash table is not found, the hash table corresponding to the tag sequence with the sequence length of 6+6 is searched. If the hash table is still not found, the hash table corresponding to the tag sequence with the sequence length of 6 is searched.
[0112] Here, the more the number of sequencing objects is, the higher the probability of the tag sequence of the sequencing object is. Therefore, searching the hash table in the order of the number of sequencing objects from large to small can reduce the number of searches as much as possible and improve the search efficiency.
[0113] In another possible implementation, when at least one to-be-processed key is queried in the hash table, K threads can be established, and the corresponding hash table is queried in parallel based on each thread and the at least one to-be-processed key, where K is the number of hash tables.
[0114] By searching the plurality of hash tables in parallel, the search efficiency of the hash table can be further improved, and the data splitting speed can be improved.
[0115] Alternatively, in another possible implementation, the number of keys contained in each hash table can be determined, and the search order of each hash table can be determined based on the number of keys. Each to-be-processed key is searched in each hash table in the search order.
[0116] The above method of determining the search order of the hash table is only illustrative, and the present disclosure is not limited to other methods of determining the search order of the hash table.
[0117] In a possible implementation, after at least one to-be-processed key is queried in the hash table, if a key matching any to-be-processed key in the at least one to-be-processed key is queried from the hash table, a mapping value is obtained according to the mapping relationship of the hash table, and the sample source of the sequencing data is determined according to the mapping value. If a key matching any to-be-processed key in the at least one to-be-processed key is not queried from the hash table, it is determined that the sample source of the sequencing data is an unknown sample.
[0118] After the sample source of the sequencing data is determined, at least one sequencing data from the same biological sample can be split from the plurality of sequencing data according to the identified sample source.
[0119] Here, after determining the sample source of the sequencing data, the sequencing data of the nucleic acid sequence and the sequencing data of the tag sequence in the sequencing data can be accurately distinguished, that is, the splitting of the sequencing data is completed.
[0120] The above data splitting method will be introduced in combination with the overall flowchart. Referring to FIG. 2, an overall flowchart of a data splitting method provided by an embodiment of the present disclosure is shown, which includes the following steps:
[0121] Step 201, identify the sequence length type of the reference tag sequence corresponding to the mixed sample data to be split according to the configuration.
[0122] Here, the sequence length type is the type of the sequence length, which may, for example, include double tag sequence 8+8, double tag sequence 6+6, single tag sequence 6, and multiple length types.
[0123] Step 202, construct different hash tables according to the sequence length types of different reference tag sequences.
[0124] Step 203, identify the part of the Index sequence in each sequencing data obtained by sequencing and encode it.
[0125] Step 204, query the encoding value of the Index sequence in the hash table.
[0126] Step 205, determine whether the query is successful.
[0127] If yes, execute step 206, and if no, execute step 207.
[0128] Step 206, determine the sample label and output the sample label.
[0129] Step 207, determine that it is an unknown sample.
[0130] For detailed description of the above steps, refer to the above embodiment, which will not be repeated here.
[0131] In the data splitting method provided by the embodiments of the present disclosure, after obtaining the sequencing data of the sequencing objects derived from multiple biological samples through sequencing, in the case that the length of the tag sequence is uncertain, the sequencing data of the tag sequence can be first cut out from the sequencing data according to at least one sequence length, and the cut-out sequencing data of the tag sequence and the reference tag sequence are matched, the reference tag sequence can be a standard tag sequence added to the sequencing objects derived from multiple biological samples during library preparation, the sample source of the sequencing data is determined based on the matching result, and the sequencing data derived from the same biological sample is split from the multiple sequencing data according to the identified sample source. Through this method, the tag sequence part of each sequencing data in the mixed sample data can be directly identified, and the efficiency and convenience of data splitting are improved.
[0132] Those skilled in the art can understand that, in the above method of the specific implementation, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process, and the specific execution order of each step should be determined by its function and possible internal logic.
[0133] Based on the same technical concept, the embodiments of the present disclosure also provide a computer device. Referring to FIG. 3, it is a structural schematic diagram of a computer device 300 provided by the embodiments of the present disclosure, which includes a processor 301, a memory 302, and a bus 303. The memory 302 is used to store execution instructions, including an internal memory 3021 and an external memory 3022; the internal memory 3021 is also called an internal storage, which is used to temporarily store operation data in the processor 301 and exchange data with the external memory 3022 such as a hard disk, the processor 301 exchanges data with the external memory 3022 through the internal memory 3021, and when the computer device 300 is running, the processor 301 and the memory 302 communicate through the bus 303, so that the processor 301 performs the splitting process of the mixed sample data, the mixed sample data includes multiple sequencing data derived from multiple sequencing objects, any sequencing object in the multiple sequencing objects includes a tag sequence and a nucleic acid sequence, the sequence length of the tag sequence in the multiple sequencing objects is not completely the same, and for any sequencing data in the multiple sequencing data, the splitting process of the mixed sample data includes:
[0134] obtaining the sequencing data, the sequencing data including sequencing data of a tag sequence and sequencing data of a nucleic acid sequence;
[0135] cutting out the sequencing data of the tag sequence from the sequencing data according to at least one sequence length, to obtain at least one sequencing data of the tag sequence; wherein the at least one sequence length covers the sequence length of at least part of the tag sequences in the multiple sequencing objects.
[0136] matching the at least one sequencing data of the tag sequence with the reference tag sequence respectively, and determining the sample source of the sequencing data according to the matching result;
[0137] splitting at least one sequencing data derived from the same biological sample from the plurality of sequencing data according to the identified sample source.
[0138] In a possible implementation, in the splitting process of the mixed sample data performed by the processor 301, the at least one sequencing data of the tag sequence is matched with the reference tag sequence respectively, and the sample source of the sequencing data is determined according to the matching result, including:
[0139] encoding the at least one sequencing data of the tag sequence respectively, converting into a key in a hash table through the encoding, and obtaining at least one to-be-processed key;
[0140] querying the at least one to-be-processed key in the hash table;
[0141] if a key matching any to-be-processed key in the at least one to-be-processed key is queried from the hash table, obtaining a mapping value according to the mapping relationship of the hash table, and determining the sample source of the sequencing data according to the mapping value;
[0142] wherein the key in the hash table is the encoding data of the reference tag sequence, and the value is a number indicating the sample source.
[0143] In a possible implementation, the processor 301 is further configured to perform:
[0144] if a key matching any to-be-processed key in the at least one to-be-processed key is not queried from the hash table, determining that the sample source of the sequencing data is an unknown sample.
[0145] In a possible implementation, the key in the hash table further includes the encoding data of the fault-tolerant tag sequence of the reference tag sequence.
[0146] the fault-tolerant tag sequence is obtained by replacing the bases at at most a preset number of base positions in the reference tag sequence with base types other than the current base type of the base positions respectively, to obtain the fault-tolerant tag sequence of the reference tag sequence; wherein the preset number is a base fault-tolerant value.
[0147] In a possible implementation, in the splitting process of the mixed sample data performed by the processor 301, the at least one sequencing data of the tag sequence is encoded respectively, converted into a key in a hash table through the encoding, and at least one to-be-processed key is obtained, including:
[0148] determining the occurrence frequency of each base, and determining the encoding value corresponding to each base based on the occurrence frequency;
[0149] encoding the at least one sequencing data of the label sequence based on the encoding value corresponding to each base, to obtain at least one processing key.
[0150] In a possible implementation, in the splitting process of the mixed sample data performed by the processor 301, the at least one processing key is queried in the hash table, including:
[0151] querying the at least one processing key in the hash table corresponding to each sequence length; wherein the reference label sequences corresponding to the keys in the same hash table have the same sequence length.
[0152] In a possible implementation, in the instructions performed by the processor 301, the querying the at least one processing key in the hash table corresponding to each sequence length includes:
[0153] establishing K threads, and querying the corresponding hash table based on each thread and the at least one processing key in parallel; wherein K is the number of the hash tables; or,
[0154] determining the number of keys contained in each hash table, and determining the searching order of each hash table based on the number of the keys, and sequentially querying each hash table based on the at least one processing key according to the searching order.
[0155] In a possible implementation, the computer device 300 is a computer device in a nucleic acid sequencer, or a computer device in wireless communication with the nucleic acid sequencer.
[0156] The disclosure also provides a computer readable storage medium, which stores a computer program. When the computer program is run by a processor, the steps of the data splitting method described in the above method embodiments are performed. The storage medium can be a volatile or non-volatile computer readable storage medium.
[0157] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system and device described above can refer to the corresponding process in the foregoing method embodiment, and will not be repeated here. In several embodiments provided in the present disclosure, it should be understood that the disclosed system, device and method can be implemented in other ways. The device embodiments described above are only schematic, for example, the division of the units is only a logical function division, and another division can be made in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some communication interfaces, devices or units, and can be electrical, mechanical or other forms.
[0158] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0159] In addition, each functional unit in each embodiment of the present disclosure can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit.
[0160] If the functions are realized in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer readable storage medium executable by a processor. Based on this understanding, the technical solutions of the present disclosure essentially or say the part of the prior art or the part of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions for making a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in each embodiment of the present disclosure. The foregoing storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), magnetic disk or optical disk, and various program code storage media.
[0161] Finally, it should be noted that the above-described embodiments are merely specific embodiments of the present disclosure, which are used to illustrate the technical solutions of the present disclosure, but not to limit the present disclosure, and the protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, it should be understood by those skilled in the art that any modification or easy-to-think change or equivalent replacement of part of the technical features of the technical solutions recorded in the foregoing embodiments can be made within the technical range disclosed by the present disclosure by those skilled in the art familiar with the technical field; and these modifications, changes or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. A computer device, characterized by, The processor, the memory and the bus, the memory stores machine readable instructions executable by the processor, when the computer equipment runs, the processor and the memory are communicated through the bus, the machine readable instructions are executed when the processor executes the splitting process of mixed sample data, the mixed sample data includes a plurality of sequencing data from a plurality of sequencing objects, any sequencing object in the plurality of sequencing objects includes a tag sequence and a nucleic acid sequence, the tag sequence is used to indicate the sample source of the sequencing object, the sequence length of the tag sequence in the plurality of sequencing objects is not completely same, for any sequencing data in the plurality of sequencing data, the splitting process of mixed sample data includes: obtain the sequencing data, the sequencing data includes the sequencing data of the tag sequence and the sequencing data of the nucleic acid sequence; The sequencing data of the tag sequence is cut out from the sequencing data according to at least one sequence length, and at least one sequencing data of the tag sequence is obtained;Wherein, the at least one sequence length covers the sequence length of at least part of the tag sequence in the plurality of sequencing objects; The at least one sequencing data of the tag sequence is matched with the reference tag sequence respectively, and the sample source of the sequencing data is determined according to the matching result; According to the identified sample source, at least one sequencing data from the same biological sample is split from the plurality of sequencing data. In the splitting process of mixed sample data executed by the processor, the at least one sequencing data of the tag sequence is matched with the reference tag sequence respectively, and the sample source of the sequencing data is determined according to the matching result, including:
2. The computer device of claim 1, wherein, The at least one sequencing data of the tag sequence is encoded respectively, and the key in the hash table is obtained by encoding, and at least one key to be processed is obtained; In the hash table, the at least one key to be processed is queried; If the key matched with any one of the at least one key to be processed is queried from the hash table, the mapping value is obtained according to the mapping relationship of the hash table, and the sample source of the sequencing data is determined according to the mapping value; Wherein, the key in the hash table is the encoding data of the reference tag sequence, and the value is the number indicating the sample source. The processor is further used to execute:
3. The computer device of claim 2, wherein, If the key matched with any one of the at least one key to be processed is not queried from the hash table, the sample source of the sequencing data is determined as unknown sample. The key in the hash table also includes the encoding data of the fault tolerant tag sequence of the reference tag sequence; The fault tolerant tag sequence is obtained by replacing the bases at most preset number of base positions in the reference tag sequence with other base types except the base type currently corresponding to the base position, wherein the preset number is the base fault tolerance value.
4. The computer device of claim 2, wherein, In the splitting process of mixed sample data executed by the processor, the at least one sequencing data of the tag sequence is encoded respectively, and the key in the hash table is obtained by encoding, and at least one key to be processed is obtained, including: 5. The computer device of claim 2 or 4, wherein, determining the occurrence frequency of each base, and determining the encoding value corresponding to each base based on the occurrence frequency; encoding each of the at least one sequencing data of the tag sequence based on the encoding value corresponding to each base, to obtain at least one to-be-processed key.
6. The computer device of claim 2 or 4, wherein, In the splitting process of the mixed sample data performed by the processor, the at least one to-be-processed key is queried in the hash table, including: querying the at least one to-be-processed key in each hash table corresponding to a sequence length; wherein the keys in the same hash table correspond to the same sequence length of the reference tag sequence.
7. The computer device of claim 6, wherein, In the splitting process of the mixed sample data performed by the processor, the at least one to-be-processed key is queried in each hash table corresponding to a sequence length, including: establishing K threads, and querying the corresponding hash table based on each thread and the at least one to-be-processed key in parallel; wherein K is the number of hash tables; or, determining the number of keys contained in each hash table, and determining the search order of each hash table based on the number of keys, and sequentially querying each hash table based on the at least one to-be-processed key according to the search order.
8. The computer device of claim 1, wherein, The computer device is a computer device on-board a nucleic acid sequencer, or a computer device in wireless communication with a nucleic acid sequencer.
9. A data splitting method, characterized by, The data splitting method is applied to mixed sample data, the mixed sample data includes a plurality of sequencing data from a plurality of sequencing objects, any sequencing object in the plurality of sequencing objects includes a tag sequence and a nucleic acid sequence, the tag sequence is used to indicate the sample source of the sequencing object, the sequence lengths of the tag sequences in the plurality of sequencing objects are not completely the same, and for any sequencing data in the plurality of sequencing data, the data splitting method includes: obtaining the sequencing data, the sequencing data including sequencing data of a tag sequence and sequencing data of a nucleic acid sequence; cutting out at least one sequencing data of the tag sequence from the sequencing data according to at least one sequence length, to obtain at least one sequencing data of the tag sequence; wherein the at least one sequence length covers the sequence length of at least part of the tag sequences in the plurality of sequencing objects; matching the at least one sequencing data of the tag sequence with a reference tag sequence respectively, and determining the sample source of the sequencing data according to the matching result; splitting at least one sequencing data from the same biological sample from the plurality of sequencing data according to the identified sample source.
10. The method of claim 9, wherein, The matching the at least one sequencing data of the tag sequence with a reference tag sequence respectively, and determining the sample source of the sequencing data according to the matching result, includes: encoding each of the at least one sequencing data of the tag sequence, converting into a key in a hash table through encoding, to obtain at least one to-be-processed key; querying the at least one to-be-processed key in the hash table; if a key matching any of the at least one to-be-processed key is queried from the hash table, obtaining a mapping value according to the mapping relationship of the hash table, and determining the sample source of the sequencing data according to the mapping value; The keys in the hash table are encoded data of reference label sequences, and the values are numbers indicating sample sources.
11. The method of claim 10, wherein, Further comprising: If no key matching any of the at least one to-be-processed key is found in the hash table, it is determined that the sample source of the sequencing data is an unknown sample.
12. The method of claim 10, wherein, The keys in the hash table further include encoded data of fault-tolerant label sequences of the reference label sequences. The fault-tolerant label sequence is obtained by replacing the bases at up to a preset number of base positions in the reference label sequence with base types other than the base type currently corresponding to the base position, respectively; wherein the preset number is a base fault-tolerant value.
13. The method according to claim 10 or 12, characterized in that, The encoding of the at least one sequencing data of the label sequence, respectively, to obtain at least one to-be-processed key, comprises: Determining the occurrence frequency of each base, and determining the encoding value corresponding to each base based on the occurrence frequency; Encoding the at least one sequencing data of the label sequence based on the encoding value corresponding to each base, respectively, to obtain at least one to-be-processed key.
14. The method of claim 10 or 12, wherein, The query of the at least one to-be-processed key in the hash table comprises: Querying based on the at least one to-be-processed key in the hash table corresponding to each sequence length; wherein the keys in the same hash table correspond to reference label sequences with the same sequence length.
15. The method of claim 14, wherein, The query based on the at least one to-be-processed key in the hash table corresponding to each sequence length comprises: Establishing K threads, and querying the corresponding hash table in parallel based on each thread and the at least one to-be-processed key; wherein K is the number of hash tables; or Determining the number of keys contained in each hash table, and determining the search order of each hash table based on the number of keys, and sequentially querying each hash table based on the at least one to-be-processed key according to the search order.
16. A computer readable storage medium characterized by: The computer readable storage medium stores a computer program, which, when executed by a processor, performs the steps of the data splitting method according to any one of claims 9 to 15.
Citation Information
Patent Citations
Method and device for determining sample source of reading segments in mixed sequencing data
CN104232760A
Tag sequence library mixing method and device for improving sequencing platform library resolution rate
CN108018607A
Nucleic acid sequence detection method and device, computer equipment and storage medium
CN115862735A
Hash index-based short nucleic acid sequence full-length matching method and system
CN117393048A
Data splitting method, computer equipment and computer storage medium
CN118248219A