Methods, devices, and software products for index sequence orientation identification in gene sequencing

By automatically identifying and correcting the orientation of index sequences in gene sequencing, the problem of complex index sequence orientation determination is solved, improving the accuracy of sample splitting and operational efficiency, and simplifying the user operation process.

CN122090949APending Publication Date: 2026-05-26SHANGHAI SAILU LIFE SCIENCES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI SAILU LIFE SCIENCES CO LTD
Filing Date
2025-07-07
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

In existing technologies, determining the orientation of the index sequence during gene sequencing is complex, leading to a high error rate for customers. Furthermore, the index orientation rules are inconsistent under different sequencing modes and library construction methods, affecting the accuracy of sample splitting.

Method used

By acquiring the design index sequence and sequencing index sequence data, a reference index sequence is obtained through cluster analysis or quality screening. A transformation operation is then performed to match the transformed index sequence with the reference index sequence, and the indexing direction of the design index sequence is automatically identified.

Benefits of technology

It eliminates the need for users to manually determine index direction rules, automatically identifies and corrects index direction, improves the accuracy and efficiency of sample splitting, can promptly detect database creation failures, and simplifies the operation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122090949A_ABST
    Figure CN122090949A_ABST
Patent Text Reader

Abstract

This invention discloses a method, device, and program product for index sequence orientation identification in gene sequencing. The method includes: acquiring a designed index sequence and sequencing index sequence data; obtaining a reference index sequence based on the sequencing index sequence data; performing a transformation operation based on the designed index sequence to obtain a transformed index sequence; and matching the transformed index sequence under various transformation modes with the reference index sequence to determine the index orientation of the designed index sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of gene technology, and in particular to an index sequence orientation identification method, computing device, and computer program product for gene sequencing. Background Technology

[0002] Currently, most sequencing instrument companies require customers to input index information according to their predefined rules. Users need to transform the original designed index sequence according to the index orientation before inputting it into the sample information. This greatly increases the workload for customers and is prone to errors, such as converting or pasting the designed index sequence, leading to incorrect information. Moreover, the index orientation may differ under the same library preparation method and sequencing mode, making the situation quite complex and prone to customer errors. Summary of the Invention

[0003] To address the existing technical problems, this invention provides a method, computing device, and computer program product for identifying the index sequence orientation in gene sequencing. This method can automatically identify the index orientation of the designed index sequence without requiring the user to manually determine complex index orientation rules.

[0004] In a first aspect, a method for identifying the orientation of an index sequence for gene sequencing is provided, comprising: acquiring a designed index sequence and sequencing index sequence data; obtaining a reference index sequence based on the sequencing index sequence data; performing a transformation operation based on the designed index sequence to obtain a transformed index sequence; and matching the transformed index sequence under various transformation modes with the reference index sequence to determine the index orientation of the designed index sequence.

[0005] In a second aspect, a computing device is provided, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor causes the processor to perform the steps of the index sequence orientation identification method for gene sequencing provided in the embodiments of this application.

[0006] Thirdly, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the index sequence orientation identification method for gene sequencing provided in the embodiments of this application.

[0007] This application uses sequencing index sequence data to obtain a representative reference index sequence, transforms the designed index sequence to obtain a transformed index sequence, and matches the transformed index sequence with the reference index sequence to obtain the index direction of the designed index sequence. It can automatically identify the index direction of the designed index sequence without requiring the user to judge complex index direction rules. Attached Figure Description

[0008] Figure 1 This is an application environment diagram of an index sequence orientation identification and correction method for gene sequencing in one embodiment; Figure 2 This is a flowchart of an index sequence orientation identification and correction method for gene sequencing in one embodiment; Figure 3 This is a flowchart of an index sequence orientation identification method for gene sequencing in one embodiment; Figure 4 This is a flowchart illustrating the method for determining a reference index sequence in a gene sequencing index sequence orientation identification and correction process, as described in one embodiment. Figure 5 This is a schematic diagram of an index sequence orientation identification and correction device for gene sequencing in one embodiment; Figure 6 This is a schematic diagram of the structure of a computing device in one embodiment. Detailed Implementation

[0009] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0010] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the specification of this invention is for the purpose of describing particular embodiments only and is not intended to limit the scope of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0011] In the following description, the expression “some embodiments” refers to a subset of all possible embodiments. However, it should be understood that “some embodiments” can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.

[0012] In gene sequencing, an index (also known as a barcode) is a specific sequence in the library adapter used to distinguish different samples. The introduction of an index allows multiple samples to be sequenced in parallel on the same sequencing chip. Later, algorithms are used to identify these samples and determine which index category each sample belongs to. However, sequencing modes, library preparation methods, and sequencer design can lead to inconsistencies in index orientation (i.e., the actual sequenced index sequence differs from the original designed index sequence), requiring prior correction before sample splitting.

[0013] For example, in paired-end sequencing, common indices are i7 (also known as Index1) and i5 (also known as Index2). Index1 is located at the P7 adapter end and is usually associated with sequencing Read2. Index2 is located at the P5 adapter end and is usually associated with sequencing Read1. Read1 and Read2 refer to a pair of complementary DNA fragment sequences in paired-end sequencing. Read1 is the sequence read from one end of the DNA fragment (e.g., the P5 adapter side). Read2 is the sequence read from the other end of the DNA fragment in reverse (e.g., the P7 adapter side), forming bidirectional coverage with Read1, which can eventually be spliced ​​into a longer continuous sequence. Index1 is attached to the 5' end of the P7 adapter, which is located at one end of the DNA library fragment. Index2 is attached to the 5' end of the P5 adapter, which is located at the other end of the library fragment.

[0014] Sequencing mode also affects index orientation. Index orientation is influenced by single-end (SE) and paired-end (PE) sequencing modes. Index orientation is also affected by the position of the index. For example, in PE sequencing mode, the index orientation will differ depending on whether it is before read1, between read1 and read2, or after read2.

[0015] The impact of library construction methods on index orientation. In transposase-based library construction (such as Nextera), DNA is directly cut and adapters are inserted using transposases, and the index is introduced via PCR amplification. Regarding index orientation, the designed index sequence in the transposase adapter may be reverse complementary to the actual sequencing direction. For example, the original designed index sequence for transposase adapter N701 is TCGCCTTA, but during sequencing analysis, the user needs to reverse complement the designed index sequence to TAAGGCGA, and then input TAAGGCGA into the sample sheet. Therefore, in current technologies, users need to identify the transposase adapter type and reverse complement the designed index sequence.

[0016] When constructing libraries for Y-type adapters (such as with Illumina TruSeq), for complete Y-type adapters, the index is already contained in a fixed position within the adapter, and the index orientation is fixed. For incomplete Y-type adapters, the index is introduced through PCR amplification, and the orientation may vary depending on the primer design. Therefore, in existing technologies, users need to determine whether the index orientation needs to be adjusted based on the adapter type (e.g., D-series or R-series).

[0017] When building a library with a chain-specific architecture (such as the dUTP method), single-chain information is preserved, and the indexing direction is opposite to that of conventional library building. In existing technologies, users need to dynamically adjust the indexing direction by combining chain-specific markers (such as Read1 / Read2 direction).

[0018] Different library preparation methods (such as transposase method, Y-adaptor method, and strand-specific library preparation) and sequencing modes (such as SE sequencing and PE sequencing) result in complex indexing direction rules. Therefore, directly splitting the sequencing sample will lead to problems such as sample splitting errors and low splitting rate.

[0019] like Figure 1 As shown, Figure 1 This diagram illustrates the application environment of an index sequence orientation identification and correction method for gene sequencing in one embodiment. A gene sequencer 10 communicates with a computing device 20. The gene sequencer 10 uses a sequencing chip to sequence sequencing fragments, obtaining a sequencing file. The sequencing file includes the sequencing sequence data of each sequenced fragment. After obtaining the sequencing file, the computing device 20 can obtain the sequencing index sequence for each sequencing fragment from the sequencing sequence data. The sequencing index sequence indicates the sequencing sequence of the designed index sequence within the sequencing fragment. The sequencing index sequence data includes multiple sequencing index sequences. The computing device 20 can provide a user interface through which the user can input the designed index sequence, which is the user's original designed index, i.e., the index without any transformation operations. The number of designed index sequences can be one or more. The sequencing index sequence data is obtained based on sequencing samples provided by the same user as the designed index sequence. Based on the designed index sequence and the sequencing index sequence data, the computing device 20 can automatically analyze the index orientation of the designed index sequence, facilitating subsequent automatic splitting of the sequencing samples.

[0020] Please see Figure 2 This is a flowchart illustrating a method for index sequence orientation identification and correction in gene sequencing according to an embodiment of this application. The method for index sequence orientation identification and correction in gene sequencing is applied in a computing device and includes the following steps: S11. Obtain the design index sequence and sequencing index sequence data.

[0021] In this embodiment, the designed index sequence refers to the index originally designed by the user, i.e., the index without any transformation operations. The sequencing index sequence data includes multiple sequencing index sequences. The sequencing index sequence indicates the indexed sequencing sequence within the sequencing fragment. Each designed index sequence can be multi-indexed or single-indexed. For example, a designed index sequence may be dual-indexed, such as including i7 and i5. A single index means that a designed index sequence includes one index. Sequencing index sequences can also be multi-indexed or single-indexed. The sequencing index sequence data can be at least a portion extracted from tens of millions of sequencing index sequences. The designed index sequences are one or more designed index sequences under the same library construction method, or one or more designed index sequences under the same sequencing method. When there are multiple designed index sequences, the indexing direction of each index sequence is the same. For example, if there are two designed index sequences, both sequences will include dual indices, and the indexing direction of index1 in the first designed index sequence is the same as the indexing direction of index1 in the second designed index sequence. Similarly, the index direction of index2 in the first design index sequence is the same as the index direction of index2 in the second design index sequence. That is, the index direction of index1 in the first sequence and the index direction of index2 form the index direction of the first design index sequence, and the index direction of index1 in the second sequence and the index direction of index2 form the index direction of the second design index sequence.

[0022] S12. Based on the sequencing index sequence data, obtain the reference index sequence.

[0023] In this embodiment, representative sequencing index sequences are selected from the sequencing index sequence data to obtain reference index sequences. For example, cluster analysis can be performed on all sequencing index sequences in the sequencing index sequence data to obtain reference index sequences. Alternatively, sequencing index sequences with high quality scores can be selected from the sequencing index sequence data to obtain reference index sequences. There are multiple reference index sequences. The number of types of reference index sequences is the same as the number of types of designed index sequences. For example, if there are three types of designed index sequences, then there are also three types of reference index sequences.

[0024] S13. Based on the designed index sequence, perform transformation operations to obtain the transformed index sequence.

[0025] In this embodiment, the transformation operations include, but are not limited to, forward, reverse, complementary, and reverse-complementary methods. The forward method means that the design index sequence remains unchanged, i.e., the transformed index sequence is the same as the design index sequence. The reverse method means that the base order of the design index sequence is flipped from the end to the beginning, but the bases at each position remain unchanged. For example, the design index sequence is ATGC, and the transformed index sequence is CGTA. The complementary method means that a new sequence complementary to the design index sequence is generated according to the base pairing principle (AT / U, CG), while the orientation remains unchanged. For example, the design index sequence is ATGC, and the transformed index sequence is TACG. The reverse-complementary method means that the design index sequence is first reversed to obtain the reversed sequence, and then a sequence complementary to the reversed sequence is generated. The final result is that both the orientation and the bases are completely opposite to the original sequence. For example, the design index sequence is ATGC, and the transformed index sequence is GCAT.

[0026] For example, when designing an index sequence as a single index, there are 4 possible transformation index sequences. When designing an index sequence as a double index, there are 16 possible transformation index sequences.

[0027] S14. Match the transformation index sequences under various transformation methods with the reference index sequence to determine the index direction of the design index sequence.

[0028] In this embodiment, the sequencing fragment contains a design index sequence. During sequencing, the design index sequence is transformed due to the influence of the sequencing direction. After sequencing, the sequencing index sequence is obtained. The reference index sequence is essentially a transformed version of the design index sequence. Since the reference index sequence is already a representative sequencing sequence and is also a transformed version of the design index sequence, the transformed index sequences under various transformation methods can be matched with the reference index sequence to determine the transformation method corresponding to the maximum matching degree, thereby obtaining the index direction of the design index sequence. When there are multiple design index sequences, there are also multiple transformed index sequences for each transformation method.

[0029] S15. Obtain the sequencing index sequence to be identified. Based on the index direction, correct the designed index sequence or the sequencing index sequence to be identified to obtain the corrected index sequence. Match the corrected index sequence with the designed index sequence or the sequencing index sequence to be identified. Based on the matching result, determine the identification result of the sequencing index sequence to be identified.

[0030] In this embodiment, the sequencing index sequence to be identified is the index sample that needs to be classified and identified. The designed index sequence can be corrected to obtain a corrected designed index sequence. This corrected designed index sequence is then used as the corrected index sequence. The corrected index sequence is matched with each sequencing index sequence to be identified, and the identification result of the sequencing index sequence to be identified is determined based on the matching results. Alternatively, the sequencing index sequence to be identified can be corrected to obtain a corrected sequencing index sequence. This corrected sequencing index sequence is then used as the corrected index sequence. The corrected sequencing index sequence is then matched with each designed index sequence, and the identification result of the sequencing index sequence to be identified is determined based on the matching results.

[0031] In the above embodiments, a representative reference index sequence is obtained based on the sequencing index sequence data. The designed index sequence is transformed to obtain a transformed index sequence. The transformed index sequence under each transformation method is matched with the reference index sequence to obtain the index direction of the designed index sequence. The index direction of the designed index sequence can be automatically identified, eliminating the need for users to judge complex index direction rules, thus facilitating the automatic splitting of sequencing samples.

[0032] It is understandable that in the embodiments of the index sequence orientation identification and correction method for gene sequencing described above, after determining the index orientation of the designed index sequence by using the designed index sequence transformation to obtain the transformed index sequence and matching it with the reference index sequence determined based on the sequencing index sequence data, the determined index orientation can also be used to transform the sequences in the sequencing sequence files obtained from gene sequencing of the same sample, and perform bioinformatics analysis such as alignment and detection based on the transformed results. On the other hand, when matching the transformed index sequences obtained from different transformation operations with the reference index sequence to determine the index orientation of the designed index sequence, in addition to the transformed index sequence corresponding to the correct index orientation forming a matching value with the reference index sequence, transformed index sequences in other orientations also have a certain probability of forming a matching value. If the library construction quality is high, then even if the matching value of the transformed index sequences in other orientations with the reference index sequence is non-zero, the matching value is usually very small. Therefore, if two or more transformed index sequences have non-zero matching values ​​with the reference index sequence, and both configuration values ​​exceed a certain threshold, it indicates that a database creation error may have occurred, resulting in a large number of transformed index sequences in directions other than the correct index direction. Therefore, this situation can be used to indicate a database creation failure, making it easier for users to detect and resolve database creation failures as early as possible.

[0033] Therefore, determining the index direction, besides its application in the embodiments of index sequence direction identification and correction methods for gene sequencing, which utilizes S15 for correction and matching to split the sequencing index sequence to be identified, can also be used for bioinformatics analysis such as sequence alignment and detection in sequencing sequence files of the same sample, or for applications such as verifying the quality of the current library. It is evident that, in the embodiments of this application, determining the index direction itself constitutes a comprehensive technical solution to address the complex index direction rules resulting from different library construction methods. That is, please refer to... Figure 3 In another aspect, this application also provides a method for identifying the orientation of index sequences for gene sequencing, including steps S11 to S14.

[0034] Furthermore, it is understood that the main difference between the index sequence orientation identification method for gene sequencing and the index sequence orientation identification and correction method for gene sequencing in this application embodiment is that the index sequence orientation identification and correction method for gene sequencing provides a means of splitting the sequencing index sequence to be identified through step S15. Therefore, the index sequence orientation identification and correction method for gene sequencing, as well as the embodiments for determining the index orientation, are also applicable to the index sequence orientation identification method for gene sequencing, and will not be repeated in the subsequent description of the embodiments in this specification.

[0035] In some embodiments, clustering methods can be used to obtain a reference index sequence. For example... Figure 4 As shown, step S12 can specifically include: S121. Obtain the current centroid sequence.

[0036] In this embodiment, the number of current centroid sequences is configured to a preset number of centroid sequences. This preset number of centroid sequences can be user-configurable. The design index sequence is designed and input by the user, who knows the number of sequences. The clustering methods include, but are not limited to, k-Means-based distance methods, hierarchical clustering methods, etc. The preset number of centroid sequences represents the number of clusters required. A current centroid sequence is the current center of a category. The current centroid sequences of each category can be iteratively updated using clustering methods, i.e., updating each category. The preset number of centroid sequences is represented by K.

[0037] S122. For each sequencing index sequence in the sequencing index sequence data, calculate the distance between the sequencing index sequence and each of the current centroid sequences, and assign the sequencing index sequence to the sequencing index sequence cluster where the current centroid sequence with the smallest distance is located, thereby obtaining each sequencing index sequence cluster.

[0038] Optionally, calculating the distance between the sequencing index sequence and each of the current centroid sequences includes: For each base position, the first base corresponding to the base position is obtained from the sequencing index sequence, and the second base corresponding to the base position in the current centroid sequence is obtained, and the first base and the second base are compared; When the first base and the second base are the same, the score corresponding to the base position is counted as the first preset score; when the first base and the second base are different, the score corresponding to the base position is counted as the second preset score. Based on the scores corresponding to each base position, the distance between the sequencing index sequence and the current centroid sequence is obtained.

[0039] In this embodiment, for a sequencing index sequence, the distance between the sequencing index sequence and a current centroid sequence can be calculated using the method described above. The distances between the sequencing index sequence and K current centroid sequences are calculated respectively, resulting in K distances. The sequencing index sequence is then assigned to the cluster to which the nearest current centroid sequence belongs. The method for calculating the distance between the sequencing index sequence and a current centroid sequence is as follows: The sequencing index sequence and the current centroid sequence are two sequences of equal length. The sequencing index sequence indexA=(s1,s2,…, sn) and the current centroid sequence KA=(k1, k2,…, kn) are defined as the Hamming distance DH(indexA, KA) between them, which is defined as the sum of the counts of dissimilar characters at the same base positions in the two sequences. in, This is an indicator function, representing the score corresponding to the base position. equal ,but It equals the first preset score, for example, equal to 0. If Not equal to ,but It equals the second preset score, for example, equal to 1.

[0040] For example, let's take indexA (ACGTTTGGA) and KA (AGGTTTGAG) as examples: so, The answer is 0+1+0+0+0+0+0+1+1=3.

[0041] S123. For each sequencing index sequence cluster, perform statistics on each sequencing index sequence in the sequencing index sequence cluster to obtain the updated current centroid sequence. Use the updated current centroid sequence as the current centroid sequence of the sequencing index sequence cluster, return to continue, and use the current centroid sequence after stopping the iteration as the reference index sequence.

[0042] In this embodiment, the updated current centroid sequence is used as the centroid sequence for the next iteration. For example, if the updated current centroid sequence is obtained in the first iteration, the current centroid sequence in the second iteration is the updated current centroid sequence obtained in the first iteration. One sequencing index sequence cluster represents a category. After obtaining the updated current centroid sequence, the process returns to obtain the current centroid sequence, assigns the updated current centroid sequence as the current centroid sequence, and executes subsequent steps. For any sequencing index sequence cluster, the updated current centroid sequence can be obtained in two ways. One way is to count the mode at each base position.

[0043] Optionally, for each sequencing index sequence cluster, the step of statistically analyzing each sequencing index sequence in the sequencing index sequence cluster to obtain the updated current centroid sequence includes: For each sequencing index sequence cluster, the number of base types at each base position in the sequencing index sequence cluster is counted. For each base position, the base corresponding to the maximum number of base types is taken as the base at the base position in the updated current centroid sequence, thus obtaining the updated current centroid sequence.

[0044] The length of the sequencing index sequence is denoted as `indexLen`. There are a total of `indexNum` sequencing index sequences in the sequencing index sequence cluster. The base at the i-th base position (`indexPos_i`) of the current centroid sequence is updated. The bases at `indexPos_i` of all sequencing index sequences (`indexNum`) in the sequencing index sequence cluster are counted, and the base with the most frequent occurrences is used as the new centroid base at `indexPos_i`. For example, if base A appears most frequently at the third base position, then base A is used as the base at the third base position in the updated current centroid sequence.

[0045] Alternatively, the step of statistically analyzing each sequencing index sequence within a sequencing index sequence cluster to obtain the updated current centroid sequence includes: For each sequencing index sequence cluster, the frequency of occurrence of each sequencing index sequence in the sequencing index sequence cluster is calculated; Based on the occurrence frequency, target sequencing index sequences that meet the occurrence frequency criteria are selected, and the sequence distance between each target sequencing index sequence and each sequencing index sequence in the sequencing index sequence cluster is calculated. The minimum sequence distance and the corresponding target sequencing index sequence are used as the updated current centroid sequence.

[0046] The length of the sequencing index sequence is represented by `indexLen`, and there are a total of `indexNum` sequences in the sequencing index sequence cluster. The frequencies of the `indexNum` sequencing index sequences are counted and sorted from highest to lowest. The sequencing index sequences with the highest frequencies are selected, and the highest frequencies are chosen. The highest frequencies can be adjusted, for example, 10 (`indexMax1`, `indexMax2`, ..., `indexMax10`). For any `indexMaxi`, the Hamming distances between `indexMaxi` and each of the sequences in the sequencing index sequence cluster (`indexNum`) are calculated. The sum of the Hamming distances of `indexNum` is then collected as `disi`, which is used as the sequence distance sum corresponding to `indexMaxi`. This yields 10 sequence distance sums. The `indexMaxi` with the smallest sequence distance sum is then selected as the target sequencing index sequence. This target sequencing index sequence is used as the updated current centroid sequence for this sequencing index sequence cluster.

[0047] In the above embodiments, based on the clustering method, representative reference index sequences are found through iterative calculation among numerous sequencing index sequences. During the clustering process, the base information or the frequency of sequence occurrence in the sequencing index sequence clusters are statistically analyzed, which can more accurately update the current centroid sequence. This results in more accurate sequencing index sequence clusters, improves the accuracy of obtaining reference index sequences, and thus improves the accuracy of finding the index direction of the designed index sequence.

[0048] In some embodiments, obtaining the reference index sequence based on the sequencing index sequence data includes: Obtain the quality score of each sequencing index sequence in the sequencing index sequence data; Based on the quality score, sequencing index sequences that meet the quality score criteria are selected, and the reference index sequence is obtained based on the sequencing index sequences that meet the quality score criteria.

[0049] In this embodiment, the sequencing file may also include quality assessment data for each sequencing sequence, for example, expressed as a quality score. Sequencing index sequences with quality scores higher than a preset quality score value are selected as sequencing index sequences that meet the quality score criteria. For example, the lowest quality score of the sequencing index sequences is greater than 20, the average quality score is greater than 30, and so on.

[0050] In some embodiments, matching the transformation index sequence under various transformation modes with the reference index sequence to determine the index direction of the design index sequence includes: Calculate the matching degree between the transformation index sequence and each of the reference index sequences under various transformation methods to obtain the matching degree corresponding to each transformation method; From the matching degrees corresponding to each of the transformation methods, the maximum matching degree is determined, and the maximum matching degree is compared with the original matching degree to determine the index direction of the design index sequence, wherein the original matching degree represents the matching degree corresponding to the forward method.

[0051] In this embodiment, when the indexing direction is forward, it means that the designed index sequence is the same as the sequence obtained by sequencing the designed index sequence. The original matching degree is the matching degree corresponding to the transformed index sequence in the forward direction.

[0052] There are multiple ways to calculate the matching degree between the transform index sequence and each reference index sequence. Optionally, calculating the matching degree between the transform index sequence and each reference index sequence under various transform modes to obtain the matching degree corresponding to each transform mode includes: For each transformation method, the distance between the transformation index sequence and each of the reference index sequences under the transformation method is calculated to obtain multiple distances corresponding to each transformation index sequence. From the multiple distances corresponding to each transformation index sequence, the minimum distance corresponding to each transformation index sequence is obtained. The minimum distances corresponding to each of the transformation index sequences are summed to obtain the minimum distance sum, which is used as the matching degree corresponding to the transformation method.

[0053] In this embodiment, the distance calculation method between the transformed index sequence and each of the reference index sequences is the same as the Hamming distance calculation method described above, and will not be repeated here.

[0054] For example, if there are three design index sequences, then there are also three transformation index sequences under the reverse transformation method, namely transformation index sequences A1, A2, and A3, and three reference index sequences, namely reference index sequences B1, B2, and B3. Then the distances between A1 and B1 are C11, between A1 and B2 are C12, and between A1 and B3 are C13, with C12 being the smallest; the distances between A2 and B1 are C21, between A2 and B2 are C22, and between A2 and B3 are C23, with C23 being the smallest; the distances between A3 and B1 are C31, between A3 and B2 are C32, and between A3 and B3 are C33, with C33 being the smallest; therefore, the matching degree corresponding to the reverse transformation method is C12 + C23 + C33.

[0055] Optionally, another method can be used to calculate the matching degree between the transformation index sequence and each of the reference index sequences under various transformation modes, and to obtain the matching degree corresponding to each transformation mode, including: For each transformation method, the distance between the transformation index sequence and each of the reference index sequences under the transformation method is calculated to obtain multiple distances corresponding to each transformation index sequence; The number of distances that are preset distance values ​​among the multiple distances corresponding to each of the transformation index sequences is counted, and the number is used as the matching degree corresponding to the transformation method.

[0056] In this embodiment, for example, if there are three design index sequences, then there are also three transformation index sequences under the reverse transformation method, namely transformation index sequences A1, A2, and A3, and three reference index sequences, namely reference index sequences B1, B2, and B3. The distances between A1 and B1 are then calculated as follows: A1 = 1, A1 = 0, and A1 = 2; A2 = 0, A2 = 3, and A2 = 1; A3 = 1, A3 = 1, and A3 = 1. If the preset distance value is 0, and 0 appears twice, then the matching degree corresponding to the reverse transformation method is 2.

[0057] Optionally, comparing the maximum matching degree with the original matching degree to determine the indexing direction of the designed index sequence includes: When the maximum matching degree is greater than or equal to a preset multiple of the original matching degree, the transformation method corresponding to the maximum matching degree is used as the index direction of the designed index sequence; When the maximum matching degree is less than a preset multiple of the original matching degree, the forward direction is used as the index direction of the designed index sequence.

[0058] In this embodiment, the preset multiple is a number greater than 1. When the maximum matching degree is greater than or equal to the preset multiple of the original matching degree, it indicates that the maximum matching degree is much greater than the original matching degree. For example, the maximum matching degree... If 0.8 > the original matching degree, it indicates that the transformation method corresponding to the maximum matching degree is the optimal transformation method. The transformation method corresponding to the maximum matching degree is used as the index direction of the designed index sequence. When the designed index sequence includes multiple indices, the transformation method corresponding to the maximum matching degree includes the index direction of each index. If the maximum matching degree is less than a preset multiple of the original matching degree, it indicates that the maximum matching degree is not significantly greater than the original matching degree, and the forward direction is used as the index direction of the designed index sequence.

[0059] On the one hand, after determining the indexing direction of the designed index sequence, the designed index sequence can be transformed, or the sequencing index sequence to be identified can be transformed. Based on the transformed results, bioinformatics analyses such as alignment and detection can be performed.

[0060] On the other hand, when matching the transformed index sequences under each transformation method with the reference index sequence to determine the index direction of the designed index sequence, the matching degree corresponding to each transformation method can be determined. That is, in addition to determining the matching degree corresponding to the transformation method representing the correct index direction, the matching degree corresponding to other transformation methods can also be determined. If the database construction quality is high, the matching degree corresponding to other transformation methods, even if it is non-zero, is usually very small. Therefore, if at least two transformation methods have non-zero matching degrees, and at least two transformation methods have matching degrees exceeding a certain threshold, it indicates that a database construction error may have occurred, resulting in a large number of transformed index sequences in directions other than the correct index direction. Therefore, this situation can be used to indicate a database construction failure, allowing users to detect and resolve database construction problems as early as possible.

[0061] In the above embodiments, the matching degree between the transformed index sequence and each of the reference index sequences is calculated to obtain the matching degree corresponding to each of the transformed index sequences. From the matching degrees corresponding to each of the transformed index sequences, the maximum matching degree is determined. Based on the transformation method corresponding to the maximum matching degree, the index direction of the designed index sequence is determined. The index direction of the designed index sequence can be automatically identified, eliminating the need for users to judge complex index direction rules themselves, thereby facilitating the automatic splitting of sequencing samples.

[0062] In some embodiments, the Obtain the sequencing index sequence to be identified; based on the index direction, correct the designed index sequence or the sequencing index sequence to be identified to obtain a corrected index sequence; match the corrected index sequence with the designed index sequence or the sequencing index sequence to be identified; and determine the identification result of the sequencing index sequence to be identified based on the matching result, including: According to the index direction, the designed index sequence is corrected to obtain a corrected designed index sequence. The similarity between the corrected designed index sequence and each of the sequencing index sequences to be identified is calculated to obtain the similarity corresponding to each sequencing index sequence to be identified. Sequencing index sequences to be identified that meet the first similarity condition are identified as index samples along with the designed index sequence; or According to the index direction, the sequencing index sequence to be identified is corrected to obtain the corrected sequencing index sequence. The similarity between the corrected sequencing index sequence and each of the designed index sequences is calculated to obtain the similarity corresponding to each of the designed index sequences. The designed index sequences whose similarity meets the second similarity condition and the sequencing index sequence to be identified are identified as a class of index samples.

[0063] In this embodiment, the first or second similarity condition includes, but is not limited to: similarity exceeding a preset similarity threshold, and similarity ranked in the top preset positions from high to low. When the designed index sequence includes multiple indices, the index direction includes the index direction of each index. Based on the index direction of each index, the corresponding indices in the designed index sequence are transformed to obtain the corrected designed index sequence. For example, the designed index sequence given by the customer is index1: ACGTTGCG, index2: CGTTAGGT. The transformation corresponding to the maximum matching degree is that index1 is in the forward direction, i.e., unchanged, and index2 is in the reverse complementary direction. Then, the designed index sequence given by the customer is transformed into: index1New: ACGTTGCG, index2New: ACCTACG. Then, index1New and index2New are combined as the alignment sequence and compared with the sequencing sequence obtained by the gene sequencer to see which designed index sequence each read belongs to. For example, the read corresponding to the sequencing sequence that is the same as both index1New and index2New is identified as the sample of the index1+index2 designed index sequence.

[0064] In the above embodiments, the designed index sequence can be transformed according to the index direction, and the sequencing index sequence to be identified can also be transformed. The index direction of the designed index sequence can be automatically identified, eliminating the need for users to judge complex index direction rules, thereby facilitating the automatic splitting of sequencing samples.

[0065] In another aspect, this application provides a computer program product, including a computer program that, when executed by a processor, implements the index sequence orientation identification and correction method for gene sequencing as described in any embodiment of this application, or implements the index sequence orientation identification method for gene sequencing as described in any embodiment of this application.

[0066] In the computer program product, the optional implementation form of the program module architecture of the computer program that implements the index sequence orientation recognition and correction method for gene sequencing, or the steps of the index sequence orientation recognition method for gene sequencing, can be an index sequence orientation recognition and correction device for gene sequencing.

[0067] Please see Figure 5One embodiment of this application provides an index sequence orientation identification and correction device for gene sequencing, comprising: an acquisition module 41 for acquiring a design index sequence and sequencing index sequence data; a screening module 42 for obtaining a reference index sequence based on the sequencing index sequence data; a transformation module 43 for performing a transformation operation based on the design index sequence to obtain a transformed index sequence; and a determination module 44 for matching the transformed index sequence under various transformation modes with the reference index sequence to determine the index orientation of the design index sequence.

[0068] Optionally, the filtering module 42 is also used for: Get the current centroid sequence; For each sequencing index sequence in the sequencing index sequence data, calculate the distance between the sequencing index sequence and each current centroid sequence, and assign the sequencing index sequence to the sequencing index sequence cluster where the current centroid sequence with the smallest distance is located, thereby obtaining each sequencing index sequence cluster. For each sequencing index sequence cluster, the sequencing index sequences in the sequencing index sequence cluster are statistically analyzed to obtain the updated current centroid sequence. The updated current centroid sequence is used as the current centroid sequence of the sequencing index sequence cluster, and the process returns to continue. The current centroid sequence after stopping the iteration is used as the reference index sequence.

[0069] Optionally, the filtering module 42 is also used for: For each base position, the first base corresponding to the base position is obtained from the sequencing index sequence, and the second base corresponding to the base position in the current centroid sequence is obtained, and the first base and the second base are compared; When the first base and the second base are the same, the score corresponding to the base position is counted as the first preset score; when the first base and the second base are different, the score corresponding to the base position is counted as the second preset score. Based on the scores corresponding to each base position, the distance between the sequencing index sequence and the current centroid sequence is obtained.

[0070] Optionally, the filtering module 42 is also used for: For each sequencing index sequence cluster, the number of base types at each base position in the sequencing index sequence cluster is counted. For each base position, the base corresponding to the maximum number of base types is taken as the base at the base position in the updated current centroid sequence, thus obtaining the updated current centroid sequence.

[0071] Optionally, the filtering module 42 is also used for: For each sequencing index sequence cluster, the frequency of occurrence of each sequencing index sequence in the sequencing index sequence cluster is calculated; Based on the occurrence frequency, target sequencing index sequences that meet the occurrence frequency criteria are selected, and the sequence distance between each target sequencing index sequence and each sequencing index sequence in the sequencing index sequence cluster is calculated. The minimum sequence distance and the corresponding target sequencing index sequence are used as the updated current centroid sequence.

[0072] Optionally, the filtering module 42 is also used for: Obtain the quality score of each sequencing index sequence in the sequencing index sequence data; Based on the quality score, sequencing index sequences that meet the quality score criteria are selected, and the reference index sequence is obtained based on the sequencing index sequences that meet the quality score criteria.

[0073] Optionally, the determination module 44 is also used for: Calculate the matching degree between the transformation index sequence and each of the reference index sequences under various transformation methods to obtain the matching degree corresponding to each transformation method; From the matching degrees corresponding to each of the transformation methods, the maximum matching degree is determined, and the maximum matching degree is compared with the original matching degree to determine the index direction of the design index sequence, wherein the original matching degree represents the matching degree corresponding to the forward method.

[0074] Optionally, the determination module 44 is also used for: For each transformation method, the distance between the transformation index sequence and each of the reference index sequences under the transformation method is calculated to obtain multiple distances corresponding to each transformation index sequence. From the multiple distances corresponding to each transformation index sequence, the minimum distance corresponding to each transformation index sequence is selected. The minimum distances corresponding to each of the transformation index sequences are summed to obtain the minimum distance sum, which is used as the matching degree corresponding to the transformation method.

[0075] Optionally, the determination module 44 is also used for: For each transformation method, the distance between the transformation index sequence and each of the reference index sequences under the transformation method is calculated to obtain multiple distances corresponding to each transformation index sequence; The number of distances that are preset distance values ​​among the multiple distances corresponding to each of the transformation index sequences is counted, and the number is used as the matching degree corresponding to the transformation method.

[0076] Optionally, the determination module 44 is also used for: When the maximum matching degree is greater than or equal to a preset multiple of the original matching degree, the transformation method corresponding to the maximum matching degree is used as the index direction of the designed index sequence; When the maximum matching degree is less than a preset multiple of the original matching degree, the forward direction is used as the index direction of the designed index sequence.

[0077] Optionally, the split module 45 is also used for: According to the index direction, the designed index sequence is corrected to obtain a corrected designed index sequence. The similarity between the corrected designed index sequence and each of the sequencing index sequences to be identified is calculated to obtain the similarity corresponding to each sequencing index sequence to be identified. Sequencing index sequences to be identified that meet the first similarity condition are identified as index samples along with the designed index sequence; or According to the index direction, the sequencing index sequence to be identified is corrected to obtain the corrected sequencing index sequence. The similarity between the corrected sequencing index sequence and each of the designed index sequences is calculated to obtain the similarity corresponding to each of the designed index sequences. The designed index sequences whose similarity meets the second similarity condition and the sequencing index sequence to be identified are identified as a class of index samples.

[0078] It will be understood by those skilled in the art that Figure 5 The structure of the index sequence orientation recognition and correction device for gene sequencing does not constitute a limitation on the device itself. Each module can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independently of the controller in the computing device, or stored in software in the memory of the computing device, so that the controller can invoke and execute the operations corresponding to each module. In other embodiments, the index sequence orientation recognition and correction device for gene sequencing may include more or fewer modules than those shown in the figure.

[0079] Please see Figure 6 In another aspect of this application, a computing device 20 is also provided, including a memory 3011 and a processor 3012. The memory 3011 stores a computer program. When the computer program is executed by the processor, the processor 3012 performs the steps of the index sequence orientation recognition and correction method for gene sequencing provided in any of the above embodiments of this application, or performs the steps of the index sequence orientation recognition method for gene sequencing provided in any of the above embodiments of this application. The computing device may include a desktop computer, a laptop computer, a tablet computer, a handheld computer, a smart speaker, a server, etc., a mobile phone (e.g., a smartphone, a cordless phone, etc.), a wearable device (e.g., a pair of smart glasses or a smartwatch), or similar devices.

[0080] The processor 3012 is the control center, connecting various parts of the computing device via various interfaces and lines. It executes software programs and / or modules stored in the memory 3011, and calls data stored in the memory 3011 to perform various functions and process data. Optionally, the processor 3012 may include one or more processing cores; the processor 3012 includes, but is not limited to, one or more combinations of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), Field-Programmable Gate Array (FPGA), etc. Preferably, the processor 3012 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user page, and applications, and the modem processor mainly handles wireless communication. It is understood that the aforementioned modem processor may not be integrated into the processor 3012.

[0081] The memory 3011 can be used to store software programs and modules. The processor 3012 executes various functional applications and data processing by running the software programs and modules stored in the memory 3011. The memory 3011 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the computing device, etc. In addition, the memory 3011 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 3011 may also include a memory controller to provide the processor 3012 with access to the memory 3011.

[0082] In another aspect, this application provides a storage medium storing a computer program. When the computer program is executed by a processor, the processor performs the index sequence orientation identification and correction method for gene sequencing provided in any of the above embodiments of this application, or performs the steps of the index sequence orientation identification method for gene sequencing provided in any of the above embodiments of this application.

[0083] Those skilled in the art will understand that all or part of the processes in the methods provided in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0084] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. The scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for index sequence orientation identification in gene sequencing, characterized in that, include: Obtain the design index sequence and sequencing index sequence data; Based on the sequencing index sequence data, a reference index sequence is obtained; Based on the designed index sequence, a transformation operation is performed to obtain the transformed index sequence; The transformation index sequence under various transformation methods is matched with the reference index sequence to determine the index direction of the design index sequence.

2. The index sequence orientation identification method for gene sequencing as described in claim 1, characterized in that, The process of obtaining the reference index sequence based on the sequencing index sequence data includes: Get the current centroid sequence; For each sequencing index sequence in the sequencing index sequence data, calculate the distance between the sequencing index sequence and each current centroid sequence, and assign the sequencing index sequence to the sequencing index sequence cluster where the current centroid sequence with the smallest distance is located, thereby obtaining each sequencing index sequence cluster. For each sequencing index sequence cluster, the sequencing index sequences in the sequencing index sequence cluster are statistically analyzed to obtain the updated current centroid sequence. The updated current centroid sequence is used as the current centroid sequence of the sequencing index sequence cluster, and the process returns to continue. The current centroid sequence after stopping the iteration is used as the reference index sequence.

3. The index sequence orientation identification method for gene sequencing as described in claim 2, characterized in that, The calculation of the distance between the sequencing index sequence and each of the current centroid sequences includes: For each base position, the first base corresponding to the base position is obtained from the sequencing index sequence, and the second base corresponding to the base position in the current centroid sequence is obtained, and the first base and the second base are compared; When the first base and the second base are the same, the score corresponding to the base position is counted as the first preset score; when the first base and the second base are different, the score corresponding to the base position is counted as the second preset score. Based on the scores corresponding to each base position, the distance between the sequencing index sequence and the current centroid sequence is obtained.

4. The index sequence orientation identification method for gene sequencing as described in claim 2, characterized in that, For each sequencing index sequence cluster, the updated current centroid sequence is obtained by statistically analyzing each sequencing index sequence within the cluster. For each sequencing index sequence cluster, the number of base types at each base position in the sequencing index sequence cluster is counted. For each base position, the base corresponding to the maximum number of base types is taken as the base at the base position in the updated current centroid sequence, thus obtaining the updated current centroid sequence.

5. The index sequence orientation identification method for gene sequencing as described in claim 2, characterized in that, For each sequencing index sequence cluster, the updated current centroid sequence is obtained by statistically analyzing each sequencing index sequence within the cluster. For each sequencing index sequence cluster, the frequency of occurrence of each sequencing index sequence in the sequencing index sequence cluster is calculated; Based on the occurrence frequency, target sequencing index sequences that meet the occurrence frequency criteria are selected, and the sequence distance between each target sequencing index sequence and each sequencing index sequence in the sequencing index sequence cluster is calculated. The minimum sequence distance and the corresponding target sequencing index sequence are used as the updated current centroid sequence.

6. The index sequence orientation identification method for gene sequencing as described in claim 1, characterized in that, The process of obtaining the reference index sequence based on the sequencing index sequence data includes: Obtain the quality score of each sequencing index sequence in the sequencing index sequence data; Based on the quality score, sequencing index sequences that meet the quality score criteria are selected, and the reference index sequence is obtained based on the sequencing index sequences that meet the quality score criteria.

7. The index sequence orientation identification method for gene sequencing as described in claim 1, characterized in that, The step of matching the transformation index sequence under various transformation modes with the reference index sequence to determine the index direction of the design index sequence includes: Calculate the matching degree between the transformation index sequence and each of the reference index sequences under various transformation methods to obtain the matching degree corresponding to each transformation method; From the matching degrees corresponding to each of the transformation methods, the maximum matching degree is determined, and the maximum matching degree is compared with the original matching degree to determine the index direction of the design index sequence, wherein the original matching degree represents the matching degree corresponding to the forward method.

8. The index sequence orientation identification method for gene sequencing as described in claim 7, characterized in that, The calculation of the matching degree between the transformation index sequence and each reference index sequence under various transformation modes, to obtain the matching degree corresponding to each transformation mode, includes: For each transformation method, the distance between the transformation index sequence and each of the reference index sequences under the transformation method is calculated to obtain multiple distances corresponding to each transformation index sequence. From the multiple distances corresponding to each transformation index sequence, the minimum distance corresponding to each transformation index sequence is selected. The minimum distances corresponding to each of the transformation index sequences are summed to obtain the minimum distance sum, which is used as the matching degree corresponding to the transformation method.

9. The index sequence orientation identification method for gene sequencing as described in claim 7, characterized in that, The calculation of the matching degree between the transformation index sequence and each reference index sequence under various transformation modes, to obtain the matching degree corresponding to each transformation mode, includes: For each transformation method, the distance between the transformation index sequence and each of the reference index sequences under the transformation method is calculated to obtain multiple distances corresponding to each transformation index sequence; The number of distances that are preset distance values ​​among the multiple distances corresponding to each of the transformation index sequences is counted, and the number is used as the matching degree corresponding to the transformation method.

10. The index sequence orientation identification method for gene sequencing as described in claim 7, characterized in that, The step of comparing the maximum matching degree with the original matching degree to determine the indexing direction of the designed index sequence includes: When the maximum matching degree is greater than or equal to a preset multiple of the original matching degree, the transformation method corresponding to the maximum matching degree is used as the index direction of the designed index sequence; When the maximum matching degree is less than a preset multiple of the original matching degree, the forward direction is used as the index direction of the designed index sequence.

11. A computing device, characterized in that, It includes a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the method as described in any one of claims 1 to 10.

12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 10.