Method and device for detecting igk gene rearrangement, electronic equipment and storage medium

By assembling and aligning paired-end sequencing data, and combining the IGKV, IGKJ, Kde, and J_C_intron gene libraries, the shortcomings of existing tools in IGK gene rearrangement detection have been addressed, enabling automated detection of IGK gene rearrangements. This is suitable for monitoring minimal residual disease and recurrence analysis in lymphoma, improving the accuracy and efficiency of detection.

CN117133357BActive Publication Date: 2026-08-25BOE TECHNOLOGY GROUP CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210552015.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-18
Publication Date
2026-08-25
Estimated Expiration
2042-05-18

AI Technical Summary

Technical Problem

Existing gene rearrangement detection tools lack detection schemes for VJ gene rearrangement, V-Kde gene rearrangement, and J_C_intron-Kde gene rearrangement of the IGK gene, and cannot be effectively used for the diagnosis of B-cell lymphoma and the differentiation between polyclonal reactive hyperplasia and malignant proliferative diseases.

Method used

By using assembly and alignment techniques based on paired-end sequencing data, and utilizing IGKV, IGKJ, Kde, and J_C_intron gene libraries, we can automatically detect VJ gene rearrangements, V-Kde gene rearrangements, and J_C_intron-Kde gene rearrangements in the IGK gene. Combined with high-throughput sequencing technology and majority voting correction methods, we can improve the accuracy and precision of sequencing data.

Benefits of technology

It enables automated detection of IGK gene rearrangements, suitable for monitoring minimal residual disease and relapse in lymphoma, improving detection accuracy and efficiency, and supporting downstream analysis of immune repertoire sequencing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117133357B_ABST
    Figure CN117133357B_ABST
Patent Text Reader

Abstract

The application discloses a kind of IGK gene rearrangement detection method, device, electronic equipment and storage medium, detection method includes: obtaining the first end sequencing sequence and second end sequencing sequence of test sample;First end sequencing sequence and second end sequencing sequence are assembled based on, obtain assembly sequence;Based on assembly sequence, determine target alignment gene from gene reference database;Gene reference database includes IGKV gene library, IGKJ gene library, Kde gene library and J_C_intron gene library, target alignment gene includes at least one of target V gene, target J gene, target Kde gene and target J_C_intron gene;Based on target alignment gene, determine IGK gene rearrangement result in assembly sequence.The above-mentioned method can automatically detect VJ gene rearrangement, V-Kde gene rearrangement and J_C_intron-Kde gene rearrangement in IGK.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of gene detection technology, and in particular to a method, apparatus, electronic device, and storage medium for detecting IGK gene rearrangements. Background Technology

[0002] When pluripotent hematopoietic stem cells differentiate into lymphocytes, gene rearrangements occur. Each lymphocyte has a unique rearranged gene sequence; that is, normal lymphocyte genes exhibit polyclonal rearrangements. However, lymphoma cells and their progeny cells are from the same clone, sharing the same gene coding. Tumor cell DNA amplification samples show a specific band in a particular region during electrophoresis. In contrast, amplification samples from lymphocytes in patients with reactive lymph node hyperplasia and normal individuals show diffuse bands during electrophoresis. Current research indicates that gene rearrangements are acquired gene damage. Lymphoma cells are formed by the monoclonal proliferation of cells with gene abnormalities, thus exhibiting monoclonal alterations. This monoclonal gene rearrangement can serve as a specific molecular marker for detecting B-cell lymphoma and is used in the diagnosis of B-cell lymphoma. Furthermore, the detection of this clonal characteristic helps differentiate between polyclonal reactive hyperplasia and malignant proliferative diseases.

[0003] Studies have shown that immunoglobulin Kappa (IGK) gene rearrangements are found in 60% of B-cell acute lymphoblastic leukemia (B-ALL) cases in children, and these IGK gene rearrangements are related to the deletion and rearrangement of the Kappa deletion element (Kde gene). The recombination signal sequence of the Kde gene is located approximately 24 kb downstream of the C gene fragment. Specific types of Kde rearrangements include: 1) V-Kde rearrangement: The Kde recombination signal sequence can rearrange into the V gene fragment, leading to deletions of the J and C genes; 2) J_C_intron-Kde rearrangement: The recombination signal sequence in the intron between the J and C genes rearranges with the Kde gene recombination signal sequence, leading to deletion of the C gene.

[0004] Current gene rearrangement detection tools include MiGEC, Mixcr, and IgBlast, but they are all for identifying V(D)J gene rearrangements in genes such as IGH, IGK, TRB, and TRD. They lack a solution for detecting VJ gene rearrangements, V-Kde gene rearrangements, and J_C_intron-Kde gene rearrangements in the IGK gene. Summary of the Invention

[0005] In view of the above problems, the present invention proposes a method, apparatus, electronic device and storage medium for detecting IGK gene rearrangements, so as to solve or partially solve the technical problem of how to detect VJ gene rearrangement, V-Kde gene rearrangement and J_C_intron-Kde gene rearrangement in IGK gene rearrangement.

[0006] In a first aspect, the present invention provides a method for detecting IGK gene rearrangements through an embodiment, comprising: Obtain paired-end sequencing data of the test sample; the paired-end sequencing data includes the first-end sequencing sequence and the second-end sequencing sequence; The assembled sequence is obtained by assembling based on the first end sequencing sequence and the second end sequencing sequence. Based on the assembled sequence, a target alignment gene is determined from a target gene reference database; wherein, the gene reference database includes the IGKV gene library, IGKJ gene library, Kde gene library and J_C_intron gene library in germ cell lines, and the target alignment gene includes at least one of the target V gene, target J gene, target Kde gene and target J_C_intron gene; Based on the target alignment gene, the rearrangement result of the IGK gene in the assembled sequence is determined.

[0007] In some optional embodiments, the first end sequencing sequence includes a plurality of first read sequences, and the second end sequencing sequence includes a plurality of second read sequences; The assembly based on the first end sequencing sequence and the second end sequencing sequence to obtain the assembled sequence includes: Traverse the first read length sequence to determine the first similar read length sequence corresponding to the first read length sequence; perform majority voting based on each group of the first read length sequence and the first similar read length sequence to obtain the first end-corrected sequence; and traverse the second read length sequence to determine the second similar read length sequence corresponding to the second read length sequence; perform majority voting based on each group of the second read length sequence and the second similar read length sequence to obtain the second end-corrected sequence. The assembled sequence is obtained by assembling based on the first end correction sequence and the second end correction sequence.

[0008] In some optional embodiments, obtaining the first end-corrected sequence by majority voting based on each group of the first read length sequence and the first similar read length sequence includes: Based on each group of the first read length sequence and the first similar read length sequence, the number of similarities is determined; when the number of similarities is greater than a set value, a majority vote is performed on each base in the first read length sequence and the first similar read length sequence to obtain the first corrected read length sequence; based on all the first corrected read length sequences, the first end-corrected sequence is obtained. The step of obtaining the second end-corrected sequence by majority voting based on each group of the second read length sequence and the second similar read length sequence includes: Based on each group of the second read sequence and the second similar read sequence, the number of similarities is determined; when the number of similarities is greater than a set value, a majority vote is performed on each base in the second read sequence and the second similar read sequence to obtain the second corrected read sequence; based on all the second corrected read sequences, the second end-corrected sequence is obtained.

[0009] In some optional embodiments, after obtaining the first end-corrected sequence and the second end-corrected sequence, the detection method further includes: The connector sequence in the first corrected read sequence is removed to obtain a first preprocessed read sequence, and a first end preprocessed sequence is obtained based on all the first preprocessed read sequences; and the connector sequence in the second corrected read sequence is removed to obtain a second preprocessed read sequence, and a second end preprocessed sequence is obtained based on all the second preprocessed read sequences. The assembly based on the first end-corrected sequence and the second end-corrected sequence to obtain the assembled sequence includes: The assembled sequence is obtained by assembling based on the first end preprocessing sequence and the second end preprocessing sequence.

[0010] In some optional embodiments, after obtaining the first end preprocessed sequence and the second end preprocessed sequence, the detection method further includes: Delete a first preprocessed read sequence whose length is less than a first set length to obtain a first-end sequence to be assembled; and delete a second preprocessed read sequence whose length is less than the first set length to obtain a second-end sequence to be assembled. The assembly based on the first end preprocessing sequence and the second end preprocessing sequence to obtain the assembled sequence includes: The assembly sequence is obtained by assembling based on the first end sequence to be assembled and the second end sequence to be assembled.

[0011] In some optional embodiments, the first set length ranges from 10bp to 100bp.

[0012] In some optional embodiments, the assembly based on the first end sequence to be assembled and the second end sequence to be assembled to obtain the assembled sequence includes: Obtain the inverse complementary read sequence of the second preprocessed read sequence; Based on the first preprocessed read length sequence and the reverse complementary read length sequence, an overlapping sequence is determined; When the length of the overlapping sequence is not less than the second preset length, the overlapping sequence in the reverse complementary read length sequence is deleted to obtain the read length sequence to be assembled. The first preprocessed read length sequence is concatenated with the read length sequence to be assembled to obtain the assembled read length sequence; The assembled sequence is obtained based on all the assembled read sequences.

[0013] In some optional embodiments, determining the target alignment gene from the target gene reference database based on the assembled sequence includes: Based on the set alignment parameters, the target alignment gene corresponding to each assembled read sequence is determined from the target gene reference database; The set alignment parameters include: the similarity between the alignment fragment in the assembled read sequence and the target alignment gene is not less than 90%, and the length of the alignment fragment ranges from 4 to 11.

[0014] In some optional embodiments, when the target alignment gene includes only the target V gene and the target J gene, determining the IGK gene rearrangement result in the assembled sequence based on the target alignment gene includes: The nucleotide positions of phenylalanine residues in the target J gene are obtained, and the termination point in the assembled sequence is determined based on the nucleotide positions; Cysteine ​​residues in the assembled sequence are detected within a set range before the termination point, and the position of the cysteine ​​residue closest to the termination point is taken as the starting point; the set range is the assembled sequence fragment from the termination point to 60 bp to 90 bp before the termination point. Based on the starting point and the ending point, the CDR3 region in the assembly sequence is determined.

[0015] In some optional embodiments, when the target alignment gene includes only the target V gene and the target J gene, determining the IGK gene rearrangement result in the assembled sequence based on the target alignment gene includes: Cluster analysis was performed on the assembled sequence based on the target V gene and the target J gene to obtain the number of clone sequences and the proportion of clone sequences in the assembled sequence.

[0016] Secondly, the present invention provides an IGK gene rearrangement detection device through one embodiment, comprising: An acquisition module is used to obtain paired-end sequencing data of the test sample; the paired-end sequencing data includes a first-end sequencing sequence and a second-end sequencing sequence. An assembly module is used to assemble based on the first end sequencing sequence and the second end sequencing sequence to obtain an assembled sequence; The alignment module is used to determine the target alignment gene from the target gene reference database based on the assembled sequence; wherein the gene reference database includes the IGKV gene library, IGKJ gene library, Kde gene library and J_C_intron gene library in germ cell lines, and the target alignment gene includes at least one of the target V gene, target J gene, target Kde gene and target J_C_intron gene; The determination module is used to determine the IGK gene rearrangement result in the assembled sequence based on the target alignment gene.

[0017] Thirdly, the present invention provides an electronic device through one embodiment, including a processor and a memory, the memory being coupled to the processor, the memory storing instructions that, when executed by the processor, cause the electronic device to perform the steps of any of the detection methods described in the first aspect embodiment.

[0018] Fourthly, the present invention provides, through one embodiment, a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the detection method described in any one of the first aspect embodiments.

[0019] The method for detecting IGK gene rearrangements provided by this invention involves assembling an assembled sequence based on the first and second end sequencing sequences from the raw paired-end sequencing data. This assembled sequence is then compared with reference gene sequences in germline IGKV, IGKJ, Kde, and J_C_intron gene libraries to identify target alignment genes, including at least one of the target V, J, Kde, and J_C_intron genes. Based on these target alignment genes, the IGK gene rearrangement result in the assembled sequence is determined. This method provides an automated workflow for detecting VJ, V-Kde, and J_C_intron-Kde gene rearrangements in the IGK gene, suitable for downstream analysis and identification needs such as monitoring minimal residual disease and recurrence in lymphoma, and immune repercussion sequencing.

[0020] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings: Figure 1 A schematic flowchart of the detection method provided in the first aspect embodiment of the present invention is shown; Figure 2 This diagram illustrates the length distribution of the assembled read sequence provided in a first aspect embodiment of the present invention. Figure 3 A schematic diagram of a detection device provided in a second aspect embodiment of the present invention is shown; Figure 4 A schematic diagram of an electronic device provided in a third aspect embodiment of the present invention is shown; Figure 5 A schematic diagram of a computer-readable storage medium provided in a fourth aspect embodiment of the present invention is shown. Detailed Implementation

[0022] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0023] To detect VJ gene rearrangement, V-Kde gene rearrangement, and J_C_intron-Kde gene rearrangement in IGK gene rearrangement, this invention provides a method for detecting IGK gene rearrangement, the overall concept of which is as follows: Obtain paired-end sequencing data of the test sample; the paired-end sequencing data includes the first-end sequencing sequence and the second-end sequencing sequence; assemble the sample based on the first-end sequencing sequence and the second-end sequencing sequence to obtain the assembled sequence; based on the assembled sequence, determine the target alignment gene from the target gene reference database; wherein, the gene reference database includes the IGKV gene library, IGKJ gene library, Kde gene library and J_C_intron gene library in germ cell lines, and the target alignment gene includes at least one of the target V gene, target J gene, target Kde gene and target J_C_intron gene; based on the target alignment gene, determine the IGK gene rearrangement result in the assembled sequence.

[0024] The above scheme assembles an assembled sequence based on the first and second end sequencing sequences from the raw paired-end sequencing data. This assembled sequence is then compared with reference gene sequences in germline IGKV, IGKJ, Kde, and J_C_intron gene libraries to identify target alignment genes, including at least one of the target V, J, Kde, and J_C_intron genes. Based on these target alignment genes, the IGK gene rearrangement in the assembled sequence is determined. This method provides an automated workflow for detecting VJ, V-Kde, and J_C_intron-Kde gene rearrangements in the IGK gene, suitable for downstream analysis and identification needs such as minimal residual disease (MRD) and relapse monitoring in lymphoma, and immune repertoire sequencing.

[0025] The following sections will provide further details with reference to specific implementation methods.

[0026] Explanation of some key English terms involved in the specific implementation: BCR: known as B cell antigen receptor, is an immunoglobulin molecule (IG) that grows on the surface of B lymphocytes. It consists of two heavy chains (IGH) and two light chains (IGK or IGL).

[0027] The IGK light chain is composed of a constant region (C gene) and a variable region (V gene, J gene). Human IGK gene: located on the short arm of chromosome 2 (2p11.2), containing C gene, Kde gene and multiple V and J genes.

[0028] CDR3 region: The region in the variable region that determines the target for identification, containing the end of the V gene and the beginning of the J gene.

[0029] In the embodiments of the first aspect, a method for detecting IGK gene rearrangements based on high-throughput sequencing technology or "next-generation" sequencing technology (NGS) is provided. Please refer to [link to relevant documentation]. Figure 1 The process includes steps S1 to S4, as detailed below: S1: Obtain paired-end sequencing data of the test sample; paired-end sequencing data includes the first-end sequencing sequence and the second-end sequencing sequence; Specifically, the test sample is a lymphocyte sample. After nucleic acid extraction, library construction and other steps, the test sample is sent to a high-throughput sequencer to obtain paired-end sequencing data.

[0030] Paired-end sequencing involves sequencing a single strand of deoxyribonucleic acid (DNA) in both the forward and reverse directions. In this embodiment, the first-end sequencing sequence represents the nucleic acid sequence obtained by sequencing along the first direction of the DNA during paired-end sequencing, and the second-end sequencing sequence represents the nucleic acid sequence obtained by sequencing along the second direction of the DNA during paired-end sequencing. The first and second directions are opposite; for example, the first direction can be from left to right, and the second direction can be from right to left.

[0031] Taking a commonly used high-throughput sequencing data characterization standard as an example, the information of the first and second end sequencing sequences is stored in separate FASTQ files, mainly for preserving the base sequence and sequencing quality. The base sequence and sequencing quality are represented using ASCII encoding.

[0032] A fastq file stores multiple reads; a read is a single long sequence, also known as a sequencing short sequence, which is the base sequence obtained by a high-throughput sequencer in a single sequencing run.

[0033] Therefore, the first end sequencing sequence includes multiple first-read sequences, and the second end sequencing sequence includes multiple second-read sequences.

[0034] For ease of description and distinction, in this embodiment of the invention, the first-end sequencing sequence and the subsequent processed sequence are uniformly labeled as Read1, abbreviated as R1, and the second-end sequencing sequence and the subsequent processed sequence are uniformly labeled as Read2, abbreviated as R2; the reads in R1 are marked as r 1i Mark the reads in R2 as r 2i ; 1≤i≤N and are integers; where i is the read number and N is the number of reads included in the first-end sequencing sequence or the second-end sequencing sequence.

[0035] S2: Assemble the sequence based on the first and second end sequencing sequences to obtain the assembled sequence; The assembly sequence is obtained by assembling or splicing the first read from the first end of the sequencing sequence with the second read from the second end of the sequencing sequence according to the correspondence of gene sequencing or the sequence number ID of the reads, so as to obtain a complete assembled sequence. Existing tools such as Pear or Pandaseq can be used during assembly, and no specific limitation is made here.

[0036] In some optional embodiments, before assembly, the first and second end sequencing sequences are subjected to data quality checks and preprocessing to remove low-quality reads and obtain high-quality data for assembly, thereby improving the accuracy of subsequent target gene alignment.

[0037] One option for quality control and preprocessing is to correct the paired-end sequencing data, as follows: After obtaining the paired-end sequencing data of the test sample, the first read sequence is traversed to determine the first similar read sequence corresponding to the first read sequence; based on each pair of the first read sequence and the first similar read sequence, a majority vote is performed to obtain the first end-corrected sequence; the second read sequence is traversed to determine the second similar read sequence corresponding to the second read sequence; based on each pair of the second read sequence and the second similar read sequence, a majority vote is performed to obtain the second end-corrected sequence.

[0038] Specifically, the similarity between each read in the first end of the sequencing sequence and other reads is calculated, and other reads with similarity greater than a set threshold are taken as similar read length sequences.

[0039] For example, for the first read length sequence r in R1 11 Calculate r sequentially 11 With r 12 r 13 , ..., r 1N The similarity between them is then used to determine the similarity r between them that is greater than a set threshold. 1j As r 11 Similar read length sequences. Similarly, determine r sequentially. 12 r 13 , ..., r 1N The corresponding similar read sequences. Methods for calculating the similarity between base sequences are existing technology and will not be elaborated upon here. The threshold value is determined based on requirements and is not specifically limited here.

[0040] Next, a majority vote is performed on each group of first-read sequences and the first similar read sequence corresponding to that group. Majority voting involves finding the majority element in an array of n elements and replacing the minority elements with the majority element; the majority element is defined as the element that appears more than [n / 2] times in the array. This majority voting method corrects the first-end sequencing sequence, resulting in a first-end corrected sequence. This reduces sequencing errors during gene sequencing, improving the accuracy of subsequent gene alignment and analysis.

[0041] Optionally, each base in the reads can be used as a voting element for each group of first-length reads r. 1i and r 1i The corresponding first similar read length sequence r 1j The bases at the same position are taken out one by one and a majority vote is performed. Then the base determined by the majority vote is used as the corrected base at that position.

[0042] Furthermore, taking the first end sequencing sequence as an example, an optional method to obtain the first end corrected sequence by majority voting based on each group of first read sequences and first similar read sequences is as follows: Based on each set of first read length sequences and first similar read length sequences, the number of similarities is determined; when the number of similarities is greater than a set value, a majority vote is performed on each base in the first read length sequence and the first similar read length sequence to obtain the first corrected read length sequence; based on all the first corrected read length sequences, the first end-corrected sequence is obtained.

[0043] Specifically, when performing internal similarity calculations on the first-end sequencing sequences, the similarity count is determined. When a similar read sequence is found for any first-read sequence, the similarity count for that first-read sequence is automatically incremented by 1. For example, for read sequence r... 11 If similarity calculation is used to find r 11 Similar reads include: r 12 r 15 The similarity count is 2; for a read-length sequence r 12 If similarity calculation is used to find r 12 Similar reads include: r 13 r 14 and r 17 If the similarity is 3, then the number of similarities is 3.

[0044] After statistically analyzing each set of first-read sequences and their corresponding first-similar-read sequences, a majority vote is conducted on the first-read sequences and their first-similar-read sequences with a similarity count greater than a set value to obtain the corresponding first-corrected-read sequence. Conversely, the first-read sequences and their first-similar-read sequences with a similarity count less than or equal to the set value can be directly deleted and not included in subsequent sequence assembly and alignment. The set value can be between 1 and 3, with a preferred value of 2, meaning that only the first-read sequences and their first-similar-read sequences with a similarity count greater than 2 are retained for majority voting.

[0045] For example, if the total number of reads in a given set of first-length reads and first-similar-length reads is 5 > 2, then a majority vote is performed on these five reads. If the first base of the five reads is A, T, A, A, T, then according to the majority vote principle, the majority base is determined to be A, and the first base of either the first-length read or all five reads is uniformly corrected to A. Then, following the same method, a majority vote is performed on the second, third, and so on, up to the last base of each of the five reads, to determine the corrected first-length read as the first corrected read.

[0046] After completing the majority vote correction of the first read length sequence and the first similar read length sequence of all groups, the first end-corrected sequence can be obtained.

[0047] The second-end sequencing sequence is similar to the first-end sequencing sequence, as follows: Based on each group of second read sequences and second similar read sequences, the number of similarities is determined; when the number of similarities is greater than a set value, a majority vote is performed on each base in the second read sequence and the second similar read sequence to obtain the second corrected read sequence; based on all the second corrected read sequences, the second end corrected sequence is obtained.

[0048] The above method performs similarity calculations within the first and second sequencing sequences, respectively, and retains reads with a similarity count greater than a set value for majority voting. This corrects amplification errors generated during gene sequencing, thereby improving the reliability of the sequencing sequence and enhancing the accuracy of subsequent target gene alignment.

[0049] In some optional embodiments, before correcting the sequencing sequence using majority voting, reads containing unknown nucleotides (i.e., N-bases) in the first and second end sequencing sequences can be removed, as well as reads with an average base quality lower than a set quality. This further improves the data quality of the sequencing sequence and reduces the workload of correcting the sequencing sequence. The set quality can range from 20 to 25, preferably 20.

[0050] In some optional embodiments, after majority voting correction of the sequencing sequences is completed, the detection method further includes: The connector sequence in the first corrected read sequence is removed to obtain the first preprocessed read sequence, and the first end preprocessed sequence is obtained based on all the first preprocessed read sequences; the connector sequence in the second corrected read sequence is removed to obtain the second preprocessed read sequence, and the second end preprocessed sequence is obtained based on all the second preprocessed read sequences; after removing the connector sequence, the first end preprocessed sequence and the second end preprocessed sequence can be used to proceed to the subsequent assembly steps.

[0051] Specifically, adapter sequences are short, known sequences added to both ends of the target sequencing fragment during high-throughput sequencing to distinguish different test samples during mixed sequencing. Therefore, they can be removed before assembly.

[0052] Taking the first-end correction sequence as an example, the following method can be used to remove the connector sequence: By retrieving the first 4000-10000 rows of R1, adapter sequences added by different sequencing platforms were retrieved to identify and filter the adapter sequence; when a certain r was detected... 1i If the overlap between the left and right ends and the connector sequence is greater than or equal to 3 bp, then that segment is identified as the connector sequence and is removed.

[0053] In some optional embodiments, after obtaining the first end preprocessed sequence and the second end preprocessed sequence, and before assembly, the detection method further includes: The first preprocessed read sequence with a length shorter than a first set length is deleted to obtain the first end of the sequence to be assembled; and the second preprocessed read sequence with a length shorter than the first set length is deleted to obtain the second end of the sequence to be assembled.

[0054] Specifically, after removing the connector sequence, based on the lengths of reads in R1 and R2, all reads shorter than a first set length are removed. It should be noted that when a certain read in R1: r 1a When the length is lower than the first set length, delete r in R1 synchronously. 1a And in R2 with r 1a The corresponding r 2b The first set length (trim_len) is a parameter that represents the length of preprocessed single-end reads. It can be adjusted according to actual needs, with an adjustable range of 10bp to 100bp, where bp is one base pair.

[0055] Before assembly, reads with a single - end length less than the first set length in the first - end pre - processed sequence and the second - end pre - processed sequence are deleted, which can reduce the sequencing fragments in the sequencing sequence that are unrelated to the V gene, J gene, Kde gene, and J_C_intron gene. Thus, the interference of invalid gene fragments and non - target alignment gene fragments can be reduced during subsequent gene library alignment, thereby reducing the gene alignment workload and improving the gene alignment accuracy.

[0056] Next, based on the first - end to - be - assembled sequence and the second - end to - be - assembled sequence, assembly is performed to obtain an assembled sequence.

[0057] An optional assembly scheme is as follows: Obtain the reverse - complementary read length sequence of the second pre - processed read length sequence; determine the overlapping sequence according to the first pre - processed read length sequence and the reverse - complementary read length sequence; when the length of the overlapping sequence is not less than the second set length, delete the overlapping sequence in the reverse - complementary read length sequence to obtain the to - be - assembled read length sequence; splice the first pre - processed read length sequence and the to - be - assembled read length sequence to obtain the assembled read length sequence; based on all the assembled read length sequences, obtain the assembled sequence.

[0058] Specifically, all reads: r in R2 2i are transformed into their reverse - complementary reads, denoted as r 2i ’, and then r 1i is compared with r 1i corresponding to r 2i ’ to determine the overlapping sequence between the two and determine the overlapping sequence length overlap. When overlap ≥ overlap_len, the overlapping sequence in r 2i ’ is removed, and then the remaining sequences in r 1i and r 2i ’ are connected to obtain the assembled read length sequence (assembled), which is marked as query_id. Here, overlap_len is the second set length, that is, the minimum overlapping sequence length, and its selectable value range is 10bp - 40bp.

[0059] If the overlapping sequence length overlap between r 1i and r 2i ’ is less than overlap_len, then this group of r 1i and r 2i ’ are respectively saved as the assembled failure sequences (assembled_F and assembled_R), and the assembled failure sequences do not participate in the gene alignment in the subsequent steps.

[0060] The assembled sequence is obtained by concatenating all the first preprocessed read sequences with the read sequences to be assembled. The assembled sequence includes multiple assembled read sequences query_id.

[0061] S3: Based on the assembled sequence, determine the target alignment gene from the target gene reference database; wherein, the gene reference database includes the IGKV gene library, IGKJ gene library, Kde gene library and J_C_intron gene library in germ cell lines, and the target alignment gene includes at least one of the target V gene, target J gene, target Kde gene and target J_C_intron gene; Specifically, the purpose of this embodiment is to analyze the VJ gene rearrangement, V-Kde gene rearrangement, and J_C_intron-Kde gene rearrangement in the IGK gene. Each query_id sequence is locally aligned with the reference sequences of the V / J / Kde / J_C_intron genes in the germline to determine which V / J / Kde / J_C_intron genes the assembled sequence originates from, thereby determining the gene rearrangement of the assembled sequence.

[0062] One possible comparison scheme is as follows: Regarding VJ gene rearrangements: Each query_id sequence is sequentially compared with multiple V and J gene sequences of IGK in the IMGT immune repertoire data to find the target alignment gene that meets the set alignment parameters, and the gene ID is extracted and recorded as subject_id.

[0063] For V-Kde gene rearrangements and J_C_intron-Kde gene rearrangements: Each query_id sequence is sequentially compared with the IGK J_C_intron library and the Kde gene library to find the target alignment gene that meets the set alignment parameters, and the gene ID is extracted and recorded as subject_id.

[0064] Optionally, the alignment parameters can be set to include: the similarity between the aligned fragment in the assembled read sequence and the target aligned gene is not less than 90%, and the length of the aligned fragment ranges from 4 to 11. These alignment parameters can improve the speed and accuracy of deriving the target aligned genes, namely the V gene, J gene, Kde gene, and J_C_intron gene, from the gene reference database.

[0065] In practice, the blastn tool can be used. By inputting the set alignment parameters, the alignment can be performed in the IGKV, IGKJ, J-C_intron and Kde gene reference databases. If the alignment is successful, the target alignment gene ID is extracted. The alignment parameters set in the blastn tool are: 1) Similarity parameter between the alignment fragment and the target alignment gene: -perc_identity=90; 2) Length of the sequence fragment -word_size=4~11, preferably 11.

[0066] By using the above-mentioned alignment parameters, the optimal target alignment gene can be found among 114 IGKV genes, 9 IGKJ genes, as well as the Kde gene and J_C_intron gene.

[0067] S4: Based on the target alignment gene, determine the IGK gene rearrangement result in the assembled sequence.

[0068] After obtaining the target alignment gene from the gene reference database, the target alignment gene can be used to annotate the sequence fragments in the assembled sequence, thereby obtaining the rearrangement results or rearrangement status of the IGK gene, which can be used for subsequent identification and analysis of IGK gene rearrangements.

[0069] If a VJ gene rearrangement is detected in the IGK gene, it is necessary to identify the CDR3 sequence. Current methods define the CDR3 region as a sequence segment from the second conserved cysteine ​​residue at the 3' end of the V gene to the conserved phenylalanine residue in the J gene. However, studies have shown that the second conserved cysteine ​​residue may not be the last cysteine ​​residue on the V gene, necessitating a more precise method to determine the start position of the CDR3 region.

[0070] In some optional embodiments, if the target alignment gene subject_id obtained from the query_id sequence alignment only contains the V and J genes, then determining the IGK gene rearrangement result in the assembled sequence based on the target alignment gene also includes annotating its CDR3 region, as follows: The nucleotide positions of phenylalanine residues in the target J gene are obtained, and the termination point is determined in the assembled sequence based on the nucleotide positions. Cysteine ​​residues in the assembled sequence are detected within a set range before the termination point, and the position of the cysteine ​​residue closest to the termination point is taken as the start point. The set range is the assembled sequence fragment from the termination point to 60 bp to 90 bp before the termination point. The CDR3 region in the assembled sequence is determined based on the start point and the termination point.

[0071] Specifically, based on the target J gene obtained from the alignment, the nucleotide position corresponding to the phenylalanine residue "FGXG" is detected to determine the termination point of the CDR3 region in the assembled sequence. Then, cysteine ​​residues are searched within a length range of 60bp to 90bp before the CDR3 termination point, and the cysteine ​​residue closest to the termination point is taken as the start point of the CDR3 region, thereby determining the CDR3 sequence based on the start and termination points. Preferably, the assembled sequence fragment within a length range of 75bp before the termination point is set.

[0072] The above scheme determines the position of the termination point of the CDR3 region in the sequence by detecting the nucleotide position corresponding to the phenylalanine residue "FGXG" in the J gene. Then, it searches for cysteine ​​residues within a range of 60-90 bp before the termination point, and takes the last cysteine ​​residue as the start point of the CDR3 region. Searching for the cysteine ​​residue closest to the termination point within this 60-90 bp search range ensures that the found cysteine ​​residue is the last cysteine ​​residue before the phenylalanine residue, thus improving the accuracy of CDR3 region determination.

[0073] Furthermore, current gene rearrangement detection tools do not perform relevant immune repertoire functional analyses on VJ gene rearrangements in IGK, such as clonal diversity and cross-sample common clonal analysis. Therefore, it is necessary to provide an automated detection and analysis solution for VJ gene rearrangement identification and immune repertoire analysis in IGK.

[0074] In some optional embodiments, if the target alignment gene subject_id obtained from the query_id sequence alignment only contains the V and J genes, then determining the IGK gene rearrangement result in the assembled sequence based on the target alignment gene further includes performing cloning analysis, as follows: Cluster analysis was performed on the assembled sequences based on the target V and target J genes to obtain the number of clone species, the number of clone sequences, and the proportion of clone sequences in the assembled sequences. After obtaining the data on the number of clone species, the number of clone sequences, and the proportion of clone sequences, public clone analysis can be performed to further explore the relationship between the immune repertoire and diseases.

[0075] To more intuitively illustrate the clonal analysis results of the VJ gene rearrangement, an optional embodiment is provided, along with a specific implementation example: After obtaining the raw paired-end sequencing data, a test sample was sequentially subjected to the following steps: removing reads containing unknown nucleotides (N), removing reads with an average base quality of less than 20, performing majority voting correction on reads with a similarity of >2, and then assembling the sample after removing the adapter sequence to obtain the assembled sequence.

[0076] First, perform statistical visualization analysis based on the length of the assembled read sequence. (See [link to relevant documentation]). Figure 2 A schematic diagram showing the length distribution of the provided assembled read sequences. Figure 2 The vertical axis, Sequence counts, represents the percentage of assembled reads; the horizontal axis, Sequence length, represents the length of the assembled reads, in bytes (bp). Figure 2 It can be seen that there are multiple quantity peaks in the entire assembly sequence, indicating that there may be multiple clones in the assembly sequence.

[0077] Gene alignment was performed using each assembled read sequence, and examples of the alignment results with the IGKV, IGKJ, J_C_intron, and Kde gene reference databases are shown in Table 1: Table 1. Alignment results of an assembled read sequence

[0078] Based on the target alignment gene obtained from the comparison, namely the Subject_id gene, VJ rearrangement clonal analysis was performed on all assembled read sequences. The number of clone types, the number of sequences supporting the clone, and the proportion of sequences supporting the clone were calculated for the entire assembled sequence, as shown in Table 2. Table 2. Clonal analysis of VJ rearrangements

[0079] The top 10 clones show that the number of sequences in clone 1 and clone 2 is similar, and both account for more than 40% of the assembled sequences. This indicates that the IGK rearrangement of the current test sample is a polyclonal type.

[0080] Based on the same inventive concept, an embodiment of the second aspect of this invention provides a detection device for IGK gene rearrangements. Please refer to [link to relevant documentation]. Figure 3 The detection device includes: The acquisition module 10 is used to obtain paired-end sequencing data of the test sample; the paired-end sequencing data includes the first end sequencing sequence and the second end sequencing sequence. Assembly module 20 is used to assemble based on the first end sequencing sequence and the second end sequencing sequence to obtain the assembled sequence; The alignment module 30 is used to determine the target alignment gene from the target gene reference database based on the assembled sequence; wherein the gene reference database includes the IGKV gene library, IGKJ gene library, Kde gene library and J_C_intron gene library in germ cell lines, and the target alignment gene includes at least one of the target V gene, target J gene, target Kde gene and target J_C_intron gene; Module 40 is used to determine the IGK gene rearrangement result in the assembled sequence based on the target alignment gene.

[0081] Optionally, the first end sequencing sequence includes multiple first read sequences, and the second end sequencing sequence includes multiple second read sequences; Assembly module 20 is used for: Traverse the first reading sequence to determine the first similar reading sequence corresponding to the first reading sequence; perform majority voting based on each pair of the first reading sequence and the first similar reading sequence to obtain the first end-corrected sequence; and traverse the second reading sequence to determine the second similar reading sequence corresponding to the second reading sequence; perform majority voting based on each pair of the second reading sequence and the second similar reading sequence to obtain the second end-corrected sequence. The assembled sequence is obtained by assembling based on the first-end correction sequence and the second-end correction sequence.

[0082] Optionally, the assembly module is used for: Based on each set of first read length sequences and first similar read length sequences, the number of similarities is determined; when the number of similarities is greater than a set value, a majority vote is performed on each base in the first read length sequence and the first similar read length sequence to obtain the first corrected read length sequence; based on all the first corrected read length sequences, the first end-corrected sequence is obtained. Based on each group of second read sequences and second similar read sequences, the number of similarities is determined; when the number of similarities is greater than a set value, a majority vote is performed on each base in the second read sequence and the second similar read sequence to obtain the second corrected read sequence; based on all the second corrected read sequences, the second end corrected sequence is obtained.

[0083] Optionally, assembly module 20 is used for: The connector sequence in the first corrected read sequence is removed to obtain the first preprocessed read sequence, and the first end preprocessed sequence is obtained based on all the first preprocessed read sequences; and the connector sequence in the second corrected read sequence is removed to obtain the second preprocessed read sequence, and the second end preprocessed sequence is obtained based on all the second preprocessed read sequences. The assembled sequence is obtained by assembling the first and second preprocessed sequences.

[0084] Optionally, assembly module 20 is used for: Delete a first preprocessed read sequence whose length is less than a first set length to obtain a first-end sequence to be assembled; and delete a second preprocessed read sequence whose length is less than the first set length to obtain a second-end sequence to be assembled. The assembled sequence is obtained by assembling the first-end preprocessed sequence and the second-end preprocessed sequence, including: The assembled sequence is obtained by assembling the first end sequence and the second end sequence.

[0085] Furthermore, assembly module 20 is used for: Obtain the inverse complementary read sequence of the second preprocessed read sequence; The overlapping sequence is determined based on the first preprocessed read length sequence and the reverse complementary read length sequence; When the length of the overlapping sequence is not less than the second set length, the overlapping sequence in the reverse complementary read length sequence is deleted to obtain the read length sequence to be assembled. The first preprocessed read sequence is concatenated with the read sequence to be assembled to obtain the assembled read sequence; Based on all the assembled read length sequences, the assembled sequence is obtained.

[0086] Optionally, the comparison module 30 is used for: Based on the set alignment parameters, the target alignment gene corresponding to each assembled read sequence is determined from the target gene reference database; The alignment parameters include: the similarity between the alignment fragment in the assembled read sequence and the target alignment gene is not less than 90%, and the length of the alignment fragment ranges from 4 to 11.

[0087] Optionally, when the target alignment genes only include target V and target J genes, the determination module 40 is used for: Obtain the nucleotide positions of phenylalanine residues in the target J gene, and determine the termination point in the assembled sequence based on the nucleotide positions; The cysteine ​​residues in the assembled sequence are detected within a set range before the termination point, and the position of the cysteine ​​residue closest to the termination point is taken as the starting point; the set range is the assembled sequence fragment from the termination point to 60 bp to 90 bp before the termination point. Based on the start and end points, determine the CDR3 region in the assembly sequence.

[0088] Optionally, when the target alignment genes only include target V and target J genes, the determination module 40 is used for: Cluster analysis was performed on the assembled sequences based on the target V gene and the target J gene to obtain the number of clone sequences and the proportion of clone sequences in the assembled sequences.

[0089] Based on the same inventive concept, an embodiment of the third aspect of this invention provides an electronic device 400. Please refer to... Figure 4 It includes a processor 420 and a memory 410, the memory 410 being coupled to the processor 420, the memory 410 storing a computer program 411, which, when executed by the processor 420, causes the electronic device 400 to perform the steps of the control method described in the foregoing embodiments.

[0090] Specifically, electronic devices contain operating systems and third-party applications. These electronic devices can be servers, desktop computers, tablets, laptops, mobile phones, wearable devices, in-vehicle terminals, and other similar devices.

[0091] Based on the same inventive concept, please refer to the optional embodiments of the present invention. Figure 5 A computer-readable storage medium 500 is provided, on which a computer program 511 is stored, which, when executed by a processor, performs the steps of the control method described in the foregoing embodiments.

[0092] For the sake of brevity, any aspects not mentioned in the embodiments of the apparatus, electronic devices, and computer-readable storage media may be referred to the corresponding content in the foregoing embodiments of the detection method.

[0093] In summary, this invention provides a method, apparatus, electronic device, and storage medium for detecting IGK gene rearrangements. The method involves assembling an assembled sequence based on the first and second end sequencing sequences from paired-end sequencing raw data. This assembled sequence is then compared with gene reference sequences in germline IGKV, IGKJ, Kde, and J_C_intron gene libraries to identify target alignment genes, including at least one of the target V, J, Kde, and J_C_intron genes. Based on these target alignment genes, the IGK gene rearrangement result in the assembled sequence is determined. This method provides an automated workflow for detecting VJ, V-Kde, and J_C_intron-Kde gene rearrangements in the IGK gene, suitable for downstream analysis and identification needs such as monitoring minimal residual disease and recurrence in lymphoma, and immune repercussions sequencing.

[0094] It should be noted that the term "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Furthermore, the character " / " in this document generally indicates that the preceding and following related objects are in an "or" relationship; the word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of multiple such elements. This invention can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims listing several means, several of these means can be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.

[0095] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0096] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0097] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0098] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0099] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0100] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A method for detecting IGK gene rearrangements, characterized in that, The detection method includes: Obtain paired-end sequencing data of the test sample; the paired-end sequencing data includes the first-end sequencing sequence and the second-end sequencing sequence; The assembled sequence is obtained by assembling based on the first end sequencing sequence and the second end sequencing sequence; Based on the assembled sequence, a target alignment gene is determined from a target gene reference database; wherein, the gene reference database includes the IGKV gene library, IGKJ gene library, Kde gene library and J_C_intron gene library in germ cell lines, and the target alignment gene includes at least one of the target V gene, target J gene, target Kde gene and target J_C_intron gene; Based on the target alignment gene, the rearrangement result of the IGK gene in the assembled sequence is determined; When the target alignment genes only include the target V gene and the target J gene, determining the IGK gene rearrangement result in the assembled sequence based on the target alignment genes includes: The nucleotide positions of phenylalanine residues in the target J gene are obtained, and the termination point in the assembled sequence is determined based on the nucleotide positions; Cysteine ​​residues in the assembled sequence are detected within a set range before the termination point, and the position of the cysteine ​​residue closest to the termination point is taken as the starting point; the set range is the assembled sequence fragment from the termination point to 60 bp to 90 bp before the termination point, so as to ensure that the cysteine ​​residue found is the last cysteine ​​residue before the phenylalanine residue. Based on the starting point and the ending point, the CDR3 region in the assembly sequence is determined.

2. The detection method as described in claim 1, characterized in that, The first end sequencing sequence includes multiple first read sequences, and the second end sequencing sequence includes multiple second read sequences; The assembly based on the first end sequencing sequence and the second end sequencing sequence to obtain the assembled sequence includes: Traverse the first read length sequence to determine the first similar read length sequence corresponding to the first read length sequence; perform majority voting based on each group of the first read length sequence and the first similar read length sequence to obtain the first end-corrected sequence; and traverse the second read length sequence to determine the second similar read length sequence corresponding to the second read length sequence; perform majority voting based on each group of the second read length sequence and the second similar read length sequence to obtain the second end-corrected sequence. The assembled sequence is obtained by assembling based on the first end correction sequence and the second end correction sequence.

3. The detection method as described in claim 2, characterized in that, The step of obtaining a first end-corrected sequence by majority voting based on each group of the first read length sequence and the first similar read length sequence includes: Based on each group of the first read length sequence and the first similar read length sequence, the number of similarities is determined; when the number of similarities is greater than a set value, a majority vote is performed on each base in the first read length sequence and the first similar read length sequence to obtain the first corrected read length sequence; based on all the first corrected read length sequences, the first end-corrected sequence is obtained. The step of obtaining the second end-corrected sequence by majority voting based on each group of the second read length sequence and the second similar read length sequence includes: Based on each group of the second read sequence and the second similar read sequence, the number of similarities is determined; when the number of similarities is greater than a set value, a majority vote is performed on each base in the second read sequence and the second similar read sequence to obtain the second corrected read sequence; based on all the second corrected read sequences, the second end-corrected sequence is obtained.

4. The detection method as described in claim 3, characterized in that, After obtaining the first end-corrected sequence and the second end-corrected sequence, the detection method further includes: The connector sequence in the first corrected read sequence is removed to obtain a first preprocessed read sequence, and a first end preprocessed sequence is obtained based on all the first preprocessed read sequences; and the connector sequence in the second corrected read sequence is removed to obtain a second preprocessed read sequence, and a second end preprocessed sequence is obtained based on all the second preprocessed read sequences. The assembly based on the first end-corrected sequence and the second end-corrected sequence to obtain the assembled sequence includes: The assembled sequence is obtained by assembling based on the first end preprocessing sequence and the second end preprocessing sequence.

5. The detection method as described in claim 4, characterized in that, After obtaining the first preprocessed sequence and the second preprocessed sequence, the detection method further includes: Delete a first preprocessed read sequence whose length is less than a first set length to obtain a first-end sequence to be assembled; and delete a second preprocessed read sequence whose length is less than the first set length to obtain a second-end sequence to be assembled. The assembly based on the first end preprocessing sequence and the second end preprocessing sequence to obtain the assembled sequence includes: The assembly sequence is obtained by assembling based on the first end sequence to be assembled and the second end sequence to be assembled.

6. The detection method as described in claim 5, characterized in that, The first set length ranges from 10bp to 100bp.

7. The detection method as described in claim 5, characterized in that, The assembly process based on the first-end sequence to be assembled and the second-end sequence to be assembled, to obtain the assembled sequence, includes: Obtain the inverse complementary read sequence of the second preprocessed read sequence; Based on the first preprocessed read length sequence and the reverse complementary read length sequence, an overlapping sequence is determined; When the length of the overlapping sequence is not less than the second preset length, the overlapping sequence in the reverse complementary read length sequence is deleted to obtain the read length sequence to be assembled. The first preprocessed read length sequence is concatenated with the read length sequence to be assembled to obtain the assembled read length sequence; The assembled sequence is obtained based on all the assembled read sequences.

8. The detection method as described in claim 7, characterized in that, The step of determining the target alignment gene from the target gene reference database based on the assembled sequence includes: Based on the set alignment parameters, the target alignment gene corresponding to each assembled read sequence is determined from the target gene reference database; The set alignment parameters include: the similarity between the alignment fragment in the assembled read sequence and the target alignment gene is not less than 90%, and the length of the alignment fragment ranges from 4 to 11.

9. The detection method as described in claim 1, characterized in that, When the target alignment genes only include the target V gene and the target J gene, determining the IGK gene rearrangement result in the assembled sequence based on the target alignment genes includes: Cluster analysis was performed on the assembled sequence based on the target V gene and the target J gene to obtain the number of clone sequences and the proportion of clone sequences in the assembled sequence.

10. A device for detecting IGK gene rearrangements, characterized in that, The detection device includes: An acquisition module is used to obtain paired-end sequencing data of the test sample; the paired-end sequencing data includes a first-end sequencing sequence and a second-end sequencing sequence. An assembly module is used to assemble based on the first end sequencing sequence and the second end sequencing sequence to obtain an assembled sequence; The alignment module is used to determine the target alignment gene from the target gene reference database based on the assembled sequence; wherein the gene reference database includes the IGKV gene library, IGKJ gene library, Kde gene library and J_C_intron gene library in germ cell lines, and the target alignment gene includes at least one of the target V gene, target J gene, target Kde gene and target J_C_intron gene; The determination module is used to determine the IGK gene rearrangement result in the assembled sequence based on the target alignment gene; Wherein, when the target alignment genes only include the target V gene and the target J gene, determining the IGK gene rearrangement result in the assembled sequence based on the target alignment genes includes: The nucleotide positions of phenylalanine residues in the target J gene are obtained, and the termination point in the assembled sequence is determined based on the nucleotide positions; Cysteine ​​residues in the assembled sequence are detected within a set range before the termination point, and the position of the cysteine ​​residue closest to the termination point is taken as the starting point; the set range is the assembled sequence fragment from the termination point to 60 bp to 90 bp before the termination point, so as to ensure that the cysteine ​​residue found is the last cysteine ​​residue before the phenylalanine residue. Based on the starting point and the ending point, the CDR3 region in the assembly sequence is determined.

11. An electronic device, characterized in that, The device includes a processor and a memory, the memory being coupled to the processor, the memory storing instructions that, when executed by the processor, cause the electronic device to perform the steps of the detection method according to any one of claims 1-9.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the detection method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Method and device for determining pre-rearrangement V / J gene sequences

    CN107038349A