Method for assembling IGHV short-read-length double-end sequencing data

By calculating the maximum common sequence through clustering algorithm and dynamic programming algorithm, splitting and assembling IGHV short-read double-end sequencing data, the problems of insufficient sensitivity, high cost and long time in IGHV mutation detection in existing technologies are solved, and the fast, accurate and low-cost assembly of IGHV full-length sequences is achieved.

CN120673850APending Publication Date: 2025-09-19SHENZHEN NEOIMMUNE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510635389.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing technologies for IGHV mutation detection have problems such as insufficient detection sensitivity, high cost, long time and complex experiments, especially when using short-read sequencing data, it is difficult to assemble the complete IGHV full-length sequence.

Method used

By obtaining IGHV short-read paired-end sequencing data, clustering algorithm and dynamic programming algorithm are used to calculate the maximum common sequence, split and assemble the target cell sequencing data, and realize the assembly of IGHV full-length sequence.

Benefits of technology

It achieves rapid, accurate, low-cost and convenient IGHV mutation detection, adapts to the immune repertoire data generated by short-read sequencing, and improves detection speed and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673850A_ABST
    Figure CN120673850A_ABST
Patent Text Reader

Abstract

The invention discloses a method for assembling IGHV short-read-length double-end sequencing data, and the method comprises the following steps: obtaining IGHV short-read-length double-end sequencing data, and analyzing through an immune repertoire to obtain target cell sequencing data; splitting the target cell sequencing data into double-end read length data based on a plurality of different forward primers; dividing the double-end read length data of the first forward primer into forward read length data and reverse read length data, and calculating a maximum common sequence of the forward read length data and the double-end read length data of the second forward primer; assembling the forward read length data and the double-end read length data of the second forward primer according to the maximum common sequence; and repeating to obtain a full-length sequence of the IGHV. The method for acquiring the full length of the IGHV based on dynamic planning assembly can adapt to immune repertoire data generated by short read length sequencing, so that rapid, accurate, low-cost and convenient IGHV mutation detection is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of molecular biology technology, and in particular to a method for assembling IGHV short-read paired-end sequencing data. Background Art

[0002] Somatic hypermutation is a high-frequency mutation phenomenon that occurs in the human immunoglobulin heavy chain variable region (IGHV) gene. It is an adaptive mechanism of the immune system that aims to increase the diversity and affinity of antibodies. Studies have found that IGHV mutation status is an important prognostic marker for many diseases such as chronic lymphocytic leukemia and small lymphocytic lymphoma. Clinically, mutation status is divided into two types: IGHV mutation (mutation rate ≥ 2%) and IGHV no mutation (mutation rate < 2%). For patients with IGHV no mutation status, it is speculated that their tumor cells originate from primitive B cells in the pregerminal center. Patients progress rapidly and are more likely to respond poorly to immunochemotherapy, corresponding to a poor prognosis. Therefore, IGHV mutation detection is of great significance in fields such as immunology, oncology, and personalized medicine. In clinical application, there is an urgent need for accurate, stable, rapid, and highly standardized IGHV mutation detection methods to achieve the calculation of IGHV mutation rate and the determination of IGHV mutation status.

[0003] Currently, there are many technical means to solve the problem of IGHV mutation detection, such as Sanger sequencing: forward primers are designed in the Leader and FR1 regions upstream of the V region, and reverse primers are designed in the J region or C region. After amplification, the amplified products are separated and purified for Sanger sequencing. Since its read length is usually between 500 and 1000 bases, the complete IGHV gene fragment can be directly obtained. However, the first-generation sequencing has low throughput, insufficient detection sensitivity, limited clonality resolution, and lacks the ability to analyze clonal evolution and accurately detect mutations. There is a high risk of experimental failure in polyclonal data, and the experimental process is complicated, which requires more time and economic costs when analyzing large samples or polyclonal samples.

[0004] In addition to Sanger sequencing, the more commonly used method for IGHV mutation detection is to directly obtain or assemble the full-length sequence of IGHV through NGS technology, specifically using PE250 double-end sequencing to obtain a 250-base sequence in both the forward and reverse directions, and splicing the amplified data into a complete full-length. Although this method overcomes many problems of first-generation sequencing, it still relies on long-read sequencing, which results in higher costs and longer sequencing times. However, if a shorter short-read PE150 sequencing is used, the amplification product of any single region primer cannot cover the complete V region, and the multi-component data obtained by multiple PCR primers needs to be assembled, and the existing methods can usually only distinguish and assemble Leader-J and FR1-J data. Therefore, it is necessary to provide a method to realize the assembly of IGHV short-read double-end sequencing data. Summary of the Invention

[0005] The present application aims to solve at least one of the technical problems existing in the prior art. To this end, the present application proposes a method for assembling IGHV short-read paired-end sequencing data.

[0006] In a first aspect of the present application, a method for assembling IGHV short-read paired-end sequencing data is provided, the method comprising:

[0007] S100: Obtain IGHV short-read paired-end sequencing data and obtain target cell sequencing data through immune repertoire analysis;

[0008] S200: Split target cell sequencing data into paired-end read data based on multiple different forward primers;

[0009] S300: dividing the paired-end read data of the first forward primer into forward read data and reverse read data, and calculating the maximum common sequence between the forward read data and the paired-end read data of the second forward primer, where the first forward primer is located upstream of the second forward primer;

[0010] S400: Assemble the forward read data and the paired-end read data of the second forward primer according to the maximum common sequence;

[0011] S500: Repeat S300 to S400 to obtain the full-length sequence of IGHV.

[0012] The method according to the embodiment of the present application has at least the following beneficial effects:

[0013] A specific assembly method for obtaining the full-length IGHV has been developed for the short-read immune repertoire data obtained based on the design of multiplex PCR primer sets. According to the maximum common sequence obtained, several groups of read length data are assembled and continuously extended to finally obtain the full-length IGHV sequence. This method can adapt to the immune repertoire data generated by short-read sequencing, thereby realizing fast, accurate, low-cost and convenient IGHV mutation detection.

[0014] In some embodiments of the present application, the calculation in S300 is based on at least one of a clustering algorithm and a dynamic programming algorithm. Clustering or dynamic programming is used to efficiently and accurately derive the maximum common sequence and assemble it, thereby improving the detection speed and accuracy of IGHV mutation detection.

[0015] In some embodiments of the present application, S400 includes assembling the upstream sequence excluding the maximum common sequence in the forward read data, the maximum common sequence, and the downstream sequence excluding the maximum common sequence in the double-end read data of the second forward primer.

[0016] In some embodiments of the present application, the forward read data and / or the double-end read data of the second forward primer used to calculate the maximum common sequence in S300 are read sequences screened based on at least one of sequence frequency, common sequence length, and the proportion of common sequence length in the total sequence length.

[0017] In some embodiments of the present application, the target cell sequencing data in S100 is sequencing data containing high-frequency clones of abnormal heavy chains.

[0018] In some embodiments of the present application, the heavy chain abnormal high-frequency clone is a clone that meets at least one of the set conditions of cloning frequency, distribution continuity, and nucleated cell proportion.

[0019] In some embodiments of the present application, the cloning frequency is set to be greater than 3%.

[0020] In some embodiments of the present application, the set distribution continuity includes discontinuous distribution.

[0021] In some embodiments of the present application, the set nucleated cell ratio is no less than 0.2%.

[0022] In some embodiments of the present application, the multiple different forward primers in S200 include a leader peptide forward primer and a framework region forward primer.

[0023] In some embodiments of the present application, the framework region forward primer includes a forward primer for at least one framework region among the FR1 region, the FR2 region, and the FR3 region.

[0024] In some embodiments of the present application, each framework region includes at least one forward primer.

[0025] In some embodiments of the present application, the short-read double-end sequencing is any one of PE150, PE100, and PE75.

[0026] In a second aspect of the present application, a method for detecting IGHV mutations is provided, comprising the step of assembling the full-length sequence of IGHV according to the aforementioned method.

[0027] In some embodiments of the present application, the step of aligning the full-length sequence with a reference sequence is also included.

[0028] According to a third aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the above method.

[0029] In a fourth aspect of the present application, a device is provided, which includes a processor and a memory, wherein the memory stores a computer program that can be run on the processor, and the processor implements the above method when running the computer program.

[0030] In a fifth aspect of the present application, a device is provided, comprising:

[0031] The acquisition module is used to obtain IGHV short-read paired-end read data and obtain target cell sequencing data through immune repertoire analysis;

[0032] A splitting module is used to split the target cell sequencing data into paired-end read data based on multiple different forward primers;

[0033] An assembly module is used to divide the double-end read data of the first forward primer into forward read data and reverse read data, calculate the maximum common sequence of the forward read data and the double-end read data of the second forward primer; and assemble the forward read data and the double-end read data of the second forward primer according to the maximum common sequence.

[0034] In some embodiments of the present application, calculation of the assembly module is based on at least one of a clustering algorithm and a dynamic programming algorithm.

[0035] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become obvious from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 This is a flowchart of a method for assembling IGHV short-read paired-end sequencing data in one embodiment of the present application.

[0037] Figure 2Schematic diagram of the structure, length, and multiplex PCR amplification products of the antibody heavy chain (IGH) in one embodiment of the present application.

[0038] Figure 3 This is a process for detecting IGHV mutations based on PE150 in one embodiment of the present application.

[0039] Figure 4 This is the comparison result of the full-length sequence of IGHV in one embodiment of the present application with the germline genes in the IMGT database.

[0040] Figure 5 This is the alignment result of the full-length assembly based on PE150 data in one embodiment of the present application.

[0041] Figure 6 This is a comparison of the IGHV mutation detection results of different clinical samples using the method of Example 1 and the Sanger sequencing method in this application. DETAILED DESCRIPTION

[0042] The following will clearly and completely describe the concept and technical effects of this application in conjunction with the embodiments to fully understand the purpose, features and effects of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments of this application, other embodiments obtained by those skilled in the art without creative work are all within the scope of protection of this application.

[0043] The embodiments of the present application are described in detail below. The described embodiments are exemplary and are only used to explain the present application, and should not be understood as limiting the present application.

[0044] In the description of this application, "several" means more than one, "plurality" means more than two, "greater than," "less than," and "exceed" are understood to exclude the number itself, while "above," "below," and "within" are understood to include the number itself. The use of "first" and "second" in the description is solely for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance, implicitly specifying the number of the indicated technical features, or implicitly specifying the order of the indicated technical features.

[0045] Unless otherwise defined, all technical and scientific terms used in this application have the same meaning as commonly understood by those skilled in the art to which this application belongs. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application. The specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0046] The first aspect of the present application provides a method for assembling IGHV short-read paired-end sequencing data.

[0047] The term "human immunoglobulin heavy chain variable region (IGHV)" refers to a component of the human immunoglobulin heavy chain (IGH). IGH consists of a variable region and a constant region. The amino acid sequences of three regions in the variable region are highly variable. These regions form a spatial conformation complementary to the antigenic epitope and are also called complementarity determining regions (CDRs), denoted by CDR1, CDR2, and CDR3. The amino acid sequences of regions outside the CDRs in the variable region are relatively unchanged and are called framework regions (FRs). The FR region is divided into four parts by the CDR regions, namely FR1, FR2, FR3, and FR4. Therefore, the basic structure of the IGH encoding gene can be divided from upstream to downstream (5' to 3' end) into a leader peptide (Leader), FR1, CDR1, FR2, CDR2, FR3, CDR3, FR4, and C region. The V region includes Leader to FR3, the J region includes FR4, and the end of the V region, the D region, and the front of the J region constitute CDR3. CDR3 is the most diverse part of the entire variable region and is also the part that directly binds to the antigen during the antigen recognition process. As the core coding gene for complementary antigen recognition, it has high specificity. Therefore, CDR3 is generally used to define a cell clone.

[0048] The term "read length" refers to the number of consecutive bases measured from a single nucleic acid molecule in a single sequencing cycle. In some embodiments, a short read length refers to a read length of less than 300 bp, such as 250 bp, 200 bp, 150 bp, 100 bp, 75 bp, or 50 bp. In some embodiments, a short read length is 50 to 300 bp.

[0049] The term "sequencing" refers to the process of determining the order of bases in a nucleic acid molecule using specific experimental techniques. Single-end sequencing involves sequencing only one end of a nucleic acid molecule, generating a single read; paired-end sequencing involves sequencing both ends of a nucleic acid molecule, generating two reads in opposite directions. In some embodiments, the sequencing method is second-generation sequencing, specifically, any of the following: 454 Roche GS FLX, Illumina Solexa, SOLiD, Ion Torrent, DNBSEQ, etc.

[0050] The term "assembly" refers to the process of piecing together the long reads generated by sequencing to form a longer continuous sequence.

[0051] refer to Figure 1The method for assembling IGHV short-read paired-end sequencing data includes steps S100 to S500.

[0052] S100: Obtain IGHV short-read paired-end sequencing data and obtain target cell sequencing data through immune repertoire analysis.

[0053] The term "immune repertoire" refers to the combination of all lymphocyte receptor encoding genes (clones) in the body. Among them, lymphocytes include B lymphocytes and T lymphocytes. Lymphocyte receptors are proteins expressed on the surface of lymphocyte membranes that can recognize antigens and produce immune responses. According to the type of lymphocytes, they can be divided into T cell receptors (TCR) and B cell receptors (BCR). The immune repertoire specifically includes but is not limited to TRA, TRB, TRD, TRG, IGH, IGK, IGL, etc.

[0054] Because CDR3, as the core gene encoding antigen complementary recognition, is highly specific, in some embodiments, target cells are recognized through CDR3, and sequencing data containing specific CDR3 sequences constitutes target cell sequencing data. Therefore, immune repertoire analysis includes analyzing and determining CDR3 sequences in sequencing data. In some embodiments, the CDR3 sequence is determined by sequence structure analysis. In some embodiments, sequence structure analysis includes determining V(D)J sequences. In some embodiments, determining V(D)J sequences is achieved through alignment. In some embodiments, determining V(D)J sequences includes alignment and re-alignment. In some embodiments, immune repertoire analysis also includes data filtering and quality control. In some embodiments, immune repertoire analysis includes data filtering and quality control, determining V(D)J sequences through alignment and re-alignment, and determining CDR3 and amino acid sequences through sequence structure analysis. In some embodiments, immune repertoire analysis can be performed using at least one tool such as Imonitor, scRepertoire, MiXCR, Immcantation, TRUST4, IMGT / HighV-QUEST, Partis, or IgBLAST.

[0055] In some embodiments, the target cells are cells harboring heavy chain aberrant clones. In some embodiments, the target cell sequencing data includes sequencing data containing heavy chain aberrant clones. In some embodiments, heavy chain aberrant clones are identified based on the frequency of their CDR3 sequences. In some embodiments, heavy chain aberrant clones are clones with a clonal frequency greater than 3%, such as 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, or 80%. In some embodiments, clonal frequency is calculated as: clonal frequency = number of clones harboring a specific CDR3 sequence / total number of clones × 100%. In some embodiments, aberrant clones are identified based on not only clonal frequency but also distribution continuity and nucleated cell percentage. Therefore, in some embodiments, a heavy chain abnormally high-frequency clone is a clone that meets at least one of the following conditions: a set cloning frequency, a set distribution continuity, and a set nucleated cell percentage.

[0056] The distribution continuity requirement means that the clone frequency (freq) distribution satisfies the distribution discontinuity. The discontinuity judgment criterion is that the number of clones that appear within a range of 10 times the corresponding frequency (Freq) of the clone should not exceed 5. The calculation formula is:

[0057]

[0058] The calculation process of the nucleated cell ratio (TNC) is as follows:

[0059] Set the total number of cloned sequences to Na;

[0060] From the BCR / TCR gene, we can separate Nr cell line sequences with known sequences and an input of S. Based on the known sequences and their corresponding receptor chain types, we can know the amplification factor of the gene in the receptor chain. The calculation formula is:

[0061] Separate the ALB plasmid sequence Pr from the system. The number of ALB plasmids input is N. Calculate the amplification multiple of the ALB plasmid in the system. The calculation formula is:

[0062] ALB cell line sequences (Cr) were separated from the system. Since the system used the same multiplex PCR method for immune repertoire sequencing analysis, the ALB plasmid and the ALB cell line had the same amplification times. Based on the amplification times, the original number of ALB cells, i.e., the total number of nucleated cells (Cell_num), can be calculated using the following formula: The amplified product of the ALB cell line is diploid, so the result needs to be divided by 2;

[0063] Since the BCR / TCR clones in the test sample and the artificially synthesized BCR / TCR clones were subjected to immune repertoire sequencing analysis using the same multiplex PCR method, the amplification times of the two were the same. Therefore, the nucleated cell percentage (TNC) corresponding to the clone can be calculated based on the amplification times and the total number of nucleated cells. The calculation formula is:

[0064]

[0065] In some embodiments, the set distribution continuity is a discontinuous distribution.

[0066] In some embodiments, the percentage of nucleated cells is set to be no less than 0.2%.

[0067] S200: Split target cell sequencing data into paired-end read data based on multiple different forward primers.

[0068] The term "primer" refers to an oligonucleotide that initiates template-dependent nucleic acid synthesis. Typically, in the presence of a nucleic acid template, dNTPs, a polymerase, and optional cofactors, and under appropriate temperature and pH conditions, the primer can be extended to its 3' end by the addition of nucleotides by the polymerase to form an extension product. The length of the primer can vary depending on the reaction conditions and the intended purpose. In some embodiments, the primer is 8 to 25 nucleotides in length, for example, 8, 10, 15, 20, or 25 nucleotides.

[0069] A forward primer is a primer that binds to the 3' end of the template strand and generates a new strand complementary to the template strand; a reverse primer is a primer that binds to the 3' end of the complementary strand of the template strand and generates a new strand complementary to the complementary strand. In some embodiments, Read 1 of a double-end sequencing run is typically generated by the forward primer, and Read 2 is typically generated by the reverse primer.

[0070] Forward primers and reverse primers are designed using the more conservative regions on IGHV that are less prone to mutation. Therefore, in some embodiments, a plurality of different forward primers include a leader peptide forward primer and a framework region forward primer. In some embodiments, the framework region forward primer includes a forward primer for at least one framework region in the FR1 region, the FR2 region, and the FR3 region. In some embodiments, each framework region design includes at least one forward primer, specifically 1 to 5 forward primers, for example, 1, 2, 3, 4, or 5 forward primers. In some embodiments, the framework region forward primer includes a FR1 forward primer, a FR2 forward primer, and a FR3 forward primer. Similarly, in some embodiments, the reverse primer is a FR4 reverse primer designed based on FR4 of the J region. Thus, the sequencing data includes Leader-J extension product data, FR1-J extension product data, FR2-J extension product data, and FR3-J extension product data. Reference Figure 2 The Leader-J product is approximately 450 bp long, the FR1-J product is approximately 400 bp long, the FR2-J product is approximately 270 bp long, and the FR3-J product is approximately 150 bp long. Therefore, in PE150 sequencing mode, Leader-J and FR1-J cannot be sequenced, while FR2-J can be sequenced in most cases and FR3-J can be sequenced in all cases.

[0071] The length and position of the amplified products (i.e., the sequenced target sequences) ultimately obtained by different forward primers vary. Therefore, in some embodiments, the target cell sequencing data can be split into paired-end read data based on multiple different forward primers by splitting the target CDR3 data into primer data for each region based on indicators such as the target sequence length, the target sequence alignment start position, the reference sequence alignment start position, and the distance from the target sequence alignment start position to the reference sequence alignment start position (e.g., q.Start / End in BLAST software: the start / end position of the alignment region in the query sequence, s.Start / End: the start / end position of the alignment region on the target sequence). In some embodiments, the reference sequence is derived from the IMGT database (International ImMunoGeneTics Information System), which is an annotated sequence of the immune repertoire in the human genome.

[0072] In some embodiments, after splitting the target cell sequencing data into paired-end read data based on multiple different forward primers, the method further includes sorting the paired-end read data of each forward primer according to quantity.

[0073] In some embodiments, after splitting the target cell sequencing data into paired-end read data based on a plurality of different forward primers, the method further includes performing quality control filtering on the split paired-end read data.

[0074] S300: Divide the double-ended read data of the first forward primer into forward read data and reverse read data, and calculate the maximum common sequence between the forward read data and the double-ended read data of the second forward primer, where the first forward primer is located upstream of the second forward primer.

[0075] In some embodiments, the calculation is based on at least one of a clustering algorithm and a dynamic programming algorithm.

[0076] The term "clustering algorithm" refers to an unsupervised learning method that divides objects in a dataset into clusters, where objects within a cluster are highly similar and objects between clusters are highly different. Depending on the specific clustering principle, these include partition-based clustering (such as K-means and K-medoids), hierarchical clustering (such as BIRCH and CURE), density-based clustering (such as DBSCAN, OPTICS, and DENCLUE), grid-based clustering (such as STING, CLIQUE, and WaveCluster), and model-based clustering (such as GMM, LDA, and SOM neural networks).

[0077] The term "dynamic programming algorithm" refers to a method that decomposes an original problem into multiple interrelated subproblems and solves the subproblems to construct an optimal solution to the original problem using the solutions of the subproblems, such as the maximum common substring (LCS).

[0078] In some embodiments, the first forward primer and the second forward primer are adjacent forward primers on the template strand. For example, in a combination of a Leader forward primer, a FR1 forward primer, a FR2 forward primer, and a FR3 forward primer, the first forward primer and the second forward primer are respectively a Leader forward primer and a FR1 forward primer, or a FR1 forward primer and a FR2 forward primer, or a FR2 forward primer and a FR3 forward primer.

[0079] In some embodiments, the forward read data used to calculate the maximum common sequence in S300 is a read sequence screened out from all the forward read data based on at least one of sequence frequency, common sequence length, and the proportion of common sequence length in the total sequence length. In some embodiments, the read sequence screened out based on sequence frequency is a read sequence with the highest frequency of occurrence in all the forward read data. In some embodiments, the read sequence screened out based on common sequence length is a read sequence with the longest common sequence length in all the forward read data. In some embodiments, the read sequence screened out based on the proportion of common sequence length in the total sequence length is a read sequence with the largest proportion of common sequence length in the total sequence length in all the forward read data. Wherein, the total sequence length can be, for example, the sequence length of the double-end sequencing data of the first forward primer, or other equivalent sequence length.

[0080] In some embodiments, the double-end read data of the second forward primer used to calculate the maximum common sequence in S300 is a read sequence screened out from the double-end read data of all the second forward primers based on at least one of sequence frequency, common sequence length, and the proportion of common sequence length in the total sequence length. In some embodiments, the read sequence screened out based on sequence frequency is a read sequence with the highest frequency of occurrence in the double-end read data of all the second forward primers. In some embodiments, the read sequence screened out based on the common sequence length is a read sequence with the longest common sequence length in the double-end read data of all the second forward primers. In some embodiments, the read sequence screened out based on the proportion of common sequence length in the total sequence length is a read sequence with the largest proportion of common sequence length in the total sequence length in the double-end read data of all the second forward primers. Wherein, the total sequence length can be, for example, the sequence length of the double-end sequencing data of the first forward primer, or other equivalent sequence length.

[0081] S400: Assemble the forward read data and the double-end read data of the second forward primer according to the maximum common sequence.

[0082] In some embodiments, S400 includes assembling the upstream sequence excluding the maximum common sequence in the forward read data, the maximum common sequence, and the downstream sequence excluding the maximum common sequence in the paired-end read data of the second forward primer.

[0083] S500: Repeat S300 to S400 to obtain the full-length sequence of IGHV.

[0084] In some embodiments, the plurality of different forward primers include a combination of a Leader forward primer, a FR1 forward primer, a FR2 forward primer, and a FR3 forward primer, and S300 to S500 include A1 to A2:

[0085] A1) dividing the paired-end read data of the FR1 forward primer into FR1 forward read data and FR1 reverse read data, calculating the maximum common sequence between the FR1 forward read data and the paired-end read data of the FR2 forward primer based on a dynamic programming algorithm, and assembling the FR1 forward read data and the paired-end read data of the FR2 forward primer based on the maximum common sequence to obtain FR1-J assembled read data;

[0086] A2) The paired-end read data of the Leader forward primer is divided into Leader forward read data and Leader reverse read data. The maximum common sequence between the Leader forward read data and the FR1-J assembly read data is calculated based on the dynamic programming algorithm. The Leader forward read data and the FR1-J assembly read data are assembled according to the maximum common sequence.

[0087] In some embodiments, it also includes: A0) dividing the double-ended read data of the FR2 forward primer into FR2 forward read data and FR2 reverse read data, calculating the maximum common sequence of the FR2 forward read data and the double-ended read data of the FR3 forward primer based on a dynamic programming algorithm, assembling the FR2 forward read data and the double-ended read data of the FR3 forward primer according to the maximum common sequence to obtain FR2-J assembled read data, and replacing the double-ended read data of the FR2 forward primer with the FR2-J assembled read data to perform A1.

[0088] In some embodiments, the assembly further comprises removing the primer sequence.

[0089] In some embodiments, the short-read paired-end sequencing is any of PE150, PE100, and PE75. For PE100 and PE75, due to their shorter read lengths, more sets of forward sequencing primers can be used, and the full-length sequence of IGHV can be assembled through more assembly cycles, for example, using two primers in the FR1 region or other FR regions.

[0090] In the above assembly process, data loss in any one or more primer regions is allowed, which enables sequence assembly with a high degree of freedom.

[0091] In a second aspect of the present application, a method for detecting IGHV mutations is provided, comprising the step of assembling the full-length sequence of IGHV according to the aforementioned method.

[0092] In some embodiments, the step of aligning the full-length sequence with a reference sequence is also included.

[0093] In some embodiments, a reference sequence comprises the sequence of an IGH germline gene in IMGT.

[0094] In some embodiments, the alignment includes determining the position and number of mutations.

[0095] In some embodiments, the method for detecting IGHV mutations is a non-disease diagnostic or therapeutic method.

[0096] According to a third aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the above method.

[0097] In a fourth aspect of the present application, a device is provided, which includes a processor and a memory, wherein the memory stores a computer program that can be run on the processor, and the processor implements the above method when running the computer program.

[0098] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer executable programs, such as the aforementioned method in the embodiments of this application. The processor implements the aforementioned method by executing the non-transitory software programs and instructions stored in the memory.

[0099] The memory may include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function, while the data storage area may store and execute the aforementioned programs. Furthermore, the memory may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device.

[0100] In some embodiments, the memory may optionally include a memory remotely located relative to the processor, and the remote memory may be connected to the processor via a network. Examples of the aforementioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0101] The non-transitory software program and instructions required to implement the above method are stored in the memory, and when executed by one or more processors, the above method is performed.

[0102] In a fifth aspect of the present application, a device is provided, comprising:

[0103] The acquisition module is used to obtain IGHV short-read paired-end read data and obtain target cell sequencing data through immune repertoire analysis;

[0104] A splitting module is used to split the target cell sequencing data into paired-end read data based on multiple different forward primers;

[0105] An assembly module is used to divide the double-end read data of the first forward primer into forward read data and reverse read data, calculate the maximum common sequence of the forward read data and the double-end read data of the second forward primer; and assemble the forward read data and the double-end read data of the second forward primer according to the maximum common sequence.

[0106] In some embodiments, the calculation of the assembly module is based on at least one of a clustering algorithm and a dynamic programming algorithm.

[0107] In some embodiments, target cells are identified by CDR3, and sequencing data containing specific CDR3 sequences is target cell sequencing data. In some embodiments, the CDR3 sequence is determined by sequence structure analysis. In some embodiments, sequence structure analysis includes determining the V(D)J sequence. In some embodiments, determining the V(D)J sequence is achieved by alignment. In some embodiments, determining the V(D)J sequence includes alignment and re-alignment. In some embodiments, immune repertoire analysis also includes data filtering and quality control. In some embodiments, immune repertoire analysis includes data filtering and quality control, alignment and re-alignment to determine the V(D)J sequence, sequence structure analysis to determine the CDR3 and amino acid sequence, etc. In some embodiments, immune repertoire analysis can be achieved by at least one tool such as Imonitor, scRepertoire, MiXCR, Immcantation, TRUST4, IMGT / HighV-QUEST, Partis, IgBLAST, etc.

[0108] In some embodiments, the target cells are cells harboring heavy chain aberrant clones. In some embodiments, the target cell sequencing data includes sequencing data containing heavy chain aberrant clones. In some embodiments, heavy chain aberrant clones are identified based on the frequency of their CDR3 sequences. In some embodiments, heavy chain aberrant clones are clones with a clonal frequency greater than 3%, such as 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, or 80%. In some embodiments, clonal frequency is calculated as: clonal frequency = number of clones harboring a specific CDR3 sequence / total number of clones × 100%. In some embodiments, aberrant clones are identified based on not only clonal frequency but also distribution continuity and nucleated cell percentage. Thus, in some embodiments, a high-frequency heavy chain abnormal clone is a clone that meets at least one of a predetermined clonal frequency, a predetermined distribution continuity, and a predetermined nucleated cell percentage. In some embodiments, the predetermined distribution continuity is a discontinuous distribution. In some embodiments, the predetermined nucleated cell percentage is no less than 0.2%.

[0109] In some embodiments, the plurality of different forward primers include a leader peptide forward primer and a framework region forward primer. In some embodiments, the framework region forward primer includes a forward primer for at least one framework region among the FR1 region, the FR2 region, and the FR3 region. In some embodiments, each framework region design includes at least one forward primer, specifically 1 to 5 forward primers, for example, 1, 2, 3, 4, or 5 forward primers. In some embodiments, the framework region forward primer includes a FR1 forward primer, a FR2 forward primer, and a FR3 forward primer. In some embodiments, the reverse primer is a FR4 reverse primer designed based on FR4 of the J region.

[0110] In some embodiments, the target cell sequencing data can be split into paired-end reads based on multiple different forward primers by splitting the target CDR3 data into primer data for each region based on indicators such as target sequence length, target sequence alignment start position, reference sequence alignment start position, and distance from the target sequence alignment start position to the reference sequence alignment start position. In some embodiments, the reference sequence is derived from the IMGT database (International ImMunoGeneTics Information System), which is an annotated sequence of the immune repertoire in the human genome.

[0111] In some embodiments, after splitting the target cell sequencing data into paired-end read data based on multiple different forward primers, the method further includes sorting the paired-end read data of each forward primer according to quantity.

[0112] In some embodiments, after splitting the target cell sequencing data into paired-end read data based on a plurality of different forward primers, the method further includes performing quality control filtering on the split paired-end read data.

[0113] In some embodiments, the first forward primer and the second forward primer are adjacent forward primers on the template strand. For example, in a combination of a Leader forward primer, a FR1 forward primer, a FR2 forward primer, and a FR3 forward primer, the first forward primer and the second forward primer are respectively a Leader forward primer and a FR1 forward primer, or a FR1 forward primer and a FR2 forward primer, or a FR2 forward primer and a FR3 forward primer.

[0114] In some embodiments, the forward read data used to calculate the maximum common sequence in S300 is a read sequence screened out from all the forward read data based on at least one of sequence frequency, common sequence length, and the proportion of common sequence length in the total sequence length. In some embodiments, the read sequence screened out based on sequence frequency is a read sequence with the highest frequency of occurrence in all the forward read data. In some embodiments, the read sequence screened out based on common sequence length is a read sequence with the longest common sequence length in all the forward read data. In some embodiments, the read sequence screened out based on the proportion of common sequence length in the total sequence length is a read sequence with the largest proportion of common sequence length in the total sequence length in all the forward read data. Wherein, the total sequence length can be, for example, the sequence length of the double-end sequencing data of the first forward primer, or other equivalent sequence length.

[0115] In some embodiments, the double-end read data of the second forward primer used to calculate the maximum common sequence in S300 is a read sequence screened out from the double-end read data of all the second forward primers based on at least one of sequence frequency, common sequence length, and the proportion of common sequence length in the total sequence length. In some embodiments, the read sequence screened out based on sequence frequency is a read sequence with the highest frequency of occurrence in the double-end read data of all the second forward primers. In some embodiments, the read sequence screened out based on the common sequence length is a read sequence with the longest common sequence length in the double-end read data of all the second forward primers. In some embodiments, the read sequence screened out based on the proportion of common sequence length in the total sequence length is a read sequence with the largest proportion of common sequence length in the total sequence length in the double-end read data of all the second forward primers. Wherein, the total sequence length can be, for example, the sequence length of the double-end sequencing data of the first forward primer, or other equivalent sequence length.

[0116] In some embodiments, the assembly module assembles the forward read data and the double-end read data of the second forward primer according to the maximum common sequence by assembling the upstream sequence of the forward read data excluding the maximum common sequence, the maximum common sequence, and the downstream sequence of the double-end read data of the second forward primer excluding the maximum common sequence.

[0117] In some embodiments, the plurality of different forward primers include a combination of a Leader forward primer, a FR1 forward primer, a FR2 forward primer, and a FR3 forward primer, and the assembly module includes A1 to A2:

[0118] A1) dividing the paired-end read data of the FR1 forward primer into FR1 forward read data and FR1 reverse read data, calculating the maximum common sequence between the FR1 forward read data and the paired-end read data of the FR2 forward primer based on a dynamic programming algorithm, and assembling the FR1 forward read data and the paired-end read data of the FR2 forward primer based on the maximum common sequence to obtain FR1-J assembled read data;

[0119] A2) The paired-end read data of the Leader forward primer is divided into Leader forward read data and Leader reverse read data. The maximum common sequence between the Leader forward read data and the FR1-J assembly read data is calculated based on the dynamic programming algorithm. The Leader forward read data and the FR1-J assembly read data are assembled according to the maximum common sequence.

[0120] In some embodiments, the assembly module is further configured to: A0) divide the double-ended read data of the FR2 forward primer into FR2 forward read data and FR2 reverse read data, calculate the maximum common sequence of the FR2 forward read data and the double-ended read data of the FR3 forward primer based on a dynamic programming algorithm, assemble the FR2 forward read data and the double-ended read data of the FR3 forward primer according to the maximum common sequence to obtain FR2-J assembled read data, and replace the double-ended read data of the FR2 forward primer with the FR2-J assembled read data for A1.

[0121] In some embodiments, the assembly module further comprises removing the primer sequence therein after the assembly is completed.

[0122] In the above assembly process, data loss in any one or more primer regions is allowed, which enables sequence assembly with a high degree of freedom.

[0123] In the above assembly process, the device further comprises an alignment module, which is used to align the full-length sequence with the reference sequence.

[0124] In some embodiments, a reference sequence comprises the sequence of an IGH germline gene in IMGT.

[0125] In some embodiments, the alignment includes determining the position and number of mutations.

[0126] The device or system implementation described above is merely illustrative. The modules described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network elements. Some or all of these modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0127] It is understood that all or some of the steps disclosed above can be implemented as software, firmware, hardware and appropriate combinations thereof. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, and the computer-readable medium can include computer storage media (or non-transitory media) and communication media (or temporary media). It is understood that computer storage media include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules or other data). Computer storage media include but are not limited to RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, magnetic tape, disk storage or other magnetic storage device, or any other medium that can be used to store desired information and can be accessed by a computer.

[0128] Additionally, it will be appreciated that communication media typically embodies computer readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0129] The present application is described below with reference to specific embodiments.

[0130] Example 1: IGHV mutation detection in a patient with chronic lymphocytic leukemia

[0131] This example is based on short-read (PE150) immune repertoire data obtained by multiplex PCR primer set design (Leader-J, FR1-J, FR2-J, FR3-J). A method for detecting IGHV mutations based on dynamic programming assembly to obtain the full-length IGHV was developed. This method is adaptable to the immune repertoire data generated by short-read PE150-NGS, achieving rapid, accurate, low-cost, and convenient IGHV mutation detection. The specific process is as follows:

[0132] 1. Construction of immune repertoire

[0133] (1) A real chronic lymphocytic leukemia patient was enrolled in the study. 1 mL of peripheral blood was drawn. DNA was extracted using the HiPure Blood DNA Midi Kit I (Magen, D3112-02) according to the instructions. DNA was then quantified using the Qubit BR. The DNA concentration was 254 ng / μL.

[0134] (2) 500 ng of DNA was used to construct a BCR IGH VDJ library using a two-step PCR method (refer to Front Immunol. 2020 Aug 6:11:1631.), and the sequencing was performed on an Illumina novaseq6000 with a read length of PE150 to obtain the raw data.

[0135] The multiplex PCR primers include forward primers (Leader-Primer, FR1-Primer, FR2-Primer, and FR3-Primer) and reverse primer (J-Primer), located at the 5' end of each IGHV region. Multiplex PCR primers were designed based on common principles, including primer length, Tm value, and GC content, for the conserved regions of the aforementioned IGHV regions. These primers were designed using PrimerPlex software.

[0136] refer to Figure 2 The multiplex PCR amplification products are Leader-J, FR1-J, FR2-J, and FR3-J, respectively. The product length of Leader-J is approximately 450bp, the product length of FR1-J is approximately 400bp, the product length of FR2-J is approximately 270bp, and the product length of FR3-J is approximately 150bp. Under the PE150 sequencing mode, Leader-J and FR1-J cannot be detected, FR2-J can be detected in most cases, and FR3-J can be detected.

[0137] (3) The data from the machine are processed using the immune repertoire analysis software IMonitor, which includes data filtering and quality control, alignment and re-alignment to determine the V(D)J sequence, sequence structure analysis to determine the CDR3 and amino acid sequence, etc. After processing, valid data is obtained and used as valid immune repertoire data for subsequent analysis and processing.

[0138] 2. Determine high-frequency abnormal clones in samples

[0139] The most frequent CDR3 clone sequence of the IGH chain in this sample is:

[0140] TGTGCGAAAGTTAATGATTACTATGGTTCGGGAATGGACGTCTGG (SEQ ID NO. 1);

[0141] The read length of the CDR3 clone sequence is 73910, and the clone frequency is 36.98%, which meets the judgment criteria of abnormally high-frequency clones. Therefore, this CDR3 is the target CDR3.

[0142] 3. Split the target CDR3 data

[0143] Sequence data of CDR3 clone sequences containing heavy chain abnormal high-frequency clones (clone frequency >3%, distribution discontinuity, and nucleated cell proportion not less than 0.2%) were screened out. Then, these sequence data were classified into corresponding Leader-J, FR1-J, FR2-J, and FR3-J data based on indicators such as sequence length, target sequence alignment starting position, reference sequence alignment starting position, and the distance from the target sequence alignment starting position to the reference sequence alignment starting point. They were sorted by quantity and quality control filtered for the split Leader-J, FR1-J, FR2-J, and FR3-J data.

[0144] 4. Selection of assembly anchor points

[0145] The highest-frequency FR1-J data that passed quality control was selected as the assembly anchor point and divided into FR1-R1 (forward read data) and FR1-R2 (reverse read data). The sequences of FR1-R1 and FR1-R2 are as follows. There is a gap region between FR1-R1 and FR1-R2 that cannot be covered by sequencing:

[0146] >FR1-R1

[0147] GGTGCAGCTGGTGGAGTCTGGAGGAGACTTGGTACAGTCTGGGGGGTCCCTGAGA CTCTCCTGTGCAGCGTCTGGATTCACCTTGAGCAGTTGTGATTTGACCTGGGTCCGCCA GGCTCCAGGGAAGGGGCTGGAGTGGGTCTCACTTAT (SEQ ID NO. 2);

[0148] >FR1-R2

[0149] GACAATTCCAAGAACACATTGTATCTGCGAATGAACAGTTTGAGAATCGACGACAC GGCCGTATATTATTGTGCGAAAGTTAATGATTACTATGGTTCGGGAATGGACGTCTGGGG CAAAGGGGCCACGGTCACCGTCTCCTCAGGTAAG (SEQ ID NO. 3).

[0150] 5. Completing the gap in FR1-J

[0151] The longest common substring algorithm based on dynamic programming is used to calculate the maximum overlap area between FR1-R1 and FR2-J data. During the calculation process, the best matching FR2 data is selected based on indicators such as the frequency ranking of FR2-J data, the length of the overlap area between FR1-R1 and FR2-J, and the proportion of the overlap area in the total FR1-J sequence, thereby obtaining the longest common sequence (longest common substring) with the maximum overlap area:

[0152] CCAGGGAAGGGGCTGGAGTGGGTCTCACTTAT(SEQ ID NO.4);

[0153] The FR1-R1 sequence before the longest common sequence, the longest common sequence, and the FR2-J sequence after the longest common sequence were selected to form a new sequence as the complete sequence from the FR1 region to the J region.

[0154] At the same time, if there is still a gap region in the FR2-J sequence, the gap region can be filled using a similar method with the FR2-R1 and FR3-J data.

[0155] >FR2-J sequence

[0156] CATCCAGATGACCCAGTCT CCAGGGAAGGGGCTGGAGTGGGTCTCACTTAT AAGGCGTAGTGGTGGTAGCACATATTATGCTGACTCCGTGAAGGGCCCGCTTCACCATTTCCAGAGACAATTCCAAGAACACATTGTATCTGCGAATGAACAGTTTGAGAATCGACGACACGGCCGTATATTATTGTGCGAAAGTTAATGATTACTATGGTTCGGGAATGGACGTCTGGGGCAAAGGGGCCACGGTCACCGTCTCCTCAGGTAAG (SEQ ID NO. 5);

[0157] >Complete FR1-J sequence after gap filling

[0158] GGTGCAGCTGGTGGAGTCTGGAGGAGACTTGGTACAGTCTGGGGGGTCCCTGAGACTCTCCTGTGCAGCGTCTGGATTCACCTTGAGCAGTTGTGATTTGACCTGGGTCCGCCAGGCT CCAGGGAAGGGGCTGGAGTGGGTCTC ACTTATAAGGCGTAGTGGTGGTAGCACATATTATGCTGACTCCGTGAAGGGCCCGCTTCACCATTTCCAGAGACAATTCCAAGAACACATTGTATCTGCGAATGAACAGTTTGAGAATCGACGACACGGCCGTATATTATTGTGCGAAAGTTAATGATTACTATGGTTCGGGAATGGACGTCTGGGGCAAAGGGGCCACGGTCACCGTCTCCTCAGGTAAG (SEQ ID NO. 6).

[0159] 6.Filling the 5' end gap of the complete FR1-J sequence

[0160] The primer sequence at the 5' end of the complete FR1-J sequence needs to be removed, as it cannot completely cover the entire IGHV region. Therefore, refer to step 5:

[0161] The longest common substring algorithm based on dynamic programming is used to calculate the maximum overlap area of ​​Leader-R1 and FR1-J data. During the calculation process, the best matching Leader-R1 data is selected based on indicators such as the frequency ranking of Leader-R1 data, the length of the overlap area between Leader-R1 and FR1-J, and the proportion of the overlap area in the total Leader-J sequence. This results in the longest common sequence (longest common substring) with the maximum overlap area:

[0162] TCTGGAGGAGACTTGGTACAGTCTGGGGGGTCCCTGAGACTCTCCTGTGCAGCGTC TGGATTCACCTTGAGCAGTTGTGATTTGACCTGGGTCCGCCAGGCTCCAGG (SEQ ID NO. 7);

[0163] The Leader-R1 sequence before the longest common sequence, the longest common sequence, and the FR1-J sequence after the longest common sequence were selected to form a new sequence as the Leader-J complete sequence covering the entire IGHV region.

[0164] >Leader-R1

[0165] CTCTTTGTTTGCAGGTGTCCAATGTGAGGTGCAGCTGGTGGAA TCTGGAGGAGACT TGGTACAGTCT GGGGGGTCCCTGAGACTCTCCTGTGCAGCGTCTGGATTCACCTTGAGC AGTTGTGATTTGACCTGGGTCCGCCAG GCTCCAGG (SEQ ID NO.8);

[0166] >Leader-J complete sequence

[0167] CTCTTTGTTTGCAGGTGTCCAATGTGAGGTGCAGCTGGTGGAA TCTGGAGGAGACTTGGTACAGTCTG GGGGGTCCCTGAGACTCTCCTGTGCAGCGTCTGGATTCACCTTGAGCAGTTGTGATTTGACCTGGGTCCGCCAGGC TCCAGG GAAGGGGCTGGAGTGGGTCTCACTTATAAGGCGTAGTGGTGGTAGCACATATTATGCTGACTCCGTGAAGGGCCGCTTCACCATTTCCAGAGACAATTCCAAGAACACATTGTATCTGCGAA TGAACAGTTTGAGAATCGACGACACGGCCGTAATTATTGTGCGAAAGTTAATGATTACTATGGTTCGGGAATGGACGTCTGGGGCAAAGGGGCCACGGTCACCGTCTCCTCAGGTAAG(SEQ ID NO.9).

[0168] 7. Remove part of the primer region to obtain the full-length IGHV sequence.

[0169] >IGHV full-length sequence

[0170] GAGGTGCAGCTGGTGGAATCTGGAGGAGACTTGGTACAGTCTGGGGGGTCCCTGAGACTCTCCTGTGCAGCGTCTGGATTCACCTTGAGCAGTTGTGATTTGACCTGGGTCCGCCAGGCTCCAGGGAAGGGGCTGGAGTGGGTC TCACTTATAAGGCGTAGTGGTGGTAGCACATATTATGCTGACTCCGTGAAGGGCCGCTTCACCATTTCCAGAGACAATTCCAAGAACACATTGTATCTGCGAATGAACAGTTTGAGAATCGACGACACGGCCGTATATTAT(SEQ IDNO.10).

[0171] 8. Compare the full-length IGHV sequence with the germline genes in the IMGT database. The detailed alignment results are as follows: Figure 4 As shown, the V gene aligned is IGHV3-23*04, the alignment length is 285bp, the number of mutations is 31, the IGHV mutation rate is 10.88%, the mutation rate is ≥2%, the mutation status is "mutation occurred", and the alignment of the assembled sequence with the reference sequence is shown in Figure 5 shown.

[0172] Example 2: Clinical data verification of 49 samples

[0173] In addition, 49 real clinical samples were collected and their IGHV mutation rates and mutation status were detected using Sanger sequencing and the method of Example 1. Figure 6 As shown in Table 1, Figure 6 The horizontal axis is the IGHV mutation rate detected by the method based on Example 1, and the vertical axis is the IGHV mutation rate of the clinical standard Sanger method. The correlation R between the IGHV mutation rate of this method and the clinical results is shown in Figure 1. 2 The value was 0.99578; in Table 1, the IGHV mutation status of the samples was judged according to the 2% mutation rate threshold, which were M (mutation), U (no mutation), and NA (uncertain), and the consistency rate of the mutation status results of the 49 clinical samples was 100%.

[0174] Table 1. Validation of Example 1 and Sanger sequencing results in clinical data

[0175] NGS-M NGS-U NGS-NA Sanger-M 19 0 0 Sanger-U 0 21 0 Sanger-NA 0 0 9

[0176] From the above examples, it can be seen that the method provided in this application realizes IGHV mutation detection in NGS immune repertoire data with shorter read lengths, and the detection results of mutation status are consistent with those of first-generation sequencing, indicating that it has good specificity and sensitivity.

[0177] The present application has been described in detail above with reference to the embodiments. However, the present application is not limited to the above embodiments. Various modifications can be made within the scope of knowledge possessed by a person of ordinary skill in the art without departing from the purpose of the present application. In addition, the embodiments of the present application and the features of the embodiments can be combined with each other unless there is a conflict.

Claims

1. A method for assembling IGHV short-read paired-end sequencing data, characterized in that: The following steps are involved: S100: Obtain the IGHV short-read paired-end sequencing data, and obtain target cell sequencing data through immune repertoire analysis; S200: splitting the target cell sequencing data into paired-end read data based on a plurality of different forward primers; S300: Dividing the paired-end read data of a first forward primer into forward read data and reverse read data, and calculating a maximum common sequence between the forward read data and the paired-end read data of a second forward primer, where the first forward primer is located upstream of the second forward primer; S400: Assembling the forward read data and the double-end read data of the second forward primer according to the maximum common sequence; S5 00: repeat S300 to S400 to obtain the full-length sequence of IGHV; Optionally, the calculation in S300 is based on at least one of a clustering algorithm and a dynamic programming algorithm.

2. The method according to claim 1, characterized in that S400 includes assembling the upstream sequence of the forward read data excluding the maximum common sequence, the maximum common sequence, and the downstream sequence of the double-end read data of the second forward primer excluding the maximum common sequence.

3. The method according to claim 1, characterized in that The forward read data and / or the double-end read data of the second forward primer used to calculate the maximum common sequence in S300 are read sequences screened based on at least one of sequence frequency, common sequence length, and the proportion of common sequence length to total sequence length.

4. The method according to claim 1, wherein The target cell sequencing data in S100 is sequencing data containing heavy chain abnormal high-frequency clones; Optionally, the heavy chain abnormally high-frequency clone is a clone that meets at least one of the following conditions: a set cloning frequency, a set distribution continuity, and a set nucleated cell ratio; Optionally, the cloning frequency is set to be greater than 3%; Optionally, the distribution continuity is set to be discontinuous distribution; Optionally, the nucleated cell ratio is set to be no less than 0.2%.

5. The method according to claim 1, wherein The multiple different forward primers in S200 include a leader peptide forward primer and a framework region forward primer; Optionally, the framework region forward primer includes a forward primer for at least one framework region among FR1 region, FR2 region, and FR3 region; Optionally, each framework region includes at least one forward primer.

6. The method according to claim 1, characterized in that The short-read double-end sequencing is any one of PE150, PE100, and PE75.

7. A method for detecting IGHV mutations, characterized in that comprising the steps of assembling the full-length sequence of IGHV according to the method of any one of claims 1 to 6; Optionally, the method further comprises the step of aligning the full-length sequence with a reference sequence.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the method according to any one of claims 1 to 7.

9. The device, characterized in that The method comprises a processor and a memory, wherein the memory stores a computer program that can be run on the processor, and the processor implements the method according to any one of claims 1 to 7 when running the computer program.

10. The device, characterized in that include: An acquisition module is used to acquire the IGHV short-read paired-end read data and obtain target cell sequencing data through immune repertoire analysis; A splitting module, configured to split the target cell sequencing data into paired-end read data based on a plurality of different forward primers; An assembly module, the assembly module being configured to separate the double-ended read data of the first forward primer into forward read data and reverse read data, calculate a maximum common sequence between the forward read data and the double-ended read data of the second forward primer; and assemble the forward read data with the double-ended read data of the second forward primer according to the maximum common sequence; Optionally, the calculation of the assembly module is based on at least one of a clustering algorithm and a dynamic programming algorithm.