Primer design, disease typing diagnosis method and device, composition, kit and application thereof

By designing primers and using high-throughput sequencing technology, combined with upstream and downstream bait oligonucleotides of TRB, a standard quality plasmid was designed for quantitative analysis. This solved the problems of false positives, false negatives, and PCR amplification bias in TRB gene rearrangement detection, and enabled efficient and accurate lymphoma detection and monitoring of small lesions.

CN121970118APending Publication Date: 2026-05-01BOE TECHNOLOGY GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BOE TECHNOLOGY GROUP CO LTD
Filing Date
2024-08-29
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies for TRB gene rearrangement detection suffer from false positives, false negatives, low detection sensitivity, and inability to detect minute lesions. Furthermore, high-throughput sequencing methods are subject to PCR amplification bias and incomplete coverage of clone types.

Method used

A primer design method is designed to determine conservation scores by obtaining reference data and aligning sites, generate conserved regions, screen and evaluate primer combinations, combine high-throughput sequencing technology, use upstream and downstream decoy oligonucleotides of TRB for amplification, design standard quality plasmids and perform quantitative analysis, and construct sequencing libraries.

Benefits of technology

It improves the detection rate and accuracy of lymphoma, solves the problems of low sensitivity and PCR amplification bias in traditional methods, and realizes accurate detection of TRB gene rearrangement and monitoring of microlesions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121970118A_ABST
    Figure CN121970118A_ABST
Patent Text Reader

Abstract

The invention relates to primer design, a disease typing diagnosis method and device, a composition, a kit and application thereof, and the primer design method comprises the following steps: obtaining a plurality of reference data including a TRBV reference gene sequence and a TRBJ reference gene sequence; aiming at each kind of reference data, performing the following operations: aligning a plurality of reference gene sequences according to sites, determining a conservative score list of each site, obtaining conservative scores of a plurality of conservative intervals according to the conservative score list of each site, selecting K conservative intervals with higher conservative scores, K being a natural number greater than or equal to 1, and selecting K being a natural number greater than or equal to 1; generating K groups of primer combinations aiming at the K conservative intervals; and screening the primers in the K groups of primer combinations, evaluating the screened K groups of primer combinations, and obtaining a final primer combination according to an evaluation result.
Need to check novelty before this filing date? Find Prior Art

Description

Primer design, disease typing and diagnostic methods and devices, compositions, kits and their uses

[0001] This disclosure relates to, but is not limited to, the field of biotechnology, and particularly to primer design, disease typing and diagnostic methods and apparatus, compositions, kits and their uses.

[0002] The T cell receptor (TCR) is a membrane receptor on the surface of T lymphocytes, composed of two polypeptide chains: α(TRA) and β(TRB) or γ(TRG) and δ(TRD). α(TRA) and γ(TRG) are light chains, while β(TRB) and δ(TRD) are heavy chains, with α(TRA) and β(TRB) forming the vast majority of TCRs. The TRB gene consists of a variable region (V), a diversit (D), a joining region (J), and a constant region (C). Specifically, the TRB V / D / J gene clusters each contain multiple V, D, or J gene segments. During lymphocyte development, a gene segment is randomly selected from each of the V, D, or J gene clusters, and under the action of recombinase, it is cleaved and linked together to form a complete gene encoding the TRB heavy chain function—a process known as gene rearrangement. Because the V, D, or J segments that make up the TRB gene are diverse, and varying numbers of bases are randomly inserted or deleted between DJ or V-DJ, TRB proteins exhibit diversity, i.e., polyclonal TRB gene rearrangements. Lymphocytes carry specific TRB rearrangement sequences. During lymphoma development, a particular lymphocyte undergoes malignant proliferation accompanied by the proliferation of its specific TRB rearrangement sequence, i.e., lymphoma TRB gene rearrangement monoclonalism. Both polyclonal and monoclonal TRB gene rearrangements provide important auxiliary means for lymphoma diagnosis.

[0003] Traditional detection methods for TRB gene rearrangements involve capillary electrophoresis combined with fluorescent fragment analysis based on first-generation sequencing platforms. This involves designing specific PCR primers and labeling them with fluorescence at the 5' end, obtaining the target fragment through PCR amplification, and then separating the amplification products by capillary electrophoresis to form a fragment size distribution peak map, thereby determining whether TRB is monoclonal or polyclonal and aiding in the diagnosis of lymphoma. However, this method has the following drawbacks:

[0004] (a) False positive results exist: This analytical method is based on the size of PCR product fragments, which leads to fragments with different sequences but the same length being mixed together to form false positive peaks;

[0005] (b) False negative results exist: the fragment distribution peak diagram is limited to a certain range, causing positive peaks outside the range to be ignored;

[0006] (c) Limited clinical application: This method is mainly used to determine tumors or hyperplasia in lymphatic system diseases. Because it is impossible to sequence the specific sequence of each clone, it cannot be used to monitor small residual lesions, etc.

[0007] (d) Low detection sensitivity: The inability to accurately assess and correct PCR amplification bias leads to low detection sensitivity.

[0008] In recent years, high-throughput sequencing technology has been widely used in the detection of TRB gene rearrangements. This involves specific amplification of TRB fragments using multiplex PCR primers, construction of sequencing libraries using adapter ligation or PCR amplification, and identification of monoclonal or polyclonal TRB gene rearrangements through TRB clone sequence analysis and frequency statistics. However, currently available products suffer from problems such as significant PCR amplification bias and incomplete coverage of clone types, leading to low detection accuracy.

[0009]

[0010] The following is an overview of the subject matter described in detail herein. This overview is not intended to limit the scope of the claims.

[0011] This disclosure provides a primer design method, including:

[0012] Obtain one or more reference data, each of which includes multiple reference gene sequences;

[0013] For each type of reference data, the following operations are performed: align the multiple reference gene sequences by site, determine the conservation score list for each site, obtain the conservation scores of multiple conservation intervals based on the conservation score list for each site, select K conservation intervals with higher conservation scores, where K is a natural number greater than or equal to 1, generate K primer combinations for the K conservation intervals; screen the primers in the K primer combinations and evaluate the screened K primer combinations, and obtain the final primer combination based on the evaluation results.

[0014] This disclosure also provides a primer design apparatus, including a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to execute the steps of the primer design method described in any embodiment of this disclosure based on the instructions stored in the memory.

[0015] This disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the primer design method described in any embodiment of this disclosure.

[0016] This disclosure also provides a program product including instructions that, when executed by a computer, perform a primer design method as described in any embodiment of this disclosure.

[0017] The primer design method and apparatus of this disclosure calculate conservation scores to obtain multiple conservation intervals, and then design, screen and evaluate primers for the conservation intervals to finally obtain a set of primers for efficient amplification. By using this primer set for high-throughput sequencing, the problem of low sensitivity of traditional capillary electrophoresis + fragment analysis is solved, as well as the problems of high bias and low coverage of PCR amplification in previous high-throughput sequencing methods, which effectively improves the detection rate and accuracy of lymphoma.

[0018] This disclosure also provides a standard quality grain design method, including:

[0019] Obtain a first gene cluster and a second gene cluster, wherein the first gene cluster includes multiple first gene sequence fragments, the second gene cluster includes multiple second gene sequence fragments, and both the first gene cluster and the second gene cluster contain multiple functional fragments.

[0020] Multiple first gene sequence fragments and multiple second gene sequence fragments are combined to obtain multiple fragment groups, each fragment group containing one first gene sequence fragment and one second gene sequence fragment;

[0021] A non-human sequence is inserted between the first and second gene sequence fragments in each fragment group to obtain multiple plasmid sequences.

[0022] The ratio of each plasmid sequence was determined to obtain the designed standard plasmid.

[0023] This disclosure also provides a composition comprising, obtained by the methods described herein:

[0024] TRB upstream decoy oligonucleotides, wherein the TRB upstream decoy oligonucleotides are derived from at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 97%, or 100% of any one or more sequences in SEQ ID NO: 1-47; and

[0025] TRB downstream decoy oligonucleotides, wherein the TRB downstream decoy oligonucleotides are derived from at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 97%, or 100% of any one or more sequences in SEQ ID NO:48-60.

[0026] In some exemplary embodiments, the upstream decoy oligonucleotide of the TRB is derived from at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 97%, or 100% of several sequences in SEQ ID NO:1-47; and the downstream decoy oligonucleotide of the TRB is derived from at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 97%, or 100% of several sequences in SEQ ID NO:48-60.

[0027] In some exemplary embodiments, the upstream decoy oligonucleotide of the TRB is derived from at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 97%, or 100% of all sequences in SEQ ID NO:1-47; and the downstream decoy oligonucleotide of the TRB is derived from at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 97%, or 100% of all sequences in SEQ ID NO:48-60.

[0028] In some exemplary embodiments, the upstream decoy oligonucleotide of the TRB is selected from all sequences in SEQ ID NO:1-47; the downstream decoy oligonucleotide of the TRB is selected from all sequences in SEQ ID NO:48-60.

[0029] In some exemplary embodiments, the bait oligonucleotide is one or more of the primer and probe.

[0030] In some exemplary embodiments, the bait oligonucleotides may be primers for amplifying the TRB gene (e.g., a TRB-specific primer set, including primers for the TRB V region and the TRB J region), and may be divided into TRB upstream bait oligonucleotides and TRB downstream bait oligonucleotides. The TRB upstream bait oligonucleotides include specific primer sequences that are complementary to the upstream of the TRB V region, and the TRB downstream bait oligonucleotides include specific primer sequences that are complementary to the downstream of the TRB J region.

[0031] In some exemplary embodiments, the TRB upstream decoy oligonucleotide further comprises a forward adapter primer sequence, and the TRB downstream decoy oligonucleotide further comprises a reverse adapter primer sequence; the forward adapter primer sequence and the reverse adapter primer sequence are complementary to the adapter primer sequence used for sequencing.

[0032] In some exemplary embodiments, the adapter primers used for sequencing are selected from one or more of the Nextera series adapters and the TruSeq series adapters.

[0033] In some exemplary embodiments, the forward adapter primer sequence and the reverse adapter primer sequence are located at both ends of each pair of bait oligonucleotides in this application, for subsequent addition of primers to both ends of the PCR product. In some exemplary embodiments, the forward adapter primer sequence is shown in SEQ ID NO:61, and the reverse adapter primer sequence is shown in SEQ ID NO:62.

[0034] In some exemplary embodiments, the TRB upstream decoy oligonucleotide is selected from one or more sequences shown in SEQ ID NO:63-109; and

[0035] The TRB downstream decoy oligonucleotide is selected from one or more sequences shown in SEQ ID NO:110-122.

[0036] This disclosure also provides the use of the compositions described herein in amplifying the TRB gene and / or detecting TRB gene rearrangements.

[0037] In some exemplary embodiments, the decoy oligonucleotides described herein (e.g., TRB upstream decoy oligonucleotides and TRB downstream decoy oligonucleotides) can be used to amplify the TRB gene to obtain rearranged PCR products, including one or more of VJ rearrangement products and VDJ rearrangement products.

[0038] In some exemplary embodiments, adapter primers can be added to both ends of the decoy oligonucleotides described herein to obtain PCR-amplified decoy oligonucleotides. PCR is then performed using these PCR-amplified decoy oligonucleotides, and the resulting PCR products can be sequenced to obtain the sequence of each rearrangement product. The rearrangement status of the TRB gene can thus be determined more accurately and efficiently.

[0039] This disclosure also provides a kit comprising the compositions described herein.

[0040] In some exemplary embodiments, the kit further comprises:

[0041] 56 standard quality particles, each of which contains a UMI sequence, and the UMI sequence contained in each standard quality particle is different, so that each standard quality particle can be uniquely identified by the UMI sequence;

[0042] The UMI sequence is 16 bp in length. The first 4 bp segment consists of the last 4 bases of the TRB D region sequence, the last 4 bp segment consists of the first 4 bases of the TRB J region sequence, and the middle 8 bp segment is a non-human random sequence.

[0043] Each of the standard quality grains further comprises a TRB V region sequence, a TRB D region sequence, and a TRB J region sequence, wherein the TRB V region sequence, the TRB D region sequence, the UMI sequence, and the TRB J region sequence are sequentially linked end-to-end in each of the standard quality grains.

[0044] In some exemplary embodiments, the first 4 bp segment of the UMI sequence is selected from GGGC, GGGG, or AGGG, and the last 4 bp segment is selected from TACT, GCAC, TGA, AGAT, AGGA, ATTC, GATT, CCAT, CAAA, ACCC, GTTC, CCCA, CGGC, TAGC, ATAG, CCAA, GTTG, CGTC, TACG, TCAT, CAAG, CGGA, AATA, TTCT, AAAC, or CCCC.

[0045] In some exemplary embodiments, the UMI sequence is as shown in SEQ ID NO:123-178.

[0046] In some exemplary embodiments, the 56 standard quality grains are standard quality grains that each contain the following sequences:

[0047] TRBV10-2, TRBD1 and TRBJ2-2; TRBV10-3, TRBD2*01 and TRBJ2-7; TRBV11-1, TRBD*02 and TRBJ2-2; TRBV11-3, TRBD1 and TRBJ2-4; TRBV12-3, TRBD2*01 and TRBJ2-2; TRBV12-5, TRBD*02 and TRBJ1-6; TRBV13, TRBD1 and TRBJ1-3; TRBV14, TRBD2*01 and TRBJ2-6; TRBV15, TRBD*02 and TRBJ1-4; TRBV16, TRBD1 and TRBJ2-1; TRBV18, TRB D2*01 and TRBJ1-6; TRBV19, TRBD*02 and TRBJ2-1; TRBV2, TRBD1 and TRBJ1-6; TRBV2, TRBD2*01 and TRBJ1-3; TRBV20-1, TRBD*02 and TRBJ2-4; TRBV29-1, TRBD1 and TRBJ2-3; TRBV24-1, TRBD2*01 and TRBJ1-4; TRBV25-1, TRBD*02 and TRBJ1-6; TRBV27, TRBD1 and TRBJ2-6; TRBV28, TRBD2*01 and TRBJ1-1; TRBV3-1, TRBD*02 and TRBJ1-5; TRBV3-1, TRBD1, and TRBJ1-4; TRBV30, TRBD2*01, and TRBJ2-3; TRBV4-1, TRBD*02, and TRBJ1-1; TRBV4-3, TRBD1, and TRBJ1-2; TRBV5-1, TRBD2*01, and TRBJ2-5; TRBV5-1, TRBD*02, and TRBJ1-5; TRBV5-4, TRBD1, and TRBJ2-4; TRBV5-5, TRBD2*01, and TRBJ1-1; TRBV5-6, TRBD*02, and TRBJ1-2; TRBV5-8, TRBD1, and TRBJ1-2; TRBV6-1, TRBD2* 01 and TRBJ1-3; TRBV6-2, TRBD*02 and TRBJ2-5; TRBV6-4, TRBD1 and TRBJ1-5; TRBV6-6, TRBD2*01 and TRBJ2-7; TRBV6-6, TRBD*02 and TRBJ1-6; TRBV6-8, TRBD1 and TRBJ2-6; TRBV6-9, TRBD2*01 and TRBJ1-1; TRBV7-8, TRBD*02 and TRBJ1-3; TRBV7-3, TRBD1 and TRBJ1-6; TRBV7-4, TRBD2*01 and TRBJ2-1; TRBV7-6, TRBD*02 and TRBJ1-5;TRBV7-7, TRBD1 and TRBJ1-4; TRBV7-9, TRBD2*01 and TRBJ1-2; TRBV7-9, TRBD*02 and TRBJ1-6; TRBV9, TRBD1 and TRBJ2-7; TRBV7-8, TRBD2*01 and TRBJ1-3; TRBV7-2, TRBD*02 and TRBJ1-6; TRBV7-2, TRBD1 and TRBJ2-1; TR BV11-2, TRBD2*01 and TRBJ2-3; TRBV11-3, TRBD*02 and TRBJ2-2; TRBV5-1, TRBD1 and TRBJ2-3; TRBV19, TRBD2*01 and TRBJ2-5; TRBV30, TRBD*02 and TRBJ2-7; TRBV20, TRBD1 and TRBJ2-6; or TRBV13, TRBD2*01 and TRBJ1-1.

[0048] In some exemplary embodiments, the 56 standard quality particles are mixed in an equimolar ratio to obtain a uniformity standard.

[0049] By mixing one or more of the 56 described standard quality particles in a high proportion and the other standard quality particles in a low proportion, an experimental standard for simulating monoclonal or multiclonal experiments is obtained. In some exemplary embodiments, the high proportion of standard quality particles in the experimental standard can be set as needed, and can be one or more types (two, three, four, five or more). The concentration of the high proportion of standard quality particles can be much higher than that of the low proportion of standard quality particles, for example, 100 times, 1000 times, 10000 times, 100,000 times, 1 million times or more of the concentration of the low proportion of standard quality particles. In some exemplary embodiments, the concentration of the high proportion of standard quality particles in the experimental standard can be slightly higher than that of the low proportion of standard quality particles, for example, 1.1 times, 1.25 times, 1.5 times, 2 times, 4 times, 8 times, 10 times or more of the concentration of the low proportion of standard quality particles.

[0050] This disclosure also provides the use of the kit described herein in evaluating the amplification efficiency of multiple primers used to amplify the TRB gene.

[0051] This disclosure also provides embodiments of the kit described herein for use in detecting TRB gene rearrangements.

[0052] In some exemplary embodiments, in the kit described herein, the molar ratios between the primers shown in SEQ ID NO:1-60 are as follows:

[0053] 2:2:1.5:1.5:1.5:2:2.5:1.5:1:2:1.5:1.5:1:1.5:1.5:1:2:1:2:1.5:1:2:1.5:1:1:1:0.5:1.5:1:0.5:1:1:2:1:1:1:1:1:1.5:0.5:1:2:0.5:0.5:1:1.5:1.5:0.5:0.5:2:1.5:0.5:0.5:1:1:1:1.5:2:1:1.

[0054] This disclosure also provides a disease typing diagnostic method for detecting TRB gene rearrangements, including the following steps:

[0055] 1) Obtain the genomic DNA of the sample to be tested;

[0056] 2) Perform PCR on the genomic DNA obtained in step 1) using the composition described herein to obtain PCR products;

[0057] 3) Sequencing the PCR products obtained in step 2) and analyzing the sequencing results to determine whether the TRB gene rearrangement in the sample is monoclonal or polyclonal.

[0058] In some exemplary embodiments, step 1) further includes incorporating the homogeneity standard described herein into the genomic DNA. In some exemplary embodiments, different standards (e.g., homogeneity standards or experimental standards) may be incorporated in step 1) for different experimental purposes.

[0059] This disclosure also provides a standard plasmid design apparatus, including a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to execute the steps of the standard plasmid design method according to any embodiment of this disclosure based on the instructions stored in the memory.

[0060] This disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the standard quality grain design method described in any embodiment of this disclosure.

[0061] This disclosure also provides a program product including instructions that, when executed by a computer, perform a standard quality grain design method as described in any embodiment of this disclosure.

[0062] The standard plasmid design method and apparatus of this disclosure obtain multiple plasmid sequences by inserting a non-human sequence between the first gene sequence fragment and the second gene sequence fragment in each fragment group. The plasmid sequences can be identified and effectively separated from the mixed sample, thereby enabling the quantification of the mixed sample.

[0063] This disclosure also provides a method for quantitative analysis of standard quality grains, including:

[0064] Obtain paired-end sequencing data corresponding to the standard quality plasmid, wherein the paired-end sequencing data includes Reads1 and Reads2 sequences;

[0065] Determine whether the Reads1 and Reads2 sequences include a UMI sequence identifier;

[0066] When both the Reads1 and Reads2 sequences include UMI sequence identifiers and the included UMI sequence identifiers are the same, it is determined that the plasmid sequences corresponding to the Reads1 and Reads2 sequences have been identified.

[0067] When only one of the Reads1 and Reads2 sequences includes a UMI sequence identifier, the Hamming distance between the Reads sequence that does not include the UMI sequence identifier and the UMI sequence identifier is determined; when the determined Hamming distance is less than or equal to a, the plasmid sequence corresponding to the Reads1 and Reads2 sequences is identified, where a is a natural number less than or equal to 2.

[0068] This disclosure also provides a standard plasmid quantification analysis apparatus, including a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to execute the steps of the standard plasmid quantification analysis method according to any embodiment of this disclosure based on the instructions stored in the memory.

[0069] This disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the standard plasmid quantitative analysis method described in any embodiment of this disclosure.

[0070] This disclosure also provides a program product including instructions that, when executed by a computer, perform a standard plasmid quantitative analysis method as described in any embodiment of this disclosure.

[0071] The standard plasmid quantitative analysis method and apparatus of this disclosure can quantify mixed samples by identifying the plasmid sequence corresponding to the read sequence based on the UMI sequence identifier and Hamming distance. For example, it can identify whether the experimental mixing ratio meets expectations (such as whether the plasmid addition ratio is consistent with the sequencing detection ratio).

[0072] This disclosure also provides a primer characterization method, including:

[0073] Obtain sequencing data corresponding to the primer to be tested, and determine the 5' end position of the primer to be tested based on the sequencing data;

[0074] Multiple sequences are designed based on the 5' end position of the primer to be detected. Each of the multiple sequences starts at the 5' end position of the primer to be detected, and the 3' end cutoff positions of two adjacent sequences differ by a single base.

[0075] Design multiple nucleic acid structures based on multiple designed sequences;

[0076] Identify the nucleic acid structures among the multiple nucleic acid structures that can complementarily pair with the primers to be detected;

[0077] The sequence of the primer to be detected is determined based on the nucleic acid structure that can complementarily pair with the primer to be detected.

[0078] The primer characterization method of this disclosure determines the 5' end position of the primer to be detected based on sequencing data, designs multiple sequences and multiple corresponding nucleic acid structures based on the 5' end position of the primer to be detected, then determines the nucleic acid structure that can complementarily pair with the primer to be detected among the multiple nucleic acid structures, and determines the sequence of the primer to be detected based on the nucleic acid structure that can complementarily pair with the primer to be detected, thus enabling the sequencing of any unknown primer.

[0079] This disclosure also provides a method for constructing a sequencing library, including:

[0080] Extract DNA from the genome to be tested;

[0081] Design specific primers for the genome to be detected, mix the specific primers with the extracted DNA, perform the first round of PCR amplification to obtain the first round of PCR product, and purify the first round of PCR product;

[0082] The purified first-round PCR product was mixed with universal adapter primers and subjected to a second-round PCR amplification to obtain a second-round PCR product. The second-round PCR product was then purified to obtain the constructed sequencing library.

[0083] The sequencing library construction method of this disclosure uses a two-step PCR method to ligate sequencing adapters to the 5' and 3' ends of the target fragment, respectively, to convert the target DNA to be sequenced into a form that can be read by the sequencer.

[0084] This disclosure also provides a method for disease classification and diagnosis, including:

[0085] Obtain sequencing data of the sample to be tested, and perform data preprocessing on the obtained sequencing data;

[0086] The preprocessed data is compared with the reference gene sequences of multiple pre-defined clone species to obtain the sequence proportion corresponding to each clone species.

[0087] Sort the proportions of multiple sequences from largest to smallest, and label the proportions of multiple sequences in the sorting order as the proportion of the first sequence to the proportion of the M1th sequence, where M1 is the number of clone species;

[0088] The method detects whether the differences between the proportions of the first and third sequences, and between the proportions of the second and third sequences, exceed preset difference thresholds. When both the differences between the proportions of the first and third sequences and between the proportions of the second and third sequences exceed preset difference thresholds, the sample to be tested is determined to be an oligoclonal sample. When the differences between the proportions of the first and third sequences exceed preset difference thresholds but the differences between the proportions of the second and third sequences do not exceed preset difference thresholds, the sample to be tested is determined to be a monoclonal sample. When neither the differences between the proportions of the first and third sequences nor the differences between the proportions of the second and third sequences exceed preset difference thresholds, the sample to be tested is determined to be a polyclonal sample.

[0089] This disclosure also provides a disease typing diagnostic apparatus, including a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to execute the steps of the disease typing diagnostic method described in any embodiment of this disclosure based on the instructions stored in the memory.

[0090] This disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the disease classification and diagnosis method described in any embodiment of this disclosure.

[0091] This disclosure also provides a program product including instructions that, when executed by a computer, perform a disease subtyping diagnosis method as described in any embodiment of this disclosure.

[0092] The disease typing diagnosis method and apparatus of this disclosure identify each TRB clone sequence and determine whether the sample to be tested is a monoclonal sample, oligoclonal sample, or polyclonal sample based on whether the difference between the proportions of the top three sequences with the largest proportion exceeds a preset difference threshold. It can determine the specific sequence of each clone, thereby accurately typing and diagnosing the sample to be tested, and thus realizing the tracking and monitoring of residual small lesions of lymphoma.

[0093] After reading and understanding the accompanying diagrams and detailed descriptions, other aspects can be understood.

[0094] Overview of the attached figures

[0095] The accompanying drawings are provided to further illustrate the technical solutions of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the technical solutions of this disclosure and do not constitute a limitation on the technical solutions of this disclosure. The shapes and sizes of the components in the drawings do not reflect actual proportions and are only intended to illustrate the content of this disclosure.

[0096] Figure 1 is a flowchart illustrating a primer design method provided by an exemplary embodiment of this disclosure;

[0097] Figure 2 is a schematic diagram of the TRB sequence structure;

[0098] Figure 3 is a schematic diagram of TRB detection results for a PBMC negative sample provided by an exemplary embodiment of this disclosure;

[0099] Figure 4 is a schematic diagram of TRB detection results for a lymphoma-positive sample provided by an exemplary embodiment of this disclosure;

[0100] Figure 5 is a flowchart illustrating a standard quality grain design method provided by an exemplary embodiment of this disclosure;

[0101] Figure 6 is a graph showing the detection results of the proportion of TRB VJ plasmid obtained by using uniformity standards and TRB multiple primers in an exemplary embodiment of this disclosure.

[0102] Figure 7 is a schematic flowchart of a standard quality grain quantitative analysis method provided by an exemplary embodiment of the present disclosure;

[0103] Figures 8A and 8B are schematic diagrams showing the clone types and corresponding sequencing sequence numbers of two uniformity standard samples provided in the exemplary embodiments of this disclosure;

[0104] Figures 8C to 8D are schematic diagrams showing the clone types and corresponding sequencing sequence numbers of two experimental standard samples provided in the exemplary embodiments of this disclosure;

[0105] Figure 9 is a flowchart illustrating a primer characterization method provided in an exemplary embodiment of this disclosure;

[0106] Figure 10 is a schematic diagram of a set (10) nucleic acid structures provided in an exemplary embodiment of this disclosure;

[0107] Figure 11 is a schematic diagram of the ligation product of the nucleic acid structure shown in Figure 10 and the primer;

[0108] Figure 12 is a schematic diagram of the process of performing Sanger fragment analysis on the ligation products shown in Figure 11;

[0109] Figure 13 is a flowchart illustrating a sequencing library construction method provided by an exemplary embodiment of this disclosure;

[0110] Figure 14 is a schematic diagram of the lymphoma TRB gene rearrangement detection library construction process provided by an exemplary embodiment of this disclosure;

[0111] Figure 15 is a flowchart illustrating a disease classification and diagnosis method provided by an exemplary embodiment of this disclosure;

[0112] Figures 16A to 16C are schematic diagrams showing the clone types and corresponding sequencing sequence numbers of three TRB samples provided in the exemplary embodiments of this disclosure;

[0113] Figure 17 is a schematic diagram of a primer design device provided in an exemplary embodiment of the present disclosure;

[0114] Figure 18 is a schematic diagram of a standard quality grain design device provided by an exemplary embodiment of the present disclosure;

[0115] Figure 19 is a schematic diagram of a standard quality grain quantitative analysis device provided by an exemplary embodiment of the present disclosure;

[0116] Figure 20 is a schematic diagram of the structure of a disease subtyping diagnostic device provided by an exemplary embodiment of the present disclosure.

[0117] Detailed Explanation

[0118] To make the objectives, technical solutions, and advantages of this disclosure clearer, the embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be arbitrarily combined with each other.

[0119] Unless otherwise defined, the technical or scientific terms used in the embodiments of this disclosure shall have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms "first," "second," and similar terms used in the embodiments of this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" indicate that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, but do not exclude other elements or objects.

[0120] As shown in Figure 1, this disclosure provides a primer design method, including:

[0121] Step 101: Obtain one or more reference data sets, each of which includes multiple reference gene sequences;

[0122] Step 102: For each type of reference data, perform the following operations: Align multiple reference gene sequences by site, determine the conservation score list for each site, obtain the conservation scores of multiple conservation intervals based on the conservation score list for each site, select K conservation intervals with higher conservation scores, where K is a natural number greater than or equal to 1, generate K primer combinations for the K conservation intervals; screen the primers in the K primer combinations and evaluate the screened K primer combinations, and obtain the final primer combination based on the evaluation results.

[0123] The primer design method of this disclosure involves aligning multiple reference gene sequences by site to determine a conservation score list for each site. Based on the conservation score list for each site, multiple conservation intervals are obtained with conservation scores. K conservation intervals with high conservation scores are selected, where K is a natural number greater than or equal to 1. K primer combinations are generated for these K conservation intervals. The primers in the K primer combinations are screened and evaluated. Based on the evaluation results, the final primer combination is obtained. This method can design PCR primers with high specificity and good uniformity at each target site, thereby effectively solving the problems of large PCR amplification bias and incomplete coverage of clone types in current products, which leads to low detection accuracy.

[0124] In this embodiment of the disclosure, a conservative interval refers to the interval in which the sequence similarity between multiple sequences exceeds a preset similarity score threshold. Within a conservative interval, different sequences exhibit a high degree of similarity, which is typically quantified by the percentage of alignment scores. For example, if multiple sequences have a similarity score percentage of over 90% in a region, then this region can be considered very conservative, i.e., this region is a conservative interval.

[0125] In this embodiment of the disclosure, the conservation score list for each site represents the proportion of different base types at each site in multiple sequences, and the conservation score for each conservation interval represents the overall similarity between different sites in each conservation interval and different sequences.

[0126] In some exemplary embodiments, the reference data may be: TRBV reference gene sequence and TRBJ reference gene sequence.

[0127] Studies have shown that TRB rearrangement detection can be used to diagnose and treat various immune-related diseases, such as lymphoma and acute lymphoblastic leukemia. For example, in children with acute B-lymphoblastic leukemia, the EFS (efficacy response time) was significantly lower in the IgH bis / oligoclonal subgroup than in the IgH monoclonal subgroup, indicating that TRB rearrangement status is associated with disease prognosis.

[0128] Using TRB multiplex amplification technology to identify molecular subtypes of lymphoma has the following advantages compared to other technologies:

[0129] (1) High specificity: TRB multiplex amplification can detect clonal immunoglobulin genes in lymphoma cells, thereby determining the molecular subtype of lymphoma, which has high specificity;

[0130] (2) High sensitivity: TRB multiplex amplification can detect very small amounts of lymphoma cells, even in low concentrations of mixed cell samples;

[0131] (3) Fast speed: Multiplex amplification using TRB can be performed quickly, usually yielding results within a few hours, which helps to determine the molecular subtype of lymphoma as early as possible;

[0132] (4) High reliability: The developed TRB multiplex amplification has high accuracy and reliability, which can provide reliable diagnostic and treatment guidance for clinicians.

[0133] For example, when the reference data includes the TRBV reference gene sequence and the TRBJ reference gene sequence, the final primer combination designed is a TRBV-J region multiple specific primer, which can be used to detect the VDJ gene rearrangement of the TRB gene.

[0134] In other words, the primer design method of this disclosure can design a set of highly efficient TRB amplification primers, and design ultramultiplex PCR primers targeting the TRB V and TRB J regions, covering all TRB gene rearrangement clonal types, and simultaneously identifying VDJ rearrangement types. By using the primer set designed in this disclosure, TRB sequences (including amplification of TRB VDJ gene rearrangements) can be amplified multiplexed, which can be used for lymphoma diagnosis and typing.

[0135] In this embodiment of the disclosure, the TRBV reference gene sequence and the TRBJ reference gene sequence can be obtained from the TRB VJ sequence downloaded from the Gemerline database of the IGMT database.

[0136] In some exemplary embodiments, the conservation score list for each site includes five base types A, T, G, C, and N, as well as the percentage of each base type in multiple reference gene sequences.

[0137] As shown in Figure 2, we first need to determine the conservation score list for each site across the entire TRBV and TRBJ intervals. First, for the TRBV interval, we align multiple TRBV reference gene sequences by site, and let P... iFor the position i, there are five base types in multiple TRBV reference gene sequences: A, T, G, C, and N, where N represents an unknown base type. The five base types at each position are sorted from highest to lowest percentage. Therefore, the conservation score list can be represented as a set, where each element contains a base type and its percentage in the sequence. For example, the conservation score list for a certain position {'A': 0.25, 'T': 0.20, 'G': 0.18, 'C': 0.15, 'N': 0.12} represents the percentages of the five base types A, T, G, C, and N at the corresponding positions in multiple reference gene sequences, which are 0.25, 0.20, 0.18, 0.15, and 0.12, respectively. After this calculation, the conservation score list for each position in the TRBV interval is obtained.

[0138] Similarly, a list of conservation scores for each site in the TRBJ interval can be calculated.

[0139] In some exemplary embodiments, the conservation scores of multiple conservation intervals are obtained based on a list of conservation scores for each site, including:

[0140] Based on the pre-set initial conservative interval [start, end] and the sliding window step size W, multiple conservative intervals are obtained through the sliding window method;

[0141] The conservatism score for each conservatism interval is calculated using the following formula: W i =log1 / R i C i =W i ×R i , Among them, W i R represents the weight of the i-th site in each conservatism interval. i C represents the percentage of the most prevalent base type at the i-th site in each conserved interval. i Let S be the conservatism score of the i-th position in each conservatism interval, where S is the conservatism score of each conservatism interval, and n is the length of each conservatism interval, and n = end - start + 1.

[0142] In this embodiment of the disclosure, the conservative range can be obtained by sliding window method or not, and this disclosure does not limit it.

[0143] In this embodiment of the disclosure, the length n of each conservatism interval can be the primer length defined experimentally.

[0144] In this embodiment of the disclosure, the weight W of each site is first calculated based on the proportion of the highest-proportion base type at each site in each conservative interval. Then, the conservative score of each site is calculated based on the weight of each site and the proportion of the highest-proportion base type at each site. Finally, the sum of the conservative scores of each site in the entire interval is taken as the conservative score of the entire conservative interval.

[0145] After obtaining the conservatism scores of all conservatism intervals, all conservatism intervals can be sorted from high to low according to their conservatism scores. The top K conservatism intervals with the highest conservatism scores are selected, where K is a natural number greater than or equal to 1. A primer combination is generated for each of these K conservatism intervals, that is, K primer combinations are generated.

[0146] In some exemplary embodiments, generating K primer combinations for K conserved regions includes:

[0147] For each of the K conservative intervals, perform the following operation:

[0148] Determine the possible base types at each site in the conservatism interval, wherein the possible base types at each site are base types whose proportion is greater than or equal to a preset proportion threshold;

[0149] A primer set is generated based on the possible base types at each site within the conserved region. The number of primers in the primer set is m, where m = ∏m. i m i denoted as the number of possible base types at the i-th site in the conservative interval, ∏ as the quadrature symbol, i being between 1 and n, and n being the length of the conservative interval.

[0150] In this embodiment of the disclosure, for each site in each conservatism interval, an indicator function f can be defined. i (j), f i (j) indicates whether the proportion of the j-th base type at the i-th site is greater than or equal to a preset proportion threshold, where j is between 1 and 4. Since N bases generally have a low proportion, they are not considered here. If the proportion of the j-th base type at the i-th site is greater than the preset proportion threshold θ, then f i (j) = 1; otherwise f i (j) = 0. When generating primer combinations, it is necessary to consider all f values ​​at each site. i For base types where (j) = 1, all f at each site... i By arranging and combining the base types (j) = 1, we can obtain all possible primer sequences for each conserved region, that is, generate a set of primer combinations for each conserved region.

[0151] For example, suppose that the first position of a certain conserved region contains three base types, such as ['A', 'T', 'G'], the second position contains two base types, such as ['G', 'C'], the third position contains only one base type ['T'], and so on. Then the primer combinations generated for this conserved region are:

[0152] [['ACT…'],

[0153] ['AGT…']

[0154] ['TCT…']

[0155] ['TGT…']

[0156] ['GCT…']

[0157] ['GGT…']

[0158] ...

[0159] ]

[0160] In some exemplary embodiments, primers in the K primer combinations are screened based on at least one of the following: dimer, hairpin structure, annealing temperature, and GC content.

[0161] In this embodiment, primers in the K-group primer combination can be screened based on factors such as dimer composition, hairpin structure, Tm temperature (annealing temperature), and GC content. However, this disclosure does not limit this, and users can also screen primers in the K-group primer combination based on other conditions.

[0162] The selection criteria for primers based on dimer formation are as follows: at the experimental temperature Tt, primers should avoid dimer formation as much as possible, where T is the preset first experimental temperature and t is the preset first temperature difference. For example, assuming the preset first experimental temperature T is 60℃ and the preset first temperature difference t is 15℃, if the primer does not form a dimer at an experimental temperature of 45℃, then the primer is retained; if the primer forms a dimer at an experimental temperature of 45℃, then the primer is deleted. Dimers are polymers formed by the combination of complementary bases on two primers during a PCR reaction. The presence of dimers is equivalent to a reduction in the amount of raw material chains that could be used for amplification, thus reducing amplification efficiency. Therefore, it is best to avoid the formation of such substances.

[0163] The selection criteria for primers based on hairpin structure are as follows: at the experimental temperature Tt, the primers should avoid forming hairpin structures as much as possible, where T is the preset first experimental temperature and t is the preset first temperature difference. For example, assuming the preset first experimental temperature T is 60℃ and the preset first temperature difference t is 15℃, if the primer does not form a hairpin structure at the experimental temperature of 45℃, then the primer is retained; if the primer forms a hairpin structure at the experimental temperature of 45℃, then the primer is deleted.

[0164] The selection criteria for primers based on Tm temperature are as follows: the annealing temperature of the primers should be within the preset experimental temperature range. For example, suppose the preset experimental temperature range is [Tm]. low T high ], where T low The lowest temperature, T high The highest temperature is [T]. If the primer annealing temperature is [T], low T high If the primer is within the range of [T], then retain the primer; if the primer's annealing temperature is not within [T], then retain the primer. low T high If the primer is within the specified range, then delete it. For example, [T] low T high The temperature can be [50℃, 60℃], however, this disclosure does not limit it.

[0165] The selection criteria for primers based on GC content are as follows: the GC content of the primers should be within a preset GC content range. For example, assuming the preset GC content range is [G... low G high ], where G low For the lowest GC content, G high The highest GC content is indicated by the primer's GC content being [G]. low G high If the GC content of the primer is within the range of [G], then retain the primer; if the GC content of the primer is not within the range of [G], then retain the primer. low G high If the primer is within the specified range, then delete it. For example, [G] low G high The percentage can be [40%, 60%], however, this disclosure does not limit it.

[0166] In some exemplary embodiments, the K primer combinations are evaluated based on at least one of the following: dimer, hairpin structure, amplification coverage, nonspecific amplification rate (or specificity).

[0167] In this embodiment of the disclosure, when screening primers in a set of primer combinations, one or more redundant primer sequences in the set of primer combinations will be deleted; and when evaluating K sets of primer combinations, one or more sets of primer combinations with poor evaluation results will be deleted, or in other words, the best or better set of primer combinations will be selected from multiple sets of primer combinations.

[0168] Complementarity, dimers, or hairpin structures at the 3' ends of primers can all lead to PCR reaction failure. Therefore, when evaluating a primer combination, if a dimer or hairpin structure is formed in the combination, the combination should be deleted.

[0169] Amplification coverage and specificity are the two most important metrics for evaluating primer effectiveness. Amplification coverage refers to the proportion of target sequences captured by the target primers in an existing database. Specificity refers to the proportion of amplified sequences targeted by a primer combination, i.e., the ratio of specifically amplified sequences to the total sequence.

[0170] For example, for TRB gene rearrangement detection, the final TRB-specific primer combinations are shown in Table 1 (SEQ ID NO:1-47 and SEQ ID NO:48-60) by screening primers in the K primer combination and evaluating the screened K primer combination.

[0171] Table 1

[0172] Secondary structure and hairpin structure analysis showed that this TRB-specific primer combination did not generate secondary structures or hairpin structures under experimental temperatures greater than or equal to 45 degrees Celsius. Amplification simulation analysis confirmed that this TRB-specific primer combination could amplify all functional TRBV and TRBJ genes, with an amplification coverage of 100%.

[0173] The following analysis examines the accuracy of primer and / or kit-based TRB rearrangement detection using a peripheral blood mononuclear cell (PBMC) negative sample and a lymphoma positive sample as examples.

[0174] The PCR reaction system was configured as follows: Thermo AmpliTaq Gold was selected as the PCR amplification reagent. TM360 DNA polymerase or equivalent amplification reagents were used. The primer components and ratios in the primer combination are shown in Table 2. This primer combination and ratio can reduce PCR amplification bias and thus improve detection accuracy. PCR primers were added to DNase-Free & RNase-Free water according to the synthesis report in Table 2 (SEQ ID NO:64-110), with a primer concentration of 100 μM. TRB V and TRB J primers were mixed to form the TRB primer pool.

[0175] Table 2

[0176] The italicized, underlined regions are the linker sequences, while the regular, ununderlined regions are the primer sequences in Table 1. Specifically, TRB V1 to TRB V47 correspond to primers V1 to V47 in Table 1, and TRB J1 to TRB J13 correspond to primers J1 to J13 in Table 1.

[0177] Lymphoma-positive and PBMC-negative samples were selected as PCR templates. Genomic DNA was extracted using the Meiji Bio Universal DNA Extraction Pre-packed Kit or an equivalent kit. The PCR amplification system configuration is shown in Table 3 (First Round PCR Amplification System Table) and Table 4 (Second Round PCR Amplification System Table). The PCR amplification conditions are shown in Table 5 (First Round PCR Amplification Conditions Table) and Table 6 (Second Round PCR Amplification Conditions Table).

[0178] Table 3

[0179] In Table 3, AmpliTaq Gold 360 buffer is a buffer for PCR amplification, dNTPs mixture is a mixture, AmpliTaq Gold 360 DNA polymerase is a polymerase, TRB primer pool is a primer mixture prepared according to the primer concentrations in Table 2, DNase and RNase-free water is nuclease-free water (DNase-free and RNase-free water), X is the volume calculated from 100 ng of template, and T represents Total.

[0180] Table 4

[0181] In Table 4, VAHTS HiFi amplification mixture is a single mixture, P5 adapter primer is a P5 adapter primer, P7 adapter primer is a P7 adapter primer, the first round PCR product is the product purified after the first PCR, and T represents Total.

[0182] Table 5

[0183] Table 6

[0184] The testing steps are as follows:

[0185] 1) AmpliTaq Gold TM 360 DNA polymerase, buffer, MgCl2, dNTPs, polymerase, GC enhancer, homogeneity standard, TRB primer pool, DNase-Free & RNase-Free water were taken out and dissolved on ice;

[0186] 2) Take a PCR reaction tube, add each component according to Table 3, and mix well by pipetting.

[0187] 3) Place the PCR reaction tubes in the PCR instrument, set the PCR amplification program according to Table 5, and run the PCR instrument;

[0188] 4) After PCR amplification, the PCR product is purified using magnetic beads, with the ratio of magnetic beads being 0.8 times the amplification volume;

[0189] 5) Take a new PCR reaction tube, add each component according to Table 4, and mix well by pipetting.

[0190] 6) Place the PCR reaction tubes in the PCR instrument, set the PCR amplification program according to Table 6, and run the PCR instrument;

[0191] 7) After PCR amplification, the PCR product is purified using magnetic beads at a ratio of 0.8 times the amplification volume. The purified product is the sequencing library.

[0192] 8) The library was subjected to high-throughput sequencing using an Illumina NovaSeq sequencer with a read length of PE150.

[0193] Figure 3 shows the TRB detection results for PBMC-negative samples, and Figure 4 shows the TRB detection results for lymphoma-positive samples. In Figures 3 and 4, the horizontal axis represents the CDR3 sequence length (i.e., the rearrangement position), and the vertical axis represents the proportion of detected plasmid sequences. The same vertical bar includes the proportion of plasmid sequences with the same CDR3 sequence length but different CDR3 sequences. The proportion of plasmid sequences with a certain CDR3 sequence length is the ratio of the number of detected plasmid sequences of that CDR3 sequence length to the total number of detected sequences. As can be seen from Figures 3 and 4, the proportion of sequences detected in PBMC-negative samples with different CDR3 sequence lengths is relatively close, showing the polyclonal nature of TRB gene rearrangement. In contrast, lymphoma-positive samples show the proliferation of a specific TRB rearrangement sequence at CDR3 sequence length 61, showing monoclonal nature. This indicates that the rearrangement detection accuracy of this TRB multiple primer combination and / or kit is high.

[0194] As shown in Figure 5, this disclosure also provides a standard quality grain design method, including:

[0195] Step 501: Obtain the first gene cluster and the second gene cluster. The first gene cluster includes multiple first gene sequence fragments, and the second gene cluster includes multiple second gene sequence fragments. Both the first gene cluster and the second gene cluster contain multiple functional fragments.

[0196] Step 502: Combine multiple first gene sequence fragments and multiple second gene sequence fragments to obtain multiple fragment groups, each fragment group containing one first gene sequence fragment and one second gene sequence fragment;

[0197] Step 503: Insert a non-human sequence between the first and second gene sequence fragments in each fragment group to obtain multiple plasmid sequences;

[0198] Step 504: Determine the ratio of each plasmid sequence to obtain the designed standard plasmid.

[0199] In conventional analytical methods, plasmid sequences do not carry UMI sequence identifiers. Plasmid sequences are identified through sequence alignment. However, this analytical method has high identification rate in samples containing only plasmid sequences. But when peripheral blood samples are mixed, only most plasmid sequences can be identified. It is difficult to identify whether some sequences are from peripheral blood samples or plasmid sequences.

[0200] The standard plasmid design method of this disclosure involves inserting a non-human sequence between the first and second gene sequence fragments in each fragment group to obtain multiple plasmid sequences. All plasmid sequences can be distinguished from rearranged sequences derived from peripheral blood samples by the non-human sequence, thus enabling accurate identification and quantification of plasmid sequences.

[0201] In this embodiment of the disclosure, the length of the inserted non-human sequence is n3 bp, where n3 is between 6 and 10. For example, n3 = 8.

[0202] In this embodiment of the disclosure, the terminal n1 bp of the first gene sequence fragment, the inserted non-human sequence, and the terminal n2 bp of the second gene sequence fragment constitute a (n1+n2+n3) bp unique molecular identifier (UMI) sequence identifier. The UMI sequence identifiers in different plasmid sequences have at least one site with different base types.

[0203] UMI sequence identifiers, as unique identifiers for sequences, are used for sequence identification and classification in subsequent analyses. Specifically, different UMI sequence identifiers distinguish DNA templates from different sources, differentiating between false-positive mutations caused by random errors during PCR amplification and sequencing, and mutations truly carried by the patient, thereby improving the sensitivity and specificity of the detection.

[0204] In this embodiment of the disclosure, n1 is between 2 and 6, and n2 is between 2 and 6. For example, n1 = 4, n2 = 4.

[0205] In some exemplary embodiments, the first gene cluster may be the TRBD gene cluster, and the second gene cluster may be the TRBJ gene cluster.

[0206] For example, the first gene cluster can be a TRBD gene cluster containing all functional fragments, and the second gene cluster can be a TRBJ gene cluster containing all functional fragments. The TRB gene sequence can be sourced from the IMGT database. All TRBD and TRBJ fragments are randomly combined, and an 8 bp non-human sequence is added to each combination to obtain multiple plasmid sequences. In each plasmid sequence, the 4 bp terminal sequence of the TRBD fragment, the 8 bp non-human sequence, and the 4 bp terminal sequence of the TRBJ fragment together form a 16 bp UMI sequence. The generated UMI sequence combinations are shown in Table 7 (SEQ ID NO: 123-178).

[0207] Table 7

[0208] Table 7 shows that the non-human random sequence inserted into each UMI sequence is just an example. Users can redesign the plasmid sequence and the inserted non-human random sequence according to their needs, as long as the base types of at least one site are different in different UMI sequence identifiers.

[0209] In some exemplary embodiments, when the standard plasmid is a uniformity standard, the ratio of each plasmid sequence is a uniform ratio of equal concentration.

[0210] In some exemplary embodiments, when the standard plasmid is an experimental standard, the proportion of one or more plasmid sequences is greater than the proportion of the remaining plasmid sequences (plasmid sequences other than one or more plasmid sequences).

[0211] In this embodiment of the disclosure, the uniformity standard is a standard prepared by mixing each plasmid in the same proportion. The experimental standard is a standard prepared by mixing one or more plasmids in a certain high proportion. In some exemplary embodiments, a control standard may also be provided, which is a standard in which no plasmid is added to the sample and pure water is used instead.

[0212] Homogeneity standards can be used to verify the amplification efficiency of primer combinations in a single experiment; control standards can be used to verify whether there is contamination in a single experiment; experimental standards are used to simulate polyclonal or monoclonal experiments; a certain proportion of homogeneity standards can be added to quantify unknown experimental samples to detect the rearrangement type and quantification of the sample itself; homogeneity standards or experimental standards can be used to verify the influence of different experimental reagents and conditions on experimental results.

[0213] The following example uses a PCR amplification uniformity experiment using uniformity standards to verify the amplification efficiency of TRB multiple primer combinations and / or kits.

[0214] The PCR reaction system was configured as follows: Thermo AmpliTaq Gold was selected as the PCR amplification reagent. TM 360 DNA polymerase or equivalent amplification reagents; PCR primers were prepared by adding DNase-free and RNase-free water according to the synthesis report in Table 2, with a primer concentration of 100 μM; TRB V and TRB J primers were mixed to form the TRB primer pool; 58 standard TRB VDJ plasmid sequences were designed, each plasmid sequence including a random fragment of the TRB V / D / J sequence and an 8-base UMI sequence identifier, as shown in Table 7 above.

[0215] The test sample was a homogeneity standard, with a total volume of approximately 10,000 copies. The plasmid concentration was quantified and its molar concentration calculated using Qubit 4.0. The plasmids were mixed equimolarly, amplified using M13F and M13R primers (universal primers), and a sequencing library was constructed. The number and proportion of each plasmid were counted using UMI (Uniqueness Index), and the molar ratio of each plasmid was adjusted to 0.95-1.05, which constituted the homogeneity standard.

[0216] The configuration of the PCR amplification system is shown in Table 3 (first round PCR amplification system) and Table 4 (second round PCR amplification system) above, and the PCR amplification conditions are shown in Table 5 (first round PCR amplification conditions) and Table 6 (second round PCR amplification conditions) above.

[0217] The testing steps are as follows:

[0218] 1) AmpliTaq Gold TM 360 DNA polymerase, buffer, MgCl2, dNTPs, polymerase, GC enhancer, homogeneity standard, TRB primer pool, DNase-Free & RNase-Free water were taken out and dissolved on ice;

[0219] 2) Take a PCR reaction tube, add each component according to Table 3, and mix well by pipetting.

[0220] 3) Place the PCR reaction tubes in the PCR instrument, set the PCR amplification program according to Table 5, and run the PCR instrument;

[0221] 4) After PCR amplification, the PCR product is purified using magnetic beads, with the ratio of magnetic beads being 0.8 times the amplification volume;

[0222] 5) Take a new PCR reaction tube, add each component according to Table 4, and mix well by pipetting.

[0223] 6) Place the PCR reaction tubes in the PCR instrument, set the PCR amplification program according to Table 6, and run the PCR instrument;

[0224] 7) After PCR amplification, the PCR product is purified using magnetic beads at a ratio of 0.8 times the amplification volume. The purified product is the sequencing library.

[0225] 8) Perform high-throughput sequencing on the sequencing library using an Illumina NovaSeq sequencer with a read length of PE150.

[0226] Figure 6 is a detection result diagram of the proportion of TRB plasmids obtained by using uniformity standards and TRB multiple primers according to an exemplary embodiment of the present disclosure. In Figure 6, the horizontal axis represents the plasmid sequence number, and the vertical axis represents the proportion of detected plasmid sequences. The proportion of plasmid sequences with a certain number is the ratio of the number of detected plasmid sequences with that number to the total number of detected sequences. As can be seen from Figure 6, the PCR amplification uniformity of the TRB multiple primer combination and / or kit is good.

[0227] As shown in Figure 7, this disclosure also provides a method for quantitative analysis of standard quality grains, including:

[0228] Step 701: Obtain the paired-end sequencing data corresponding to the standard quality plasmid, which includes the Reads1 sequence and the Reads2 sequence;

[0229] Step 702: Determine whether the Reads1 and Reads2 sequences include a UMI sequence identifier;

[0230] Step 703: When both Reads1 and Reads2 sequences include UMI sequence identifiers and the included UMI sequence identifiers are the same, it is determined that the plasmid sequences corresponding to Reads1 and Reads2 sequences have been identified.

[0231] Step 704: When only one of the Reads1 and Reads2 sequences contains a UMI sequence identifier, determine the Hamming distance between the Reads sequence that does not contain a UMI sequence identifier and the UMI sequence identifier; when the determined Hamming distance is less than or equal to a, determine that the plasmid sequence corresponding to the Reads1 and Reads2 sequences has been identified, where a is a natural number less than or equal to 2.

[0232] The standard plasmid quantitative analysis method disclosed herein determines whether Reads1 and Reads2 sequences include a UMI sequence identifier. When both Reads1 and Reads2 sequences include the same UMI sequence identifier, the plasmid sequence corresponding to Reads1 and Reads2 sequences is identified. When only one of Reads1 and Reads2 sequences includes the UMI sequence identifier, the Hamming distance between the Reads sequence without the UMI sequence identifier and the UMI sequence identifier is determined. When the determined Hamming distance is less than or equal to 'a', the plasmid sequence corresponding to Reads1 and Reads2 sequences is identified, where 'a' is a natural number less than or equal to 2. This method can accurately perform quantitative analysis of standard plasmids, and further, the amplification results of samples can be accurately inferred based on the quantitative analysis results of standard plasmids.

[0233] In this embodiment of the disclosure, when neither the Reads1 sequence nor the Reads2 sequence contains a UMI sequence identifier, or when only one of the Reads1 sequence and the Reads2 sequence contains a UMI sequence identifier and the Hamming distance between the Reads sequence containing a UMI sequence identifier and the UMI sequence identifier is greater than a, it is determined that no plasmid sequence corresponding to the Reads1 sequence and the Reads2 sequence has been identified.

[0234] In this embodiment of the disclosure, a can be equal to 1. When a is equal to 1, the identification method is more rigorous, which allows for more accurate quantitative analysis of the standard quality particles.

[0235] Paired-end sequencing performs sequencing from both ends of the insert fragment. The ATCG sequence read from each end is called a read. Each insert fragment will generate two reads, namely reads1 and reads2. The reads1 and reads2 data corresponding to a sample are stored in two compressed packages.

[0236] In this embodiment of the disclosure, when the sequencing data corresponding to the standard plasmid is single-end sequencing data, it is determined whether the Reads sequence includes a UMI sequence identifier; when the Reads sequence includes a UMI sequence identifier, it is determined that the plasmid sequence corresponding to the Reads sequence has been identified; when the Reads sequence does not include a UMI sequence identifier, it is determined that the plasmid sequence corresponding to the Reads sequence has not been identified.

[0237] As shown in Figures 8A and 8B, in one experiment, TRB multiple primers were used for amplification. Both samples A-1 and A-2 were homogeneity standards, and the sequencing data of these two samples showed that the proportions of each plasmid were relatively close.

[0238] As shown in Figures 8C and 8D, in one experiment, two samples, TRB A-1 and TRB A-2, were experimental standards. In both samples, the amount of one plasmid added was higher than that of the other plasmids, and the sequencing data showed that the plasmid had a higher read coverage.

[0239] As shown in Figure 9, this disclosure also provides a primer characterization method, including:

[0240] Step 901: Obtain the sequencing data corresponding to the primer to be tested, and determine the 5' end position of the primer to be tested based on the sequencing data;

[0241] Step 902: Design multiple sequences based on the 5' end position of the primer to be detected. Each of the multiple sequences starts at the 5' end position of the primer to be detected, and the 3' end cutoff positions of two adjacent sequences differ by a single base.

[0242] Step 903: Design multiple nucleic acid structures based on the designed sequences;

[0243] Step 904: Identify the nucleic acid structures among multiple nucleic acid structures that can complementary pair with the primers to be detected;

[0244] Step 905: Determine the sequence of the primer to be tested based on the nucleic acid structure that can complement the primer to be tested.

[0245] The primer characterization method of this disclosure determines the 5' end position of the primer to be detected based on sequencing data, designs multiple sequences and multiple corresponding nucleic acid structures based on the 5' end position of the primer to be detected, then determines the nucleic acid structure that can complementarily pair with the primer to be detected among the multiple nucleic acid structures, and determines the sequence of the primer to be detected based on the nucleic acid structure that can complementarily pair with the primer to be detected, thus enabling the sequencing of any unknown primer.

[0246] In some exemplary embodiments, the method further includes, prior to: performing PCR amplification on the primers to be detected, and obtaining sequencing data and the corresponding 5' end position of the primers to be detected by high-throughput sequencing.

[0247] In some exemplary implementations, the length of each sequence is between 18 bp and 27 bp.

[0248] In some exemplary embodiments, the number of sequences is 10.

[0249] Since primers are typically between 18 and 27 bases in length, a set of 10 sequences is designed based on the 5' end position of the primer to be detected, with each sequence differing by a single base at the 3' end, and each sequence being between 18 and 27 bp in length.

[0250] In some exemplary embodiments, each nucleic acid structure includes a hairpin structure, and each nucleic acid structure has an inverse complementary sequence fused to its end. The 5' end of each nucleic acid structure is modified with a fluorescent label.

[0251] As shown in Figure 10, a set (10) of nucleic acid structures for detection primers were designed. Each nucleic acid structure includes an artificially designed hairpin structure, reverse complementary sequences of different lengths fused to the ends, and fluorescent labels modified with bases at the ends.

[0252] In some exemplary embodiments, identifying nucleic acid structures among multiple nucleic acid structures that can complementaryly pair with the primer to be detected includes:

[0253] For each nucleic acid structure, the following steps were performed: the primer to be tested was mixed with the nucleic acid structure, and denaturation, annealing, and ligation were performed to obtain the ligation product; the length of the ligation product was then detected.

[0254] Select the nucleic acid structure corresponding to the ligation product with a length greater than the preset length threshold as a nucleic acid structure that can complementarily pair with the primer to be detected.

[0255] For example, the primers to be detected are mixed with nucleic acid structures at equimolar concentrations according to Table 8:

[0256] Table 8

[0257] After mixing, place the mixture in a boiling water bath for 5 minutes, turn off the heating switch, and let it stand to room temperature. This step usually takes 8 to 12 hours.

[0258] Configure the connection system according to Table 9:

[0259] Table 9

[0260] Mix all components in the connection system thoroughly, centrifuge the liquid to the bottom of the tube, and react at 25°C for 30 minutes.

[0261] As shown in Figure 11, the primers to be tested anneal to the paired nucleic acid structures to form double-stranded structures, which are then ligated by T4 DNA ligase. Structures that cannot be paired cannot be ligated, resulting in nucleic acid structures and ligation products of different lengths.

[0262] In this embodiment of the disclosure, the ligation product can be detected using Sanger fragment analysis. As shown in Figure 12, the T4 DNA ligase ligates the primer to be tested to an artificially designed nucleic acid structure, which is then denatured into a single strand by formamide. The corresponding ligation product is the longest, while the unligated nucleic acid structure and the free primer are shorter. Sanger fragment analysis can determine the 3' cutoff position of the primer to be tested, thereby obtaining the full-length sequence of the primer to be tested.

[0263] As shown in Figure 13, this disclosure also provides a method for constructing a sequencing library, including:

[0264] Step 1301: Extract DNA from the genome to be tested;

[0265] Step 1302: Design specific primers for the genome to be detected, mix the specific primers with the extracted DNA, perform the first round of PCR amplification to obtain the first round of PCR products, and purify the first round of PCR products;

[0266] Step 1303: Mix the purified first-round PCR product with universal adapter primers, perform a second-round PCR amplification to obtain the second-round PCR product, purify the second-round PCR product, and obtain the constructed sequencing library.

[0267] The sequencing library construction method of this disclosure uses a two-step PCR method to ligate sequencing adapters to the 5' and 3' ends of the target fragment, respectively, to convert the target DNA to be sequenced into a form that can be read by the sequencer.

[0268] As shown in Figure 14, the universal adapter primers include the P5 adapter primer and the P7 adapter primer.

[0269] For example, the genome to be detected is the TRB genome, and the specific primers include a forward primer targeting the TRB V region and a reverse primer targeting the TRB J region.

[0270] The design methods for specific primers can refer to the primer design methods described above, and will not be repeated here.

[0271] As shown in Figure 15, this embodiment of the present disclosure also provides a disease classification and diagnosis method, including:

[0272] Step 1501: Obtain the sequencing data of the sample to be tested, and perform data preprocessing on the obtained sequencing data;

[0273] Step 1502: Compare the preprocessed data with the reference gene sequences of multiple preset clone species to obtain the sequence proportion corresponding to each clone species;

[0274] Step 1503: Sort the proportions of multiple sequences from largest to smallest, and mark the proportions of multiple sequences in the sorting order as the proportion of the first sequence to the proportion of the M1th sequence, where M1 is the number of clone species;

[0275] Step 1504: Detect whether the differences between the proportions of the first and third sequences, and between the proportions of the second and third sequences, exceed preset difference thresholds; when both the differences between the proportions of the first and third sequences and between the proportions of the second and third sequences exceed preset difference thresholds, the sample to be tested is determined to be an oligoclonal sample; when the differences between the proportions of the first and third sequences exceed preset difference thresholds but the differences between the proportions of the second and third sequences do not exceed preset difference thresholds, the sample to be tested is determined to be a monoclonal sample; when neither the differences between the proportions of the first and third sequences nor the differences between the proportions of the second and third sequences exceed preset difference thresholds, the sample to be tested is determined to be a polyclonal sample.

[0276] In this embodiment of the disclosure, the proportions of multiple sequences can be regarded as multiple clonal distribution peaks. When the differences between the proportions of the first sequence and the third sequence, as well as the differences between the proportions of the second sequence and the third sequence, both exceed a preset difference threshold, it can be considered that there are two main peaks among the multiple clonal distribution peaks, and the sample to be detected is determined to be an oligoclonal sample. When the difference between the proportions of the first sequence and the third sequence exceeds a preset difference threshold, but the difference between the proportions of the second sequence and the third sequence does not exceed a preset difference threshold, it can be considered that there is only one main peak among the multiple clonal distribution peaks, and the sample to be detected is determined to be a monoclonal sample. When the differences between the proportions of the first sequence and the third sequence, as well as the differences between the proportions of the second sequence and the third sequence, do not exceed a preset difference threshold, it can be considered that there is no main peak among the multiple clonal distribution peaks, and the sample to be detected is determined to be a polyclonal sample.

[0277] In this embodiment of the disclosure, when the sample to be tested is detected as an oligoclonal sample or a monoclonal sample, the diagnostic result of the sample to be tested can be considered as positive, and the corresponding disease subtype can be determined based on the number of clonal distribution peaks and the type of clones; when the sample to be tested is determined to be a polyclonal sample, the diagnostic result of the sample to be tested can be considered as negative.

[0278] The disease typing and diagnosis method of this disclosure identifies each TRB clone sequence and determines whether the sample to be tested is a monoclonal sample, oligoclonal sample, or polyclonal sample based on whether the difference between the proportions of the top three sequences with the largest proportion exceeds a preset difference threshold. It can determine the specific sequence of each clone, thereby accurately typing and diagnosing the sample to be tested, and thus realizing the tracking and monitoring of residual small lesions of lymphoma.

[0279] In some exemplary embodiments, the acquired sequencing data undergoes data preprocessing, including:

[0280] Filter out primer sequences, low-quality sequences, sequences with a high proportion of N bases, and sequences whose length is lower than a preset length threshold from the sequencing data;

[0281] The sequencing data were subjected to adapter sequence detection and adapter sequence removal.

[0282] Remove the low-quality bases at both ends of the sequence.

[0283] In this embodiment of the disclosure, during data preprocessing, the sequencing data is first statistically analyzed, including the amount of sequencing data and its quality value. Then, the sequencing data is filtered, including filtering out sequences containing adapter sequences and with a 5' end of polyN, sequences with an average quality value lower than a preset quality threshold Q, sequences with an excessively high proportion of N bases, and sequences with a length lower than a preset length threshold. For example, the preset quality threshold Q can be 25. When the quality value of a base is greater than or equal to 25, it can be considered reliable sequencing, with a corresponding error rate of approximately 0.3%. When the proportion of N bases in a sequence is higher than a preset proportion threshold M2 (for example, M2 can be set to 10%), it is considered to have a high proportion of N bases. When the average quality value of a sequence is lower than the preset quality threshold Q or the proportion of N bases in a sequence is higher than the preset proportion threshold M, the sequence is filtered out. Furthermore, when the sequence length is lower than a preset length threshold L, the sequence is filtered out.

[0284] Data preprocessing also includes the removal of adapter sequences and low-quality bases: adapter sequences are detected and removed from the sequence, using general-purpose software such as trim_galore or cutadapter. When the base quality values ​​at both ends of the sequence are lower than a preset quality threshold Q, those bases are removed.

[0285] By preprocessing data, we can obtain high-quality data for subsequent analysis.

[0286] In some exemplary embodiments, the reference gene sequences for multiple clone species are: the TRBV reference gene sequence and the TRBJ reference gene sequence.

[0287] Assume the proportion of the i-th sequence is A i In some exemplary embodiments, detecting whether the difference between the proportion of the first sequence and the proportion of the third sequence exceeds a preset difference threshold includes: detecting whether the ratio A1 / A3 of the proportion of the first sequence to the proportion of the third sequence is greater than or equal to X; when A1 / A3≥X, it is determined that the difference between A1 and A3 exceeds the preset difference threshold; when the ratio A1 / A3 of the proportion of the first sequence to the proportion of the third sequence is <X, it is determined that the difference between A1 and A3 does not exceed the preset difference threshold, where X is the preset difference threshold and X is greater than or equal to 2.

[0288] Similarly, detecting whether the difference between the proportion of the second sequence and the proportion of the third sequence exceeds a preset difference threshold includes: detecting whether the ratio of the proportion of the second sequence to the proportion of the third sequence, A2 / A3, is greater than or equal to X. When A2 / A3≥X, it is determined that the difference between A2 and A3 exceeds the preset difference threshold; when the ratio of the proportion of the second sequence to the proportion of the third sequence, A2 / A3<X, it is determined that the difference between A2 and A3 does not exceed the preset difference threshold, where X is the preset difference threshold, and X is greater than or equal to 2.

[0289] In this embodiment of the disclosure, the preprocessed sequence is subjected to TRB rearrangement identification. To determine if the sequence is a VDJ rearrangement, the sequence is aligned to the IMGT database, and the V, D, and J fragments contained in the sequence can be identified. When the sequence contains the V, D, and J sequences of TRB, the sequencing data is considered to contain a TRB VDJ gene rearrangement. If no TRB sequence is identified in the sequence, it is considered a non-specific amplification sequence. The number of detected clone sequences corresponding to each rearrangement type is determined.

[0290] The obtained clone sequences are used to identify the clonal distribution of the sample to be tested, mainly including polyclonal, oligoclonal and monoclonal sequences.

[0291] Monoclonal sample: A single main peak is detected in the amplification product (taking X=2 as an example, the ratio of the height of the first highest clonal distribution peak to the height of the third highest clonal distribution peak is greater than or equal to twice, and the ratio of the height of the second highest clonal distribution peak to the height of the third highest clonal distribution peak is less than twice).

[0292] Oligoclonal sample: Two main peaks were detected in the amplification product (taking X=2 as an example, the ratio of the height of the first highest clonal distribution peak to the height of the third highest clonal distribution peak and the ratio of the height of the second highest clonal distribution peak to the height of the third highest clonal distribution peak are both greater than or equal to two).

[0293] Multiple clone samples: Multiple gene rearrangement distribution peaks were detected in the amplification products, with no main peak (taking X=2 as an example, the ratio of the height of the first highest clonal distribution peak to the height of the third highest clonal distribution peak is less than two times).

[0294] Assuming the sequencing data contains M1 clone species, let A represent the sequence proportion of the i-th clone species. i M1 represents the number of clones, where i is the index of the clone species, and 1 ≤ i ≤ M1. For example, suppose the sequencing data contains 30 TRB VJ gene rearrangements, then M1 = 30.

[0295] Sort all clone species in descending order of their sequence proportions to obtain an ordered sequence A1, A2, ..., A M1 Where A1≥A2...≥A M1 .

[0296] The method detects whether the differences between the proportions of the first sequence A1 and the third sequence A3, and between the proportions of the second sequence A2 and the third sequence A3, exceed preset difference thresholds. When both the differences between the proportions of the first sequence A1 and the third sequence A3, and between the proportions of the second sequence A2 and the third sequence A3, exceed the preset difference thresholds, the sample to be tested is determined to be an oligoclonal sample. When the differences between the proportions of the first sequence A1 and the third sequence A3 exceed the preset difference thresholds, but the differences between the proportions of the second sequence A2 and the third sequence A3 do not exceed the preset difference thresholds, the sample to be tested is determined to be a monoclonal sample. When neither the differences between the proportions of the first sequence A1 and the third sequence A3, nor the differences between the proportions of the second sequence A2 and the third sequence A3, exceed the preset difference thresholds, the sample to be tested is determined to be a polyclonal sample.

[0297] In this embodiment of the disclosure, when the sample to be tested is a polyclonal sample, the sample to be tested is determined to be a negative sample; when the sample to be tested is an oligoclonal sample or a monoclonal sample, the sample to be tested is determined to be a positive sample.

[0298] When performing multiplex primer amplification experiments, adjustments to reagents, primer combinations, and amplification temperatures may be necessary multiple times, requiring evaluation of the results of each experiment. The following indicators can be used to assess the amplification effectiveness of each experiment.

[0299] Amplification homogeneity: This assesses whether the sequence coverage is uniform due to amplification bias. Quantitative analysis is performed using known standard plasmids. Relative standard deviation (RSD) is used to quantify amplification homogeneity. Known homogeneity standards involve proportional input of all plasmids, but amplification results will still exhibit bias. RSD = (Standard Deviation / Mean) * 100. A smaller RSD value indicates higher amplification homogeneity, meaning the amplification results of multiple plasmids are more similar and the amplification efficiency is more uniform; conversely, a larger RSD value indicates less uniform amplification efficiency.

[0300] Amplification specificity: This is used to assess whether the amplification products specifically match the target sequence. It is evaluated by calculating the ratio of the number of amplification products in non-target regions to the total number of amplification products. The higher the specificity, the lower the proportion of non-target amplification.

[0301] Amplification coverage: assess the extent to which the standard plasmid is covered by the sequencing sequence to identify whether any plasmids have not been amplified or have been over-amplified.

[0302] The following experimental analysis example illustrates the disease typing and diagnosis method of this disclosure. In this example, three samples are all lymphoma samples mixed with homogeneity standards to perform typing and diagnosis of lymphoma samples, while homogeneity standards are used to quantify the samples.

[0303] Specifically, the data analysis process includes the following steps:

[0304] 1) Perform data statistics on the sequencing data of the samples to be tested, including the amount of sequencing data and the Q30 ratio, as shown in Table 10.

[0305] Table 10

[0306] 2) Subsequent data preprocessing included sequencing data statistics and low-quality data filtering. Low-quality data included sequences containing only primer sequences with polyN at the end, sequences with low average base quality, sequences containing structural sequences, and sequences with excessively high N content. In this example, the filtering conditions were: Q value below 20 for low-quality bases, removal of adapter sequences longer than 3 bp at the end, and discarding sequences with a length less than 50 bp after removal. The filtering results are shown in Table 11.

[0307] Table 11

[0308] 3) Clone identification and authentication

[0309] The clone rearrangements in the samples were identified and statistically analyzed, including the total number of clone sequences, clone types, and the percentage of clone sequences in the sequence (Ratio). The identification results are shown in Table 12, which show that the sequencing data of most samples contain a high proportion of clone sequences.

[0310] Table 12

[0311] Figures 16A to 16C are schematic diagrams illustrating the clone types and corresponding sequencing sequence numbers of three samples (TRB B-1, TRB B-2, and TRB B-3) provided in the exemplary embodiments of this disclosure. As shown in Figures 16A to 16C, the RSD of all three samples is less than 5, indicating good amplification uniformity. The main peak of all three samples represents the target clone of the lymphoma sample. Simultaneously, the amplification specificity is greater than 90%, demonstrating good amplification specificity. Sample 1 (TRB B-1), Sample 2 (TRB B-2), and Sample 3 (TRB B-3) are lymphoma samples incorporating uniformity standards; the amplification results show amplification of all standard plasmids, with a coverage rate of 100%.

[0312] This disclosure utilizes TRB multiplex amplification to achieve highly accurate and sensitive diagnosis and typing of lymphoma. It employs algorithmic design of highly specific multiplex amplification primers, standard quality granules for uniform amplification and quantification of experimental samples, and the aforementioned data analysis process for typing and diagnosis. This approach overcomes the low sensitivity issues of traditional capillary electrophoresis + fragment analysis, and addresses the high bias and low coverage problems of previous high-throughput sequencing PCR amplification methods, effectively improving the detection rate and accuracy of lymphoma.

[0313] This disclosure also provides a primer design apparatus, including a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to perform the steps of the primer design method as described in any embodiment of this disclosure based on the instructions stored in the memory.

[0314] As shown in Figure 17, in one example, the primer design device may include: a first processor 1710, a first memory 1720, a first bus system 1730, and a first transceiver 1740, wherein the first processor 1710, the first memory 1720, and the first transceiver 1740 are connected through the first bus system 1730, the first memory 1720 is used to store instructions, and the first processor 1710 is used to execute the instructions stored in the first memory 1720 to control the first transceiver 1740 to transmit and receive signals. Specifically, the first transceiver 1740, under the control of the first processor 1710, can acquire one or more reference data. Each type of reference data includes multiple reference gene sequences. For each type of reference data, the first processor 1710 performs the following operations: aligning the multiple reference gene sequences by site, determining a conservation score list for each site, obtaining conservation scores for multiple conservation intervals based on the conservation score list for each site, selecting K conservation intervals with higher conservation scores, where K is a natural number greater than or equal to 1, generating K primer combinations for the K conservation intervals; screening the primers in the K primer combinations and evaluating the screened K primer combinations, obtaining the final primer combination based on the evaluation results.

[0315] It should be understood that the first processor 1710 can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0316] The first memory 1720 may include read-only memory and random access memory, and provides instructions and data to the first processor 1710. A portion of the first memory 1720 may also include non-volatile random access memory. For example, the first memory 1720 may also store device type information.

[0317] In addition to the data bus, the first bus system 1730 may also include a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as the first bus system 1730 in Figure 17.

[0318] In implementation, the processing performed by the processing device can be accomplished through integrated logic circuits in the hardware of the first processor 1710 or through software instructions. That is, the method steps of this embodiment can be executed by the hardware processor, or by a combination of hardware and software modules within the processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other storage media. This storage medium is located in the first memory 1720. The first processor 1710 reads information from the first memory 1720 and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, further details are omitted here.

[0319] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the primer design method as described in any embodiment of this disclosure. The primer design method driven by executing executable instructions is essentially the same as the primer design method provided in the above embodiments of this disclosure, and will not be described in detail here.

[0320] In some possible implementations, various aspects of the primer design methods provided in this disclosure may also be implemented as a program product comprising program code that, when run on a computer device, causes the computer device to perform the steps in the primer design methods according to various exemplary embodiments of this disclosure as described above. For example, the computer device may execute the primer design methods described in the embodiments of this disclosure.

[0321] This disclosure also provides a standard plasmid design apparatus, including a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to perform the steps of the standard plasmid design method as described in any embodiment of this disclosure based on the instructions stored in the memory.

[0322] As shown in Figure 18, in one example, the standard quality grain design device may include: a second processor 1810, a second memory 1820, a second bus system 1830, and a second transceiver 1840, wherein the second processor 1810, the second memory 1820, and the second transceiver 1840 are connected through the second bus system 1830, the second memory 1820 is used to store instructions, and the second processor 1810 is used to execute the instructions stored in the second memory 1820 to control the second transceiver 1840 to transmit and receive signals. Specifically, the second transceiver 1840, under the control of the second processor 1810, can acquire a first gene cluster and a second gene cluster. The first gene cluster includes multiple first gene sequence fragments, and the second gene cluster includes multiple second gene sequence fragments. Both the first and second gene clusters contain multiple functional fragments. The second processor 1810 combines the multiple first gene sequence fragments and the multiple second gene sequence fragments to obtain multiple fragment groups. Each fragment group contains one first gene sequence fragment and one second gene sequence fragment. A non-human sequence is inserted between the first and second gene sequence fragments in each fragment group to obtain multiple plasmid sequences. The ratio of each plasmid sequence is determined to obtain a designed standard plasmid.

[0323] It should be understood that the second processor 1810 can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0324] The second memory 1820 may include read-only memory and random access memory, and provides instructions and data to the second processor 1810. A portion of the second memory 1820 may also include non-volatile random access memory. For example, the second memory 1820 may also store device type information.

[0325] In addition to the data bus, the second bus system 1830 may also include a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as the second bus system 1830 in Figure 18.

[0326] In implementation, the processing performed by the processing device can be accomplished through integrated logic circuits in the hardware of the second processor 1810 or through software instructions. That is, the method steps of this embodiment can be executed by a hardware processor, or by a combination of hardware and software modules within the processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other storage media. This storage medium is located in the second memory 1820. The second processor 1810 reads information from the second memory 1820 and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, further details are omitted here.

[0327] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the standard protogranule design method as described in any embodiment of this disclosure. The standard protogranule design method driven by executing executable instructions is essentially the same as the standard protogranule design method provided in the above embodiments of this disclosure, and will not be described in detail here.

[0328] In some possible implementations, various aspects of the standard protogranule design method provided in this disclosure may also be implemented as a program product comprising program code that, when run on a computer device, causes the computer device to perform the steps in the standard protogranule design method according to various exemplary embodiments of this disclosure as described above. For example, the computer device may execute the standard protogranule design method described in the embodiments of this disclosure.

[0329] This disclosure also provides a standard plasmid quantification analysis apparatus, including a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to perform the steps of the standard plasmid quantification analysis method as described in any embodiment of this disclosure based on the instructions stored in the memory.

[0330] As shown in Figure 19, in one example, the standard quality grain quantitative analysis device may include: a third processor 1910, a third memory 1920, a third bus system 1930, and a third transceiver 1940. The third processor 1910, the third memory 1920, and the third transceiver 1940 are connected through the third bus system 1930. The third memory 1920 is used to store instructions, and the third processor 1910 is used to execute the instructions stored in the third memory 1920 to control the third transceiver 1940 to transmit and receive signals. Specifically, the third transceiver 1940, under the control of the third processor 1910, can acquire paired-end sequencing data corresponding to the standard plasmid. The paired-end sequencing data includes Reads1 and Reads2 sequences. The third processor 1910 determines whether Reads1 and Reads2 sequences include UMI sequence identifiers. When both Reads1 and Reads2 sequences include UMI sequence identifiers and the included UMI sequence identifiers are the same, it is determined that the plasmid sequence corresponding to Reads1 and Reads2 sequences has been identified. When only one of Reads1 and Reads2 sequences includes a UMI sequence identifier, the Hamming distance between the Reads sequence that does not include a UMI sequence identifier and the UMI sequence identifier is determined. When the determined Hamming distance is less than or equal to a, it is determined that the plasmid sequence corresponding to Reads1 and Reads2 sequences has been identified, where a is a natural number less than or equal to 2.

[0331] It should be understood that the third processor 1910 can be a Central Processing Unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0332] The third memory 1920 may include read-only memory and random access memory, and provides instructions and data to the third processor 1910. A portion of the third memory 1920 may also include non-volatile random access memory. For example, the third memory 1920 may also store device type information.

[0333] In addition to the data bus, the third bus system 1930 may also include a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as the third bus system 1930 in Figure 19.

[0334] In implementation, the processing performed by the processing device can be accomplished through integrated logic circuits in the hardware of the third processor 1910 or through software instructions. That is, the method steps of this embodiment can be executed by a hardware processor, or by a combination of hardware and software modules within the processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other storage media. This storage medium is located in the third memory 1920. The third processor 1910 reads information from the third memory 1920 and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, further details are omitted here.

[0335] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the standard plasmid quantitative analysis method as described in any embodiment of this disclosure. The standard plasmid quantitative analysis method driven by executing executable instructions is essentially the same as the standard plasmid quantitative analysis method provided in the above embodiments of this disclosure, and will not be described in detail here.

[0336] In some possible implementations, various aspects of the standard plasmid quantification method provided in this disclosure may also be implemented as a program product comprising program code that, when run on a computer device, causes the computer device to perform the steps in the standard plasmid quantification method according to various exemplary embodiments of this disclosure as described above. For example, the computer device may execute the standard plasmid quantification method described in the embodiments of this disclosure.

[0337] This disclosure also provides a disease subtyping diagnostic apparatus, including a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to perform the steps of the disease subtyping diagnostic method as described in any embodiment of this disclosure based on the instructions stored in the memory.

[0338] As shown in Figure 20, in one example, the disease typing diagnostic device may include: a fourth processor 2010, a fourth memory 2020, a fourth bus system 2030, and a fourth transceiver 2040. The fourth processor 2010, the fourth memory 2020, and the fourth transceiver 2040 are connected through the fourth bus system 2030. The fourth memory 2020 is used to store instructions, and the fourth processor 2010 is used to execute the instructions stored in the fourth memory 2020 to control the fourth transceiver 2040 to transmit and receive signals. Specifically, the fourth transceiver 2040, under the control of the fourth processor 2010, acquires sequencing data of the sample to be tested. The fourth processor 2010 performs data preprocessing on the acquired sequencing data; compares the preprocessed data with the reference gene sequences of multiple preset clone species to obtain the sequence proportion corresponding to each clone species; sorts the multiple sequence proportions from largest to smallest, and labels the multiple sequence proportions in sorting order as the first sequence proportion to the M1th sequence proportion, where M1 is the number of clone species; detects whether the difference between the first sequence proportion and the third sequence proportion, and the difference between the second sequence proportion and the third sequence proportion, exceeds a preset difference threshold; when the differences between the first sequence proportion and the third sequence proportion, and the differences between the second sequence proportion and the third sequence proportion, both exceed the preset difference threshold, the sample to be tested is determined to be an oligoclonal sample; when the difference between the first sequence proportion and the third sequence proportion exceeds the preset difference threshold but the difference between the second sequence proportion and the third sequence proportion does not exceed the preset difference threshold, the sample to be tested is determined to be a monoclonal sample; when the differences between the first sequence proportion and the third sequence proportion, and the differences between the second sequence proportion and the third sequence proportion, both do not exceed the preset difference threshold, the sample to be tested is determined to be a polyclonal sample.

[0339] It should be understood that the fourth processor 2010 can be a Central Processing Unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0340] The fourth memory 2020 may include read-only memory and random access memory, and provides instructions and data to the fourth processor 2010. A portion of the fourth memory 2020 may also include non-volatile random access memory. For example, the fourth memory 2020 may also store device type information.

[0341] In addition to the data bus, the fourth bus system 2030 may also include a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as the fourth bus system 2030 in Figure 20.

[0342] In implementation, the processing performed by the processing device can be accomplished through integrated logic circuits in the hardware of the fourth processor 2010 or through software instructions. That is, the method steps of this embodiment can be executed by the hardware processor, or by a combination of hardware and software modules within the processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other storage media. This storage medium is located in the fourth memory 2020. The fourth processor 2010 reads information from the fourth memory 2020 and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, further details are omitted here.

[0343] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the disease subtyping and diagnosis method as described in any embodiment of this disclosure. The disease subtyping and diagnosis method driven by executing executable instructions is essentially the same as the disease subtyping and diagnosis method provided in the above embodiments of this disclosure, and will not be described in detail here.

[0344] In some possible implementations, various aspects of the disease subtyping diagnosis method provided in this disclosure may also be implemented as a program product comprising program code that, when run on a computer device, causes the computer device to perform the steps in the disease subtyping diagnosis method according to various exemplary embodiments of this disclosure as described above. For example, the computer device may execute the disease subtyping diagnosis method described in the embodiments of this disclosure.

[0345] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0346] It will be understood by those skilled in the art that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all components may be implemented as software executed by a processor, such as a digital signal processor or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0347] It should be noted that the above embodiments or implementation methods are merely exemplary and not restrictive. Therefore, this disclosure is not limited to the content specifically shown and described herein. Various modifications, substitutions, or omissions can be made to the form and details of the implementations without departing from the scope of this disclosure.

Claims

A primer design method, comprising: Acquire multiple reference data, including the TRBV reference gene sequence and the TRBJ reference gene sequence; For each type of reference data, the following operations are performed: align multiple reference gene sequences by site, determine the conservation score list for each site, obtain conservation scores for multiple conservation intervals based on the conservation score list for each site, select K conservation intervals with higher conservation scores (K is a natural number greater than or equal to 1), generate K primer combinations for the K conservation intervals, screen the primers in the K primer combinations, evaluate the screened K primer combinations, and obtain the final primer combinations based on the evaluation results. According to the method of claim 1, wherein, The process of obtaining the conservation scores of multiple conservation intervals based on the conservation score list for each site includes: obtaining multiple conservation intervals using a sliding window method based on a pre-set initial conservation interval [start, end] and a sliding window step size W; and calculating the conservation score of each conservation interval according to the following formula: W i =log 1 / R i C i =W i ×R i , Among them, W i R represents the weight of the i-th site in each conservatism interval. i C represents the percentage of the most prevalent base type at the i-th site in each conserved interval. i Let S be the conservatism score of the i-th position in each conservatism interval, where S is the conservatism score of each conservatism interval, and n is the length of each conservatism interval, and n = end - start + 1. The method according to claim 2, wherein, The step of generating K primer combinations for the K conserved regions includes: for each of the K conserved regions, performing the following operations: determining the possible base types present at each site in the conserved region, wherein the possible base types present at each site are base types whose proportion is greater than or equal to a preset proportion threshold; generating a primer combination based on the possible base types present at each site in the conserved region, wherein the number of primers in the primer combination is m, where m = πm. i m i denoted as the number of possible base types at the i-th site in the conservative interval, ∏ as the product symbol, i being between 1 and n, and n being the length of the conservative interval. A primer design apparatus includes a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to perform the steps of the primer design method as described in any one of claims 1 to 3 based on the instructions stored in the memory. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the primer design method as described in any one of claims 1 to 3. A computer program product includes instructions that, when executed by a computer, perform the primer design method as described in any one of claims 1 to 3. A composition comprising, by the method of any one of claims 1 to 3: an upstream TRB decoy oligonucleotide, said upstream TRB decoy oligonucleotide being derived from at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 97%, or 100% of any one or more sequences in SEQ ID NO: 1-47; and a downstream TRB decoy oligonucleotide, said downstream TRB decoy oligonucleotide being derived from at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 97%, or 100% of any one or more sequences in SEQ ID NO: 48-60. The composition according to claim 7, wherein, The upstream decoy oligonucleotide of the TRB is derived from at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 97%, or 100% of several sequences in SEQ ID NO:1-47; the downstream decoy oligonucleotide of the TRB is derived from at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 97%, or 100% of several sequences in SEQ ID NO:48-60. The composition according to claim 7, wherein, The upstream decoy oligonucleotide of the TRB is derived from at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 97%, or 100% of all sequences in SEQ ID NO:1-47; the downstream decoy oligonucleotide of the TRB is derived from at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 97%, or 100% of all sequences in SEQ ID NO:48-60. The composition according to claim 7, wherein, The upstream decoy oligonucleotide of the TRB is selected from all sequences in SEQ ID NO:1-47; the downstream decoy oligonucleotide of the TRB is selected from all sequences in SEQ ID NO:48-60. The composition according to claim 7, wherein, The upstream decoy oligonucleotide of the TRB further includes a forward adapter primer sequence, and the downstream decoy oligonucleotide of the TRB further includes a reverse adapter primer sequence; the forward adapter primer sequence and the reverse adapter primer sequence are complementary to the adapter primer sequence used for sequencing. The composition according to claim 11, wherein, The adapter primers used for sequencing are selected from one or more of the Nextera series adapters and the TruSeq series adapters. The composition according to claim 11, wherein, The forward adapter primer sequence is shown in SEQ ID NO:61, and the reverse adapter primer sequence is shown in SEQ ID NO:

62. The composition according to claim 11, wherein, The TRB upstream decoy oligonucleotide is selected from one or more of the sequences shown in SEQ ID NO:63-109; and the TRB downstream decoy oligonucleotide is selected from one or more of the sequences shown in SEQ ID NO:110-122. Use of the composition of any one of claims 7 to 14 in amplifying the TRB gene and / or detecting TRB gene rearrangements. A kit comprising the composition of any one of claims 7 to 14. The kit according to claim 16, wherein, The kit also includes: 56 standard plasmids, each containing a unique UMI sequence, such that each standard plasmid can be uniquely identified by its UMI sequence; wherein the UMI sequence is 16 bp in length, the first 4 bp fragment consists of the last 4 bases of the TRB D region sequence, the last 4 bp fragment consists of the first 4 bases of the TRB J region sequence, and the middle 8 bp fragment is a non-human random sequence; each standard plasmid also includes a TRB V region sequence, a TRB D region sequence, and a TRB J region sequence, wherein the TRB V region sequence, the TRB D region sequence, the UMI sequence, and the TRB J region sequence are sequentially linked end-to-end in each standard plasmid. The kit according to claim 17, wherein, The first 4 bp segment of the UMI sequence is selected from GGGC, GGGG, or AGGG, and the last 4 bp segment is selected from TACT, GCAC, TGAA, AGAT, AGGA, ATTC, GATT, CCAT, CAAA, ACCC, GTTC, CCCA, CGGC, TAGC, ATAG, CCAA, GTTG, CGTC, TACG, TCAT, CAAG, CGGA, AATA, TTCT, AAAC, or CCCC. The kit according to claim 17, wherein, The 56 standard quality grains are standard quality grains containing the following sequences respectively: TRBV10-2, TRBD1, and TRBJ2-2; TRBV10-3, TRBD2*01, and TRBJ2-7; TRBV11-1, TRBD*02, and TRBJ2-2; TRBV11-3, TRBD1, and TRBJ2-4; TRBV12-3, TRBD2*01, and TRBJ2-2; TRBV12-5, TRBD*02, and TRBJ1-6; TRBV13, TRBD1, and TRBJ1-3; TRBV14, TRBD2*01, and TRBJ2-6; TRBV15, TRBD*02, and TRBJ1-4; TRBV16 TRBD1 and TRBJ2-1; TRBV18, TRBD2*01 and TRBJ1-6; TRBV19, TRBD*02 and TRBJ2-1; TRBV2, TRBD1 and TRBJ1-6; TRBV2, TRBD2*01 and TRBJ1-3; TRBV20-1, TRBD*02 and TRBJ2-4; TRBV29-1, TRBD1 and TRBJ2-3; TRBV24-1, TRBD2*01 and TRBJ1-4; TRBV25-1, TRBD*02 and TRBJ1-6; TRBV27, TRBD1 and TRBJ2-6; TRBV28, TRBD2*01 and TRBJ1-1 TRBV3-1, TRBD*02 and TRBJ1-5; TRBV3-1, TRBD1 and TRBJ1-4; TRBV30, TRBD2*01 and TRBJ2-3; TRBV4-1, TRBD*02 and TRBJ1-1; TRBV4-3, TRBD1 and TRBJ1-2; TRBV5-1, TRBD2*01 and TRBJ2-5; TRBV5-1, TRBD*02 and TRBJ1-5; TRBV5-4, TRBD1 and TRBJ2-4; TRBV5-5, TRBD2*01 and TRBJ1-1; TRBV5-6, TRBD*02 and TRBJ1-2; TRBV5-8, TRB D1 and TRBJ1-2; TRBV6-1, TRBD2*01 and TRBJ1-3; TRBV6-2, TRBD*02 and TRBJ2-5; TRBV6-4, TRBD1 and TRBJ1-5; TRBV6-6, TRBD2*01 and TRBJ2-7; TRBV6-6, TRBD*02 and TRBJ1-6; TRBV6-8, TRBD1 and TRBJ2-6; TRBV6-9, TRBD2*01 and TRBJ1-1; TRBV7-8, TRBD*02 and TRBJ1-3; TRBV7-3, TRBD1 and TRBJ1-6; TRBV7-4, TRBD2*01 and TRBJ2-1;TRBV7-6, TRBD*02 and TRBJ1-5; TRBV7-7, TRBD1 and TRBJ1-4; TRBV7-9, TRBD2*01 and TRBJ1-2; TRBV7-9, TRBD*02 and TRBJ1-6; TRBV9, TRBD1 and TRBJ2-7; TRBV7-8, TRBD2*01 and TRBJ1-3; TRBV7-2, TRBD*02 and TRBJ1-6; TRBV7-2, TRBD TRBV11-2, TRBD2*01 and TRBJ2-3; TRBV11-3, TRBD*02 and TRBJ2-2; TRBV5-1, TRBD1 and TRBJ2-3; TRBV19, TRBD2*01 and TRBJ2-5; TRBV30, TRBD*02 and TRBJ2-7; TRBV20, TRBD1 and TRBJ2-6; or TRBV13, TRBD2*01 and TRBJ1-1. The kit according to claim 19, wherein, Mixing the 56 standard quality particles in an equimolar ratio yields a uniformity standard; mixing one or more of the 56 standard quality particles in a high proportion and the other standard quality particles in a low proportion yields an experimental standard for simulating monoclonal or polyclonal experiments. Use of the kit according to any one of claims 16 to 20 in evaluating the amplification efficiency of multiple primers used to amplify the TRB gene. Use of the kit according to any one of claims 16 to 20 in the detection of TRB gene rearrangements. The kit according to any one of claims 16 to 20, wherein, The molar ratios of the decoy oligonucleotides shown in SEQ ID NO:1-60 are as follows: 2:2:1.5:1.5:1.5:2:2.5:1.5:1:2:1.5:1.5:1:1.5:1.5:1:2:1:2:1.5:1:2:1.5:1:1:1:0.5:1.5:1:0.5:1:1:2:1:1:1:1:1:1:1.5:0.5:1:2:0.5:0.5:1:1.5:1.5:0.5:0.5:2:1.5:0.5:0.5:1:1:1:1.5:2:1:

1. A disease typing diagnostic method for detecting TRB gene rearrangements includes the following steps: 1) Obtain the genomic DNA of the sample to be tested; 2) Perform PCR on the genomic DNA obtained in step 1) using the composition of any one of claims 7 to 14 to obtain PCR products; 3) Sequencing the PCR products obtained in step 2) and analyzing the sequencing results to determine whether the TRB gene rearrangement in the sample is monoclonal or polyclonal. The method according to claim 24, wherein, It also includes, in step 1), incorporating the uniformity standard as defined in claim 20 into the genomic DNA. The method according to claim 24, wherein, The sequencing results were analyzed using the following steps: Preprocessing the sequencing data; comparing the preprocessed data with reference gene sequences of multiple preset clonal species to obtain the sequence proportion for each clonal species, where the reference gene sequences include the TRBV and TRBJ reference gene sequences; sorting the multiple sequence proportions from largest to smallest, and labeling them as the first sequence proportion to the M1 sequence proportion, where M1 is the number of clonal species; detecting whether the differences between the first and third sequence proportions and between the second and third sequence proportions exceed preset difference thresholds; when both the differences between the first and third sequence proportions and between the second and third sequence proportions exceed preset difference thresholds, the sample to be tested is determined to be an oligoclonal sample; when the difference between the first and third sequence proportions exceeds the preset difference threshold but the difference between the second and third sequence proportions does not exceed the preset difference threshold, the sample to be tested is determined to be a monoclonal sample; when neither the differences between the first and third sequence proportions nor the differences between the second and third sequence proportions exceed the preset difference threshold, the sample to be tested is determined to be a polyclonal sample. The method according to claim 26, wherein, Detecting whether the difference between the proportion of the first sequence and the proportion of the third sequence exceeds a preset difference threshold includes: detecting whether the ratio of the proportion of the first sequence to the proportion of the third sequence is greater than or equal to X; when the ratio of the proportion of the first sequence to the proportion of the third sequence is greater than or equal to X, it is determined that the difference between the proportion of the first sequence and the proportion of the third sequence exceeds the preset difference threshold; when the ratio of the proportion of the first sequence to the proportion of the third sequence is less than X, it is determined that the difference between the proportion of the first sequence and the proportion of the third sequence does not exceed the preset difference threshold, where X is the preset difference threshold and X is greater than or equal to 2; detecting whether the difference between the proportion of the second sequence and the proportion of the third sequence exceeds a preset difference threshold includes: detecting whether the ratio of the proportion of the second sequence to the proportion of the third sequence is greater than or equal to X; when the ratio of the proportion of the second sequence to the proportion of the third sequence is greater than or equal to X, it is determined that the difference between the proportion of the second sequence and the proportion of the third sequence exceeds the preset difference threshold; when the ratio of the proportion of the second sequence to the proportion of the third sequence is less than X, it is determined that the difference between the proportion of the second sequence and the proportion of the third sequence does not exceed the preset difference threshold. A disease typing diagnostic apparatus includes a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to perform the steps of the disease typing diagnostic method as described in any one of claims 24 to 27 based on the instructions stored in the memory. A computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the disease subtyping diagnostic method as described in any one of claims 24 to 27. A computer program product includes instructions that, when executed by a computer, perform the disease subtyping diagnostic method as described in any one of claims 24 to 27.