Methods, apparatus, and equipment for identifying the source primers of nonspecific amplification sequences
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-27
- Publication Date
- 2026-08-14
AI Technical Summary
[0003]但由于参与链组成的V、D、J在基因组上以基因簇的形式存在,各基因家族数目众多,因此多克隆重排差异大,且容易带来大量的非特异性扩增
[0026]上述说明仅是本公开技术方案的概述,为了能够更清楚了解本公开的技术手段,而可依照说明书的内容予以实施,并且为了让本公开的上述和其它目的、特征和优点能够更明显易懂,以下特举本公开的具体实施方式。
Smart Images

Figure CN117501371B_ABST
Abstract
Description
Technical Field
[0001] This disclosure belongs to the field of gene detection technology, and specifically relates to a method, apparatus, and equipment for identifying the source primers of a non-specific amplification sequence. Background Technology
[0002] In related technologies, the use of NGS technology to identify lymphoma requires multiplex amplification, high-throughput sequencing, and data analysis of DNA through upstream experiments. Generally, multiplex amplification, high-throughput sequencing, and data analysis are performed on chains such as B cell receptor IGH and IGK or T cell receptor TCRB and TCRD to identify the polyclonal rearrangement of lymphocytes.
[0003] However, since the V, D, and J genes involved in the sequencing exist as gene clusters on the genome, and each gene family is numerous, multiple clonal rearrangements vary greatly and are prone to causing a large amount of non-specific amplification. Currently, the target fragment accounts for less than 50% of multiplex amplification sequencing data, and conventional amplification and analysis do not take into account the low data efficiency caused by non-specific amplification. Summary of the Invention
[0004] This disclosure provides a method, apparatus, and equipment for identifying the source primers of a nonspecific amplification sequence.
[0005] This disclosure provides a method for identifying the source primers of a nonspecific amplification sequence, the method comprising: The amplified sequence data of the target gene fragment obtained by primer amplification, the source gene sequence data of the source gene to which the target gene fragment belongs, and the primer sequence data used in the primer amplification process are obtained. The amplified sequence data is compared with the source gene sequence data, and the amplified sequence data that does not match the source gene sequence data is regarded as non-specific amplified sequence data. The nonspecific amplification sequence data is compared with the primer sequence data, and the primers whose primer sequence data matches the nonspecific amplification sequence data are used as the amplification source primers for the nonspecific amplification sequence.
[0006] Optionally, after the step of comparing the amplified sequence data with the source gene sequence data and using the amplified sequence data that does not match the source gene sequence data as non-specific amplified sequence data, the method further includes: The non-specific amplified sequence data is compared with the reference genome sequence data to obtain the comparison results; The location information of the source gene of the non-specific amplified sequence on the genome is determined based on the alignment results.
[0007] Optionally, the step of determining the location information of the source gene of the non-specific amplified sequence on the genome based on the alignment result includes: Based on the distribution location of the alignment results on the reference genome sequence data, at least one of the following is statistically analyzed: the genomic origin, sequence location on the genome, and sequence characteristics of the non-specific amplified sequence on the reference genome.
[0008] Optionally, when the target gene fragment is an immune gene fragment, the gene sequence data includes: overlapping sequence data from paired-end sequencing and non-overlapping sequence data from paired-end sequencing. The step of obtaining the amplified sequence data of the target gene fragment by primer amplification includes: Obtain the offline data obtained by primer amplification of gene fragments; The first gene fragment and the second gene fragment in the sequencing data whose overlapping sequence length is greater than or equal to the first sequence length threshold and whose overlapping sequence length is greater than or equal to the second sequence length threshold are overlapped to obtain overlapping sequence data of paired-end sequencing. Furthermore, amplified sequence data in the sequencing data where the length of the overlapping sequence is less than the first sequence length threshold, or the length of the overlapping sequence is less than the second sequence length threshold, are regarded as non-overlapping sequence data of paired-end sequencing.
[0009] Optionally, the source gene sequence data includes: V gene family sequence data, D gene family sequence data, and J gene family sequence data; The step of comparing the amplified sequence data with the source gene sequence data, and treating the amplified sequence data that does not match the source gene sequence data as non-specific amplified sequence data, includes: The overlapping sequence data and the non-overlapping sequence data of the paired-end sequencing are respectively aligned to the V gene family sequence data, the D gene family sequence data, and the J gene family sequence data to obtain the consistency alignment value of the overlapping sequence data and the non-overlapping sequence data of the paired-end sequencing. The sum of the lengths of the overlapping sequence data from the paired-end sequencing data, where the alignment values with the V gene family sequence data, the D gene family sequence data, and the J gene family sequence data are greater than or equal to the alignment value threshold, is taken as the alignment length of the overlapping sequence data from the paired-end sequencing data. Overlapping paired-end sequencing data with an alignment length less than the alignment length threshold are used as non-specific amplification sequence data. And among the non-overlapping paired-end sequencing sequence data, those sequence data whose consistency alignment values with the V gene family sequence data, the D gene family sequence data, and the J gene family sequence data are all less than the consistency alignment value threshold are used as non-specific amplification sequence data.
[0010] Optionally, after the step of comparing the amplified sequence data with the source gene sequence data and using the amplified sequence data that does not match the source gene sequence data as non-specific amplified sequence data, the method further includes: Redundancy removal is performed on the non-specific amplified sequence data to remove redundant sequence data, wherein the redundant sequence data is sequence data in which the proportion of repetitive bases is greater than or equal to a proportion threshold.
[0011] Optionally, before the step of comparing the amplified sequence data and the source gene sequence data, and using the amplified sequence data that does not match the source gene sequence data as non-specific amplified sequence data, the method further includes: Remove low-quality sequence data from the amplified sequence data.
[0012] Optionally, the step of removing low-quality sequence data from the amplified sequence data includes: After removing the adapter, the adapter sequence data with a sequence end length greater than or equal to the end length threshold is removed, and the sequence data with an average sequence quality value less than the quality value threshold is also removed.
[0013] Optionally, the step of removing low-quality sequence data from the amplified sequence data includes: The low-quality segments in the amplified sequence data whose quality values are less than a quality value threshold are removed, and the low-quality sequence data in the amplified sequence data whose sequence length is less than a third sequence length threshold are also removed.
[0014] This disclosure provides an apparatus for identifying the source primers of a nonspecific amplified sequence, the apparatus comprising: The acquisition module is configured to acquire the amplified sequence data of the amplified gene obtained by primer amplification of the target gene fragment, the source gene sequence data of the source gene to which the target gene fragment belongs, and the primer sequence data used in the primer amplification process. The alignment module is configured to align the amplified sequence data with the source gene sequence data, and to treat amplified sequence data that does not match the source gene sequence data as non-specific amplified sequence data. The nonspecific amplification sequence data is compared with the primer sequence data, and the primers whose primer sequence data matches the nonspecific amplification sequence data are used as the amplification source primers for the nonspecific amplification sequence.
[0015] Optionally, the comparison module is further configured to: The non-specific amplified sequence data is compared with the reference genome sequence data to obtain the comparison results; The location information of the source gene of the non-specific amplified sequence on the genome is determined based on the alignment results.
[0016] Optionally, the comparison module is further configured to: Based on the distribution location of the alignment results on the reference genome sequence data, at least one of the following is statistically analyzed: the genomic origin, sequence location on the genome, and sequence characteristics of the non-specific amplified sequence on the reference genome.
[0017] Optionally, when the target gene fragment is an immune gene fragment, the gene sequence data includes: overlapping sequence data from paired-end sequencing and non-overlapping sequence data from paired-end sequencing. Optionally, the acquisition module is further configured to: Obtain the offline data obtained by primer amplification of gene fragments; The first gene fragment and the second gene fragment in the sequencing data whose overlapping sequence length is greater than or equal to the first sequence length threshold and whose overlapping sequence length is greater than or equal to the second sequence length threshold are overlapped to obtain overlapping sequence data of paired-end sequencing. Furthermore, amplified sequence data in the sequencing data where the length of the overlapping sequence is less than the first sequence length threshold, or the length of the overlapping sequence is less than the second sequence length threshold, are regarded as non-overlapping sequence data of paired-end sequencing.
[0018] Optionally, the source gene sequence data includes: V gene family sequence data, D gene family sequence data, and J gene family sequence data; The comparison module is further configured as follows: The overlapping sequence data and the non-overlapping sequence data of the paired-end sequencing are respectively aligned to the V gene family sequence data, the D gene family sequence data, and the J gene family sequence data to obtain the consistency alignment value of the overlapping sequence data and the non-overlapping sequence data of the paired-end sequencing. The sum of the lengths of the overlapping sequence data from the paired-end sequencing data, where the alignment values with the V gene family sequence data, the D gene family sequence data, and the J gene family sequence data are greater than or equal to the alignment value threshold, is taken as the alignment length of the overlapping sequence data from the paired-end sequencing data. Overlapping paired-end sequencing data with an alignment length less than the alignment length threshold are used as non-specific amplification sequence data. And among the non-overlapping paired-end sequencing sequence data, those sequence data whose consistency alignment values with the V gene family sequence data, the D gene family sequence data, and the J gene family sequence data are all less than the consistency alignment value threshold are used as non-specific amplification sequence data.
[0019] Optionally, the comparison module is further configured to: Redundancy removal is performed on the non-specific amplified sequence data to remove redundant sequence data, wherein the redundant sequence data is sequence data in which the proportion of repetitive bases is greater than or equal to a proportion threshold.
[0020] Optionally, the acquisition module is further configured to: Remove low-quality sequence data from the amplified sequence data.
[0021] Optionally, the acquisition module is further configured to: The adapter sequence data in the amplified sequence data after excision of the adapter are those with a sequence end length greater than or equal to the end length threshold, and the sequence data in the amplified sequence data with a sequence average quality value less than the quality value threshold are removed.
[0022] Optionally, the acquisition module is further configured to: The low-quality segments in the amplified sequence data whose quality values are less than a quality value threshold are removed, and the low-quality sequence data in the amplified sequence data whose sequence length is less than a third sequence length threshold are also removed.
[0023] Some embodiments of this disclosure provide a computing processing device, including: Memory containing computer-readable code; One or more processors, when the computer-readable code is executed by the one or more processors, the computing processing device performs the source primer identification method for the nonspecific amplification sequence as described above.
[0024] Some embodiments of this disclosure provide a computer program, including computer-readable code, which, when run on a computing processing device, causes the computing processing device to perform the source primer identification method for the nonspecific amplification sequence as described above.
[0025] Some embodiments of this disclosure provide a non-transient computer-readable medium storing a method for identifying the source primers of a nonspecific amplification sequence as described above.
[0026] The above description is merely an overview of the technical solution disclosed herein. In order to better understand the technical means of this disclosure and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this disclosure more apparent and understandable, specific embodiments of this disclosure are described below. Attached Figure Description
[0027] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 The schematic diagram illustrates a flowchart of a method for identifying the source primers of a nonspecific amplification sequence provided in some embodiments of this disclosure; Figure 2 This schematically illustrates one of the flowcharts for another method of identifying the source primers of a nonspecific amplification sequence provided in some embodiments of this disclosure; Figure 3 The second schematic diagram illustrates a flowchart of another method for identifying the source primers of a nonspecific amplification sequence provided in some embodiments of this disclosure. Figure 4 The third schematic diagram illustrates a flowchart of another method for identifying the source primers of a nonspecific amplification sequence provided in some embodiments of this disclosure; Figure 5 This illustration schematically shows one of the effect diagrams of a method for identifying the source primers of a nonspecific amplification sequence provided in some embodiments of this disclosure; Figure 6 The diagram illustrates, for example, the effect of a method for identifying the source primers of a nonspecific amplification sequence provided in some embodiments of this disclosure. Figure 7 The schematic diagram illustrates the structure of a source primer identification device for a nonspecific amplification sequence provided in some embodiments of this disclosure; Figure 8A block diagram schematically illustrates a computing processing apparatus for performing methods according to some embodiments of the present disclosure; Figure 9 A storage unit for holding or carrying program code implementing methods according to some embodiments of the present disclosure is illustrated schematically. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0030] Lymphoma is a group of monoclonal proliferative diseases originating from the lymphohematopoietic system. Over the past 30 years, its incidence has increased at a rate of 3% to 5% annually, and its global incidence has approximately doubled. In my country, there are 100,000 new cases of lymphoma each year, with an annual growth rate exceeding 5%, making it one of the fastest-growing common malignant tumors. According to the 2020 Global Cancer Statistics Report, there are approximately 600,000 new cases of lymphoma annually, accounting for 55% of all hematologic malignancies. Accurate diagnosis and classification of lymphoma are crucial to its treatment and prognosis. Molecular genetic characteristics can supplement information not provided by routine pathological examinations, becoming an important means of subtype differentiation.
[0031] Next-generation sequencing (NGS) technology, also known as high-throughput sequencing, is characterized by high throughput and high resolution. It can simultaneously read the sequences of hundreds of thousands to millions of DNA molecules, providing rich genetic information while significantly reducing sequencing costs and shortening sequencing time. With the application of NGS technology, molecular diagnostics is gradually playing a role in the precision diagnosis of diseases such as lymphoma, helping clinicians to better diagnose lymphoma, select treatment plans, determine prognosis, and detect residual microfocal lesions.
[0032] NGS technology has the following advantages when used for lymphoma detection: 1) High sensitivity, with a sensitivity of up to 10. 61) It has 100 times higher sensitivity than traditional flow cytometry. 2) Personalized testing: Due to the large individual differences in the immune genome, VDJ rearrangements can be identified through sequencing analysis. 3) Tracking new clones: As the disease progresses and medication is used, clone evolution can occur in individuals. New clones can be tracked to provide patients with more accurate test results. 4) Prognosis assessment: The prognosis can generally be judged by the supermutation of IGHV. When the IGHV mutation rate is greater than 2%, the prognosis is considered good. Accurate mutation assessment results can guide clinicians to adopt personalized treatment plans for patients.
[0033] The NGS technique for identifying lymphomas requires multiplex DNA amplification via upstream experiments, typically involving multiplex amplification, high-throughput sequencing, and data analysis. This usually involves multiplex amplification of B-cell receptor IGH and IGK chains, or T-cell receptor TCRB and TCRD chains, to identify clonal rearrangements in lymphocytes. However, because the V, D, and J gene pairs exist as gene clusters in the source genes, with numerous gene families and significant rearrangement differences, a large amount of non-specific amplification occurs. Currently, the target fragment accounts for less than 50% of the sequence data. Conventional amplification and analysis do not consider the low data validity caused by non-specific amplification. To optimize primer amplification and improve data validity, it is necessary to analyze the non-specific amplification results to identify the problems caused by primer amplification and provide direction for subsequent primer optimization.
[0034] Figure 1 The schematic diagram illustrates a flowchart of a method for identifying the source primers of a nonspecific amplification sequence provided in this disclosure, the method comprising: Step 101: Obtain the amplified sequence data of the target gene fragment obtained by primer amplification, the source gene sequence data of the source gene to which the target gene fragment belongs, and the primer sequence data used in the primer amplification process.
[0035] It should be noted that the target gene fragment is a characteristic gene fragment in the DNA (Deoxyribonucleic acid) or RNA (ribonucleic acid) of an organism that has undergone multiplex amplification using primers in upstream experiments. The amplified sequence data is obtained by high-throughput sequencing of the target gene fragment. The source gene sequence data is obtained by high-throughput sequencing of the source gene from which the target gene fragment originates. This data can be directly extracted from the source gene database by sequencing the source gene from which the primer amplified target strand recombinant. Primers are short segments of single-stranded DNA or RNA that act as the starting point for the extension of each polynucleotide chain during nucleic acid synthesis reactions. Since primers are pre-designed, primer genes can be amplified to construct a primer database, from which primer sequence data can be directly extracted during use.
[0036] Optionally, the execution subject of some embodiments of this disclosure may be a server or terminal that analyzes amplified sequence data. The server or terminal may be an electronic device with data processing and data transmission functions, such as a server, personal computer, tablet, or laptop. The solution of this disclosure will be described in detail below using a server as the execution subject as an example. Of course, the execution subject of some embodiments of this disclosure may also be other types of electronic devices, which can be set according to actual needs and are not limited here.
[0037] In this embodiment of the disclosure, in the upstream experiment, the target gene fragment can be amplified multiple times using PCR technology based on primers designed for the target gene fragment to obtain the amplified gene, and the amplified gene and the source gene from which the target gene fragment originates can be sequenced using high-throughput sequencing to obtain the amplified sequence data of the amplified gene and the source gene sequence data of the source gene.
[0038] In practical applications, operators can input amplified sequence data into the server, which will automatically extract source gene sequence data from the source gene database and primer sequence data from the primer database to trigger the execution of the steps of the non-specific amplified sequence source primer identification method provided in this disclosure to identify the source primers of non-specific amplified sequences other than the expected specific amplified gene in the amplified gene.
[0039] It is understandable that multiplex amplification is usually performed by amplifying a target gene fragment using primers. The gene that the primers are expected to amplify can be called a specific amplified gene, while the gene that is not induced to amplify by the primers can be called a non-specific amplified sequence. Non-specific amplified sequences not only interfere with the identification and analysis of subsequent sequence data, but also consume a lot of experimental resources, greatly reducing the effectiveness of gene-specific amplification.
[0040] Step 102: Compare the amplified sequence data with the source gene sequence data, and use the amplified sequence data that does not match the source gene sequence data as non-specific amplified sequence data.
[0041] In this embodiment, since the target gene fragment from which the amplified sequence data originates is the source gene, the specifically amplified gene obtained by amplifying the target gene fragment can be aligned to the source gene sequence data. After alignment, the amplified sequence data with a higher alignment length to the source gene sequence data can be considered as the specifically amplified sequence data, and the one with a lower alignment length can be considered as the non-specific amplified sequence data, for further analysis of the source primers of the non-specific amplified sequence data.
[0042] Step 103: Compare the non-specific amplification sequence data with the primer sequence data, and use the primers with primer sequence data that match the non-specific amplification sequence data as the amplification source primers for the non-specific amplification sequence.
[0043] In this embodiment, although the non-specific amplified sequence cannot be aligned to the source gene of the target gene fragment, it is obtained through primer-induced recombination amplification. Therefore, the non-specific amplified sequence can be aligned to the primer sequence data of the primers used. It is understood that amplified sequence data obtained through primer-induced recombination amplification can theoretically be aligned to primer sequence data. This disclosure utilizes this characteristic by comparing the non-specific amplified sequence data in the amplified sequence data with the primer sequence data to identify the primer sequence data that induced recombination to generate the non-specific amplified sequence. The primers corresponding to the used primer sequence data can then be used as the source primers for amplifying the non-specific amplified sequence.
[0044] In practical applications, after the server analyzes the amplification source primers corresponding to the non-specific amplification sequence data, it can output the primer sequence data of the amplification source primers as well as the location and distribution of the non-specific amplification sequence data. This allows operators to view and analyze the effectiveness of the primers, and then optimize the primer settings based on the output results to improve the amplification effect of the primers.
[0045] This embodiment of the disclosure compares the amplified sequence data of the amplified gene with the source gene sequence data from which the primer amplifies the target gene fragment to screen out non-specific amplified sequence data. Then, by comparing the non-specific amplified sequence data with the primer sequence data of the primers used for amplification, primer sequence data that matches the non-specific amplified sequence data can be screened out, thereby accurately identifying the amplification source primers of the non-specific amplified sequence in the amplified sequence.
[0046] Optionally, refer to Figure 2 After step 102, the method further includes: Step 201: The non-specific amplified sequence data is compared with the reference genome sequence data to obtain the comparison results.
[0047] It should be noted that the reference genome sequence data can be obtained by building a database using the human genome GRCh38 version and formatting it using makeblastdb (a format conversion tool). The formatted reference genome database is named GRCh38-db. Of course, the reference genome sequence data can also be other available human genome sequence data. This is just an example. The specific settings can be made according to actual needs, and there are no restrictions here.
[0048] In this embodiment, considering that nonspecific amplified sequences are obtained through recombination amplification and have low readability of their base distribution, the gene sequence on the source gene cannot be directly identified from the nonspecific amplified sequence data. Therefore, this disclosure compares the nonspecific amplified sequence data with reference genome sequence data, and determines the gene sequence fragment of the reference genome sequence data that matches the nonspecific amplified sequence data based on the comparison results.
[0049] Step 202: Determine the location information of the source gene of the non-specific amplified sequence on the genome based on the alignment results.
[0050] In this embodiment of the disclosure, the server can determine the location information of the source gene sequence of the specific amplified gene on the genome based on the location of the reference gene sequence data that matches the specific amplified sequence data in the reference genome sequence data determined by the comparison results.
[0051] Optionally, step 202 includes: based on the distribution position of the alignment results on the reference genome sequence data, statistically analyzing at least one of the following: genomic origin, sequence position on the genome, and sequence characteristics of the non-specific amplified sequence on the reference genome.
[0052] In this embodiment of the disclosure, non-specific amplified sequence data is aligned to reference genome sequence data. The standard for sequence alignment is that the identity value is greater than or equal to the sequence identity threshold, which is any value between 80% and 95% of the total sequence length, such as 80%, 85%, and 95%. Statistical analysis is performed on the identified sequence source locations. The statistical content includes the optimal genome source for each sequence alignment, the concentration of sequence sources in specific genome locations, or sequence characteristics.
[0053] This embodiment of the present disclosure overcomes the problem that the base sequence of nonspecific amplified sequences is not easy to identify, and improves the accuracy of identifying the location information of the source gene of nonspecific amplified sequences by comparing nonspecific amplified sequence data with reference genome sequence data and determining the location information of the source gene of nonspecific amplified sequences based on the comparison results.
[0054] Optionally, when the target gene fragment is an immune gene fragment, the gene sequence data includes: overlapping paired-end sequencing data and non-overlapping paired-end sequencing data, as shown in the reference. Figure 3 Step 101 includes: Step 1011: Obtain the offline data obtained by primer amplification of the gene fragment.
[0055] In this embodiment of the disclosure, in order to improve the identification accuracy of non-specific amplified sequence data, after the server obtains the offline data of the amplified gene, it can perform an overlap operation on the offline data. The overlap operation refers to combining two sequence data in the offline data according to the base pairing method.
[0056] Step 1012: The first gene fragment and the second gene fragment in the offline data whose overlapping sequence length is greater than or equal to the first sequence length threshold and whose overlapping sequence length is greater than or equal to the second sequence length threshold are overlapped to obtain overlapping sequence data of paired-end sequencing. The amplified sequence data in the offline data whose overlapping sequence length is less than the first sequence length threshold or whose overlapping sequence length is less than the second sequence length threshold are used as non-overlapping sequence data of paired-end sequencing.
[0057] In this embodiment, the first gene fragment R1 and the second gene fragment R2 can be two different paired gene fragments from the original amplified gene sequencing data. The criteria for performing the overlap operation are that the overlap (overlappable sequence) length of R1 and R2 is greater than or equal to a first sequence length threshold, which can be any value between 10bp and 20bp, such as 10bp, 15bp, or 2bp. After overlap, the sequence length is greater than or equal to a second sequence length threshold, which can be any value between 100bp and 150bp, such as 100bp, 125bp, or 150bp. Thus, the sequencing data that meets the overlap criteria can be overlapped to obtain overlapping paired-end sequencing data, while the sequencing data that does not meet the overlap criteria are used as non-overlapping paired-end sequencing data. The server can use both the overlapping and non-overlapping paired-end sequencing data as amplified sequence data for subsequent analysis.
[0058] This embodiment of the invention overlaps amplified sequence data, enabling bidirectional alignment of overlapping sequence data at both ends during subsequent alignment processes. This results in higher accuracy of the alignment results compared to non-overlapping gene sequence data.
[0059] Optionally, the source gene sequence data includes: V gene family sequence data, D gene family sequence data, and J gene family sequence data, as referred to... Figure 3 Step 102 includes: Step 1021: Align the overlapping sequence data and the non-overlapping sequence data of the paired-end sequencing to the V gene family sequence data, the D gene family sequence data, and the J gene family sequence data, respectively, to obtain the consistency alignment value of the overlapping sequence data and the non-overlapping sequence data of the paired-end sequencing.
[0060] Step 1022: The sum of the lengths of the sequence data in the overlapping sequence data of the paired-end sequencing, where the alignment values with the V gene family sequence data, the D gene family sequence data, and the J gene family sequence data are greater than or equal to the alignment value threshold, is taken as the alignment length of the overlapping sequence data of the paired-end sequencing. Step 1023: The paired-end sequencing data with an alignment length less than the alignment length threshold that can overlap are used as non-specific amplification sequence data. Step 1024: Among the non-overlapping sequence data from the paired-end sequencing, the sequence data whose alignment values with the V gene family sequence data, the D gene family sequence data, and the J gene family sequence data are all less than the alignment value threshold are used as non-specific amplification sequence data.
[0061] It should be noted that the V, D, and J gene families are the three main gene families in the reproductive gene sequence of human germ cells. These three gene families are distributed sequentially in the reproductive gene sequence in the order of V, D, and J.
[0062] In this embodiment of the disclosure, for the assembled set (overlapping sequence data from paired-end sequencing), mapped (completely matched data) can be defined as alignment identity ≥ consistency alignment threshold, which can be any value between 80 and 90, such as 80, 85, or 90. The length of the paired-end sequencing overlapping sequence data aligned to the V, D, or J gene families is greater than or equal to any value between 80% and 90% of the total sequence length, such as 80%, 85%, or 90%. Mapped (partially matched data) can be defined as alignments where the identity value is greater than or equal to the alignment threshold. The length of the overlapping sequence data from paired-end sequencing aligned to the V, D, and J gene families is greater than or equal to any value between 10% and 20% of the total sequence length, such as 10%, 15%, or 20%. Unmapped (unmatched data) can be defined as alignments where the identity value is greater than or equal to the alignment threshold. The length of the overlapping sequence data from paired-end sequencing aligned to the V, D, and J gene families is less than the length of the alignment threshold, which can be any value between 10% and 20% of the total sequence length, such as 10%, 15%, or 20%.
[0063] For the unassembled set (non-overlapping sequence data from paired-end sequencing), mapped is defined as alignment identity ≥ consistency threshold, where R1 and R2 can align to gene family V and gene family J respectively, or one of R1 and R2 can align to both gene families V and J simultaneously; partially mapped is defined as alignment identity ≥ consistency threshold, where R1 or R2 can align to one of the gene families V, D, or J of the immune gene; unmapped is defined as alignment identity ≥ consistency threshold, where R1 or R2 cannot align to any of the gene families V, D, or J of the immune gene, meaning that the consistency value of the non-overlapping sequence data from paired-end sequencing with any gene family V, D, or J is less than the consistency threshold.
[0064] Extract and store the data from the three datasets (mapped, partially mapped, and unmapped) in Fasta format, and perform data volume statistics for each dataset.
[0065] It is worth noting that the reason why only the non-overlapping sequence data of paired-end sequencing was compared with the V and J gene families of immune genes is that the non-overlapping sequence data of paired-end sequencing cannot be compared with the D gene family, which is located in the middle part of the immune gene sequence, because the middle part of the non-overlapping sequence data of paired-end sequencing is not overlapped.
[0066] In this embodiment of the disclosure, only the unmapped data after comparison with the immune gene, that is, the gene sequence data that is not matched, is used as non-specific amplified sequence data.
[0067] In this embodiment of the disclosure, referring to the above description, the sequence length of gene sequences whose alignment consistency value with immune gene sequence data is greater than the alignment value threshold and whose alignment length is less than the alignment length threshold can be used as non-specific amplification sequence data for subsequent analysis of the amplification primer source of non-specific amplification sequences.
[0068] Optionally, refer to Figure 4 Before step 102, the method further includes: Step 104: Remove low-quality sequence data from the amplified sequence data.
[0069] This embodiment of the disclosure removes low-quality sequence data from the amplified sequence, reducing the interference of low-quality sequence data on the analysis process and the processing resources required for analysis, thereby improving the efficiency and accuracy of data analysis.
[0070] Optionally, refer to Figure 4 After step 102, the method further includes: Step 105: Perform redundancy removal processing on the non-specific amplification sequence data, removing redundant sequence data from the non-specific amplification sequence data, wherein the redundant sequence data is sequence data in which the proportion of repetitive bases in the sequence is greater than or equal to a proportion threshold.
[0071] In this embodiment, non-specific amplified sequence data can be deredundanted. Redundant sequences are defined as those with a similarity greater than or equal to a similarity threshold identified by global sequence identification, calculated by dividing the number of identical bases in the sequence by the full length of the shorter sequence. Similarity identification tools such as CD-HIT can be used for deredundancy and sequence clustering. Clustering information is stored for each deredundanted non-specific amplified sequence, including the clustered sequence, its cluster information, and the number of sequences within each non-specific amplified sequence. This reduces the memory footprint of the non-specific amplified sequence data and the time required for subsequent analysis.
[0072] Optionally, step 202 includes: removing adapter sequence data in the amplified sequence data after removing the adapter whose sequence end length is greater than or equal to an end length threshold, and removing sequence data in the amplified sequence data whose average sequence quality value is less than a quality value threshold.
[0073] In the embodiments disclosed herein, the adapter sequence is a short, synthetically produced DNA fragment containing an enzyme cleavage site that can match blunt or sticky ends.
[0074] In this embodiment, adapter and low-quality data filtering is performed on the amplified sequence data. Based on the adapter sequence from the sequencing platform, adapter sequences are identified in the sequencing data and then removed. The removal criterion is that the adapter sequence data has an end length greater than or equal to a threshold value, which can be any value between 3bp and 6bp, such as 3bp, 4bp, or 6bp. Then, low-quality sequences are filtered and removed. The sequence filtering criterion is that the average quality value of the sequence is less than a quality average threshold value. The amount of data before and after adapter sequence removal is counted. This quality average threshold value can be any value between 20 and 25, such as 21, 23, or 25.
[0075] Optionally, step 202 includes: removing low-quality segments in the amplified sequence data whose quality value is less than a quality value threshold, and removing low-quality sequence data in the amplified sequence data whose sequence length is less than a third sequence length threshold after the low-quality segments have been removed.
[0076] In this embodiment, low-quality regions in the amplified sequence are removed. The removal criterion is that the quality value is less than a quality value threshold, which can be any value between 20 and 25, such as 20, 23, or 25. Then, the sequence length is filtered. The filtering criterion is that the length of the sequence data after removing the low-quality sequence is less than a third sequence length threshold, which can be any value between 40 bp and 50 bp, such as 40 bp, 45 bp, or 50 bp. The amount of data after quality value and length filtering is then calculated.
[0077] This disclosure reduces the memory usage of non-specific amplified sequence data and the time required for subsequent analysis by filtering low-quality sequence data before analyzing the amplified sequence data.
[0078] Exemplary examples are provided in this disclosure, using sequence data of a multiplex amplified sequence TCRD strand as amplified sequence data: S1 preprocesses the sequencing data, including removing adapters and low-quality sequences, and classifying reads for overlap.
[0079] The TCRD strands of 22 samples were amplified and sequenced using the PE150 sequencing strategy (paired sequencing, read length 150 bp). The sample sequencing data volume ranged from 0.1M to 0.3M reads, as shown in Table 1.
[0080]
[0081] Table 1 Where “_1” and “_2” represent R1 and R2 of the paired Read, respectively; Read Number: number of sample Reads; BaseCount: number of sample bases.
[0082] The header sequences were filtered, with the removal criterion being headers at least 3 bp in length at the end of the sequence. Low-quality sequences were also filtered and removed, with the filtering criterion being an average quality value less than 25. Low-quality regions within the sequence were removed, with the removal criterion being a quality value less than 25. Sequence length was filtered, with the filtering criterion being a length less than 50 bp after removing low-quality sequences. The data volume after filtering the header sequence content, quality value, and length can be found in Table 2.
[0083]
[0084] Table 2 Among them, adapter count: the number of adapter sequences in the sequence; Qual filter base (Ratio): the number of bases after quality filtering (the filtering ratio); Length filter count (Ratio): the number of sequences filtered out because R1 or R2 does not meet the length requirement (the filtering ratio); Left base: the amount of data of remaining bases; Overlapped: the proportion of sequences that can overlap.
[0085] The sequencing data were overlapped at the R1 and R2 ends. The overlap criteria were that the overlap length between R1 and R2 was greater than or equal to 10 bp, and the length of the sequence after overlap was greater than or equal to 100 bp. The overlapable sequences were stored in the "assembled" set, and the non-overlapping sequences were stored in the "unassembled" set. The data volume statistics of the two datasets can be found in Table 3.
[0086]
[0087] Table 3 Wherein, Read Number: the number of sequences in the dataset; Read Base: the number of bases in the dataset.
[0088] Step S2: Establish the immunomics database, genome database, and primer database.
[0089] An immune database was constructed using germ cell sequences derived from recombinant TCRD from the IMGT database. TCRD has two D genes: TRDD1, TRDD2, and TRDD3. However, the TRDD gene is less than 20 bp in length, with a maximum length of only 13 bp, and therefore was not used for constructing the immune database. The TRDV and TRDJ genes were used to construct the immune database, designated as TRD-V and TRD-J, respectively.
[0090] A database was built using the GRCh38 version of the human genome, formatted using makeblastdb, and named GRCh38-db.
[0091] Step S3: Create a database of amplification primer pairs (upstream primer, downstream primer), format it using makeblastdb, and name the database Primer-db after formatting.
[0092] Align the sequencing data to the established immune database. Identify the alignment results. The identification criteria are described in S3. Statistical analysis of mapped and unmapped values for the "Assembled" and "Unassembled" sets is shown in Table 3.
[0093] S4 involved sequence deduplication and source identification of the unmapped dataset, identifying sequences originating from specific locations on the genome. The deduplication method is shown in S4.1, and statistics were compiled for the deduplicated "clstr" dataset (see Table 4). The results show that the "Assembled" unmapped dataset exhibits significant clustering, achieving at least 70% deduplication on 22 samples. The top 1 and top 5 clstr sequences contain a large proportion of these sequences, demonstrating representativeness. The "Unassembled" set also shows significant clustering, achieving at least 65% deduplication. However, the top 1 and top 5 clstr sequences contain a small proportion of these sequences, lacking representativeness.
[0094]
[0095] Table 4 Wherein, Before clstr: the data volume statistics before redundancy removal; After clstr: the data volume statistics after redundancy removal; clstr Ratio (%): the percentage of redundancy removal; Top1: the number of sequences in the largest clstr; Top5: the number of sequences in the top 5 clstrs.
[0096] Aligning the "clstr" dataset to the GRCh38-db database, the sequences can be identified as having a unique genomic origin. Statistical analysis was performed on the identified sequence origin locations, including the optimal genomic origin for each sequence and the concentration of sequence origins in specific genomic locations. The alignment results are good, accurately determining the genomic location. Table 5 shows the top 5 most frequently located genomic locations for each "clstr" dataset. Therefore, when optimizing the multiplex primer design for this group, attention should be paid to minimizing the genomic locations mentioned in this table.
[0097]
[0098] Table 5 Step S5: Identify the sources of amplification primers in the "clstr" dataset and the Top 5 primers from the "clstr" dataset. Align the "clstr" dataset and the "clstr" dataset to Primer-db for analyzing the primer sources in the datasets (see...). Figure 5 , Figure 6 (Table 6). The top 5 sequences in each "clstr" dataset were aligned to Primer-db for primer source analysis. For the "Assembled" "clstr" dataset, V4, V5, and V6 were primer sequence sources. For the "Ussembled" clstr dataset, V1 was a relatively high primer sequence source. For the top 5 data in the "clstr" dataset, J1 and J3 were relatively high primer sequence sources. Figure 5 This is the result of aligning the "clstr" dataset of "Assembled" to the Primer-db database. The horizontal axis represents the sample number, and the vertical axis represents the number of times the dataset was aligned. Figure 6 This is the result of aligning the "clstr" dataset ("Unassembled") to the Primer-db database. The horizontal axis represents the sample number, and the vertical axis represents the number of times the dataset was aligned.
[0099]
[0100] Table 6 Of course, the above is just an illustrative description, and the specific data can be set according to actual needs. No restrictions are imposed here.
[0101] Figure 7 A schematic diagram of a source primer identification device 30 for a nonspecific amplification sequence provided in this disclosure is shown, the device comprising: The acquisition module 301 is configured to acquire the amplified sequence data of the amplified gene obtained by primer amplification of the target gene fragment, the source gene sequence data of the source gene to which the target gene fragment belongs, and the primer sequence data used in the primer amplification process. The comparison module 302 is configured to compare the amplified sequence data with the source gene sequence data, and to treat the amplified sequence data that does not match the source gene sequence data as non-specific amplified sequence data. The nonspecific amplification sequence data is compared with the primer sequence data, and the primers whose primer sequence data matches the nonspecific amplification sequence data are used as the amplification source primers for the nonspecific amplification sequence.
[0102] Optionally, the comparison module 302 is further configured to: The non-specific amplified sequence data is compared with the reference genome sequence data to obtain the comparison results; The location information of the source gene of the non-specific amplified sequence on the genome is determined based on the alignment results.
[0103] Optionally, the comparison module 302 is further configured to: Based on the distribution location of the alignment results on the reference genome sequence data, at least one of the following is statistically analyzed: the genomic origin, sequence location on the genome, and sequence characteristics of the non-specific amplified sequence on the reference genome.
[0104] Optionally, when the target gene fragment is an immune gene fragment, the gene sequence data includes: overlapping sequence data from paired-end sequencing and non-overlapping sequence data from paired-end sequencing. Optionally, the acquisition module 301 is further configured to: Obtain the offline data obtained by primer amplification of gene fragments; The first gene fragment and the second gene fragment in the sequencing data whose overlapping sequence length is greater than or equal to the first sequence length threshold and whose overlapping sequence length is greater than or equal to the second sequence length threshold are overlapped to obtain overlapping sequence data of paired-end sequencing. Furthermore, amplified sequence data in the sequencing data where the length of the overlapping sequence is less than the first sequence length threshold, or the length of the overlapping sequence is less than the second sequence length threshold, are regarded as non-overlapping sequence data of paired-end sequencing.
[0105] Optionally, the source gene sequence data includes: V gene family sequence data, D gene family sequence data, and J gene family sequence data; The comparison module 302 is further configured to: The overlapping sequence data and the non-overlapping sequence data of the paired-end sequencing are respectively aligned to the V gene family sequence data, the D gene family sequence data, and the J gene family sequence data to obtain the consistency alignment value of the overlapping sequence data and the non-overlapping sequence data of the paired-end sequencing. The sum of the lengths of the overlapping sequence data from the paired-end sequencing data, where the alignment values with the V gene family sequence data, the D gene family sequence data, and the J gene family sequence data are greater than or equal to the alignment value threshold, is taken as the alignment length of the overlapping sequence data from the paired-end sequencing data. Overlapping paired-end sequencing data with an alignment length less than the alignment length threshold are used as non-specific amplification sequence data. And among the non-overlapping paired-end sequencing sequence data, those sequence data whose consistency alignment values with the V gene family sequence data, the D gene family sequence data, and the J gene family sequence data are all less than the consistency alignment value threshold are used as non-specific amplification sequence data.
[0106] Optionally, the comparison module 302 is further configured to: Redundancy removal is performed on the non-specific amplified sequence data to remove redundant sequence data, wherein the redundant sequence data is sequence data in which the proportion of repetitive bases is greater than or equal to a proportion threshold.
[0107] Optionally, the acquisition module 301 is further configured to: Remove low-quality sequence data from the amplified sequence data.
[0108] Optionally, the acquisition module 301 is further configured to: The adapter sequence data in the amplified sequence data after excision of the adapter are those with a sequence end length greater than or equal to the end length threshold, and the sequence data in the amplified sequence data with a sequence average quality value less than the quality value threshold are removed.
[0109] Optionally, the acquisition module 301 is further configured to: The low-quality segments in the amplified sequence data whose quality values are less than a quality value threshold are removed, and the low-quality sequence data in the amplified sequence data whose sequence length is less than a third sequence length threshold are also removed.
[0110] This embodiment of the disclosure compares the amplified sequence data of the amplified gene with the source gene sequence data from which the primer amplifies the target gene fragment to screen out non-specific amplified sequence data. Then, by comparing the non-specific amplified sequence data with the primer sequence data of the primers used for amplification, primer sequence data that matches the non-specific amplified sequence data can be screened out, thereby accurately identifying the amplification source primers of the non-specific amplified sequence in the amplified sequence.
[0111] The various component embodiments of this disclosure can be implemented in hardware, or as software modules running on one or more processors, or a combination thereof. Those skilled in the art will understand that microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some or all of the components in the computing processing device according to embodiments of this disclosure. This disclosure can also be implemented as a device or apparatus program (e.g., a computer program and computer program product) for performing some or all of the methods described herein. Such an implementation of this disclosure can be stored on a non-transient computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.
[0112] For example, Figure 8 A computing processing apparatus is shown that can implement the methods according to this disclosure. This computing processing apparatus conventionally includes a processor 410 and a computer program product or non-transitory computer-readable medium in the form of a memory 420. The memory 420 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. The memory 420 has a storage space 430 for program code 431 for performing any of the method steps described above. For example, the storage space 430 for program code may include various program codes 431 respectively for implementing the various steps in the methods described above. These program codes can be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, compact discs (CDs), memory cards, or floppy disks. Such computer program products are typically as described in the references. Figure 9 The portable or fixed storage unit is described above. This storage unit may have the same characteristics as... Figure 8The memory 420 in the computing processing device is similarly arranged as storage segments, storage spaces, etc. Program code can be compressed, for example, in an appropriate form. Typically, the storage unit includes computer-readable code 431', that is, code that can be read by a processor such as 410, which, when run by the computing processing device, causes the computing processing device to perform the various steps in the methods described above.
[0113] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0114] The terms "an embodiment," "embodiment," or "one or more embodiments" as used herein mean that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment of this disclosure. Furthermore, please note that the examples of the phrase "in one embodiment" do not necessarily all refer to the same embodiment.
[0115] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of this disclosure may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0116] In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. This disclosure can be implemented by means of hardware comprising a plurality of different elements and by means of a suitably programmed computer. In a unit claim enumerating a plurality of means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words may be interpreted as names.
[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this disclosure, and are not intended to limit them. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure.
Claims
1. A method for identifying the source primers of a non-specific amplification sequence, characterized in that, The method includes: The amplified sequence data of the target gene fragment obtained by primer amplification, the source gene sequence data of the source gene to which the target gene fragment belongs, and the primer sequence data used in the primer amplification process are obtained. The amplified sequence data is compared with the source gene sequence data, and the amplified sequence data that does not match the source gene sequence data is regarded as non-specific amplified sequence data. The nonspecific amplification sequence data is compared with the primer sequence data, and the primers that match the primer sequence data are used as the amplification source primers for the nonspecific amplification sequence; when the target gene fragment is an immune gene fragment, the gene sequence data includes: paired-end sequencing overlapping sequence data and paired-end sequencing non-overlapping sequence data. The source gene sequence data includes: V gene family sequence data, D gene family sequence data, and J gene family sequence data; The step of comparing the amplified sequence data with the source gene sequence data, and treating the amplified sequence data that does not match the source gene sequence data as non-specific amplified sequence data, includes: The overlapping sequence data and the non-overlapping sequence data of the paired-end sequencing are respectively aligned to the V gene family sequence data, the D gene family sequence data, and the J gene family sequence data to obtain the consistency alignment value of the overlapping sequence data and the non-overlapping sequence data of the paired-end sequencing. The sum of the lengths of the overlapping sequence data from the paired-end sequencing data, where the alignment values with the V gene family sequence data, the D gene family sequence data, and the J gene family sequence data are greater than or equal to the alignment value threshold, is taken as the alignment length of the overlapping sequence data from the paired-end sequencing data. Overlapping paired-end sequencing data with an alignment length less than the alignment length threshold are used as non-specific amplification sequence data. And among the non-overlapping paired-end sequencing sequence data, those sequence data whose consistency alignment values with the V gene family sequence data, the D gene family sequence data, and the J gene family sequence data are all less than the consistency alignment value threshold are used as non-specific amplification sequence data.
2. The method according to claim 1, characterized in that; The step of obtaining the amplified sequence data of the target gene fragment by primer amplification includes: Obtain the offline data obtained by primer amplification of gene fragments; The first gene fragment and the second gene fragment in the sequencing data whose overlapping sequence length is greater than or equal to the first sequence length threshold and whose overlapping sequence length is greater than or equal to the second sequence length threshold are overlapped to obtain overlapping sequence data of paired-end sequencing. Furthermore, amplified sequence data in the sequencing data where the length of the overlapping sequence is less than the first sequence length threshold, or the length of the overlapping sequence is less than the second sequence length threshold, are regarded as non-overlapping sequence data of paired-end sequencing.
3. The method according to claim 1, characterized in that, After the step of comparing the amplified sequence data with the source gene sequence data and treating the amplified sequence data that does not match the source gene sequence data as non-specific amplified sequence data, the method further includes: The non-specific amplified sequence data is compared with the reference genome sequence data to obtain the comparison results; The location information of the source gene of the non-specific amplified sequence on the genome is determined based on the alignment results.
4. The method according to claim 3, characterized in that, The step of determining the location information of the source gene of the non-specific amplified sequence on the genome based on the alignment result includes: Based on the distribution location of the alignment results on the reference genome sequence data, at least one of the following is statistically analyzed: the genomic origin, sequence location on the genome, and sequence characteristics of the non-specific amplified sequence on the reference genome.
5. The method according to claim 1, characterized in that, After the step of comparing the amplified sequence data with the source gene sequence data and treating the amplified sequence data that does not match the source gene sequence data as non-specific amplified sequence data, the method further includes: Redundancy removal is performed on the nonspecific amplified sequence data to remove redundant sequence data, wherein the redundant sequence data is sequence data in which the proportion of repetitive bases in the sequence is greater than or equal to a proportion threshold.
6. The method according to claim 1, characterized in that, Before the step of comparing the amplified sequence data with the source gene sequence data and treating the amplified sequence data that does not match the source gene sequence data as non-specific amplified sequence data, the method further includes: Remove low-quality sequence data from the amplified sequence data.
7. The method according to claim 6, characterized in that, The step of removing low-quality sequence data from the amplified sequence data includes: The adapter sequence data in the amplified sequence data after excision of the adapter are those with a sequence end length greater than or equal to the end length threshold, and the sequence data in the amplified sequence data with a sequence average quality value less than the quality value threshold are removed.
8. The method according to claim 6, characterized in that, The step of removing low-quality sequence data from the amplified sequence data includes: The low-quality segments in the amplified sequence data whose quality values are less than a quality value threshold are removed, and the low-quality sequence data whose sequence length is less than a third sequence length threshold is removed from the amplified sequence data whose low-quality segments have been removed.
9. A device for identifying the source primers of a non-specific amplified sequence, characterized in that, The device includes: The acquisition module is configured to acquire the amplified sequence data of the amplified gene obtained by primer amplification of the target gene fragment, the source gene sequence data of the source gene to which the target gene fragment belongs, and the primer sequence data used in the primer amplification process. The alignment module is configured to align the amplified sequence data with the source gene sequence data, and to treat amplified sequence data that does not match the source gene sequence data as non-specific amplified sequence data. The nonspecific amplification sequence data is compared with the primer sequence data, and the primers that match the primer sequence data of the nonspecific amplification sequence are used as the amplification source primers of the nonspecific amplification sequence. When the target gene fragment is an immune gene fragment, the gene sequence data includes: overlapping sequence data from paired-end sequencing and non-overlapping sequence data from paired-end sequencing. The source gene sequence data includes: V gene family sequence data, D gene family sequence data, and J gene family sequence data; The step of comparing the amplified sequence data with the source gene sequence data, and treating the amplified sequence data that does not match the source gene sequence data as non-specific amplified sequence data, includes: The overlapping sequence data and the non-overlapping sequence data of the paired-end sequencing are respectively aligned to the V gene family sequence data, the D gene family sequence data, and the J gene family sequence data to obtain the consistency alignment value of the overlapping sequence data and the non-overlapping sequence data of the paired-end sequencing. The sum of the lengths of the overlapping sequence data from the paired-end sequencing data, where the alignment values with the V gene family sequence data, the D gene family sequence data, and the J gene family sequence data are greater than or equal to the alignment value threshold, is taken as the alignment length of the overlapping sequence data from the paired-end sequencing data. Overlapping paired-end sequencing data with an alignment length less than the alignment length threshold are used as non-specific amplification sequence data. And among the non-overlapping paired-end sequencing sequence data, those sequence data whose consistency alignment values with the V gene family sequence data, the D gene family sequence data, and the J gene family sequence data are all less than the consistency alignment value threshold are used as non-specific amplification sequence data.
10. A computing processing device, characterized in that, include: Memory containing computer-readable code; One or more processors, when the computer-readable code is executed by the one or more processors, the computing processing device performs the source primer identification method for the nonspecific amplification sequence as described in any one of claims 1-8.
11. A non-transient computer-readable medium, characterized in that, It contains a computer program for identifying the source primers of the nonspecific amplification sequence as described in any one of claims 1-8.