Pig mitochondrial whole genome targeted amplification primer set and its application in nanopore sequencing and SNP provenance tracing
By designing a specific shingled long fragment amplification primer set and a random forest model, the problems of enrichment of the whole genome of pig mitochondria and stratified origin traceability in existing technologies have been solved, realizing efficient and accurate traceability of low-quality samples. It is applicable to pig tissues, pork sausage products and pig fecal swabs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SCIENCE & TECHNOLOGY RESEARCH CENTER OF CHINA CUSTOMS
- Filing Date
- 2026-06-08
- Publication Date
- 2026-07-21
AI Technical Summary
Existing technologies cannot efficiently enrich the whole genome of pig mitochondria in low-quality, complex matrix samples, and it is difficult to perform high-accuracy stratified origin tracing through nanopore sequencing platforms, especially when distinguishing closely related geographical populations.
We designed specific swathenium long fragment amplification primer sets and trained them with a random forest model. We adopted a two-step prediction method to first distinguish between Asia or Europe, America and Oceania, and then further distinguish between pigs or pig products from specific countries or regions.
It achieves efficient enrichment and high-accuracy stratified origin traceability for low-quality samples, and is applicable to pig tissue, pork sausage products and pig fecal swabs. It overcomes the shortcomings of traditional detection methods and provides an efficient and stable integrated molecular detection and discrimination technology for the traceability of imported pig products at ports.
Smart Images

Figure CN122427918A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of animal-derived product gene origin traceability technology, specifically to a porcine mitochondrial whole genome targeted amplification primer set and its application in nanopore sequencing and SNP origin traceability. Background Technology
[0002] As the first line of defense for national biosecurity, ports of entry play a crucial role in preventing the introduction of major animal diseases from abroad, combating illegal cross-border trade in animal products, and safeguarding my country's livestock industry and public health. Currently, port animal product traceability faces practical challenges such as complex sample sources, short testing timelines, and the ineffectiveness of traditional traceability methods. Conventional physical markers are easily damaged or tampered with during cross-border circulation; identification methods based on morphology and physicochemical indicators are insufficient to distinguish between pigs and pig products of similar geographical origin, failing to meet the regulatory needs for rapid and accurate traceability at ports. Therefore, there is an urgent need to establish a high-throughput, high-specificity traceability technology system based on molecular genetic characteristics.
[0003] The mitochondrial genome, with its maternally inherited characteristics, strong sequence conservation, and ease of extraction from trace / degraded samples, is an ideal molecular marker for tracing the origin of animal products intercepted at ports. Existing sequencing methods based on the mitochondrial genome mainly include Sanger sequencing and metagenomic sequencing. Sanger sequencing obtains mitochondrial genome fragments through nucleic acid amplification, followed by sequencing to determine the base sequence. This method offers high accuracy and relatively simple operation, but suffers from short read lengths—the length that can be accurately measured in a single reaction is typically below 800 bp. Obtaining the mitochondrial genome sequence requires assembling multiple short fragments, leading to a high assembly error rate. Metagenomic sequencing offers high throughput, allowing for simultaneous sequencing of large numbers of samples at a lower cost. However, this method requires comprehensive sequencing of the entire genome before screening for mitochondrial genome sequences, resulting in a large amount of redundant sequencing data. Furthermore, its sensitivity is lower than targeted sequencing, leading to lower success rates for whole-genome sequencing of low-quality samples such as processed products and fecal swabs. Nanopore sequencing technology, which has emerged in recent years, has provided new possibilities for on-site sequencing at ports of entry due to its advantages of real-time sequencing, portability, and low requirements for sample quality. However, using this technology alone still faces the problems of the low proportion of mitochondrial genome in total DNA and low efficiency of direct sequencing.
[0004] Furthermore, existing mitochondrial-based origin tracing technologies largely rely on D-loop regions or a limited number of SNP sites. Due to the limited genetic information in these regions, they cannot effectively distinguish closely related populations. More importantly, there are currently no reports on establishing a hierarchical modeling method for pig-origin products, progressing from continent to country. Single multi-classification models are prone to feature aliasing when dealing with SNP characteristics of closely related geographical populations, leading to decreased accuracy. In summary, there is a lack of a technical solution that can efficiently enrich the entire pig mitochondrial genome in low-quality, complex matrix samples, overcome sequence variation interference, be compatible with nanopore sequencing platforms, and combine mitochondrial genome SNP information for high-accuracy hierarchical origin tracing.
[0005] Therefore, there is an urgent need to develop an efficient and easy-to-operate integrated technical solution to address the problem that existing technologies cannot meet the actual needs of port quarantine and rapid and accurate traceability. Summary of the Invention
[0006] In view of this, the purpose of this application is to provide a set of primers for targeted amplification of the entire porcine mitochondrial genome and its application in nanopore sequencing and SNP origin tracing. This application addresses the problems of uneven amplification coverage, primer interference, and low amplification efficiency in complex matrices that easily occur in the porcine mitochondrial circular genome. A set of specific shingled long-fragment amplification primers, as shown in SEQ ID NO. 1~12, has promising applications in porcine mitochondrial genome amplification, library construction, and sequencing origin tracing. Furthermore, this application utilizes the obtained whole genome sequence for SNP extraction and screening, and trains it using a random forest model to obtain four ensemble training models. A two-step prediction method is employed: first, it predicts whether the pig or its product originates from Asia, Europe, America, or Oceania; then, based on this distinction, it further predicts whether the pig or its product originates from China, Russia, the United States, or other Asian countries after Europe, America, or Oceania. This method is applicable to pig tissues, pork sausage products, and pig fecal swabs. It overcomes the shortcomings of traditional detection methods, such as low throughput, difficulty in distinguishing closely related origins, and low sequencing success rate for low-quality samples. It provides an efficient, stable, and intelligent integrated molecular detection and discrimination technology for tracing the origin of imported pig products at ports.
[0007] To achieve the above objectives, this application provides the following technical solution:
[0008] In a first aspect, this application provides a primer set for amplifying the whole genome of porcine mitochondria, the primer set comprising nucleotide sequences as shown in SEQ ID NO. 1~12.
[0009] In some preferred embodiments, the primer set is divided into two primer pools as follows:
[0010] Primer pool A consists of the sequences shown in SEQ ID NO.1, SEQ ID NO.2, SEQ ID NO.5, SEQ ID NO.6, SEQ ID NO.9, and SEQ ID NO.10;
[0011] Primer pool B consists of the sequences shown in SEQ ID NO.3, SEQ ID NO.4, SEQ ID NO.7, SEQ ID NO.8, SEQ ID NO.11, and SEQ ID NO.12.
[0012] Secondly, this application provides a kit for amplifying the whole genome of porcine mitochondria, which includes the primer set described in the first aspect.
[0013] In some embodiments, the kit further includes multiplex PCR reaction solution, high-fidelity DNA polymerase, and nuclease-free water.
[0014] Thirdly, this application provides a PCR method for amplifying the entire genome of porcine mitochondria, comprising the following steps:
[0015] Collect samples and extract nucleic acids from them;
[0016] The primers described in the first aspect are divided into primer pool A and primer pool B. Using the extracted nucleic acid as a template, multiplex PCR amplification reactions are performed using primer pool A and primer pool B respectively to obtain two sets of amplification products.
[0017] The two sets of amplification products were mixed to obtain a mixed amplification product of the entire porcine mitochondrial genome.
[0018] In some embodiments, the reaction system for the multiplex PCR amplification is as follows: 10 μL of 5×PrimeSTAR Buffer, 4 μL of dNTP Mixture (2.5 mM each), 0.5 μL of PrimeSTAR HS DNA Polymerase (2.5 U / μL), 0.2–1.4 μmol / L of primer pool A or B, 5 μL of nucleic acid template, and nuclease-free water to a final volume of 50 μL.
[0019] In some embodiments, the reaction conditions for the multiplex PCR amplification are: 95°C for 2 min; 98°C for 10 s, 55–50.5°C for 15 s (with a 0.5°C drop per cycle), 68°C for 3 min 40 s, 10 cycles; 98°C for 10 s, 50°C for 15 s, 68°C for 3 min 40 s, 22 cycles; 72°C for 5 min.
[0020] In some embodiments, the sample is derived from at least one of pig tissue, blood, meat products, processed pig by-products, or fecal swabs.
[0021] In some implementations, the method for extracting nucleic acids from the sample is the magnetic bead method or the column method.
[0022] Fourthly, this application provides a targeted sequencing method for the entire porcine mitochondrial genome, comprising the following steps:
[0023] (1) The whole genome of porcine mitochondria is targeted amplified using any of the PCR methods described in the third aspect to obtain mixed amplification products;
[0024] (2) Construct sequencing libraries from the amplification products of step (1);
[0025] (3) Perform nanopore sequencing on the sequencing library from step (2);
[0026] (4) Sequence assembly of sequencing results: The pig mitochondrial reference genome is split into corresponding fragments according to the primer amplification fragments as assembly reference sequences. Each fragment contig is extracted from the sequencing results, and after removing the primers, it is assembled again to obtain the complete full-length sequence of pig mitochondria.
[0027] Fifthly, this application provides a method for constructing a porcine mitochondrial genome-wide SNP-based origin tracing model, comprising the following steps:
[0028] (1) Dataset construction: 226 known mitochondrial whole genome sequences from different countries were screened from GenBank, and 45 mitochondrial whole genome sequences from different countries were determined using the targeted sequencing method described in the fourth aspect, forming 271 mitochondrial whole genome datasets;
[0029] (2) SNP extraction: After aligning the sequence from step (1) using the Maft software, a total of 1867 SNP sites were extracted using R language and the Biostrings package, and the default sites were replaced with NA; the SNP sites cover the entire coding and control regions of the porcine mitochondrial genome;
[0030] (3) Model training: A four-layer progressive hierarchical modeling was designed using the random forest method. First, the continents were coarsely classified, and then the regions were subdivided, gradually narrowing the discrimination range and reducing the interference of multi-class feature aliasing. Specifically, the first step was to merge the data of Asia and Europe and Oceania respectively, and use all effective mitochondrial SNP sites as the quantitative features as the model input. No feature reduction was made, and the global species genetic differentiation information was retained for binary classification training. On this basis, the sequences of the United States and Russia in the data of Europe and Oceania were further classified and trained, and the sequences of China in the data of Asia were classified and trained.
[0031] The dataset cleaning and feature preprocessing were performed using R language and the dplyr package; the SNP genotype data reading and result output were performed using the readxl and writexl packages; the model training, prediction, and accuracy evaluation were performed using the randomForest package; and the model performance visualization plot was performed using ggplot2. The interactive source tracing analysis system was built using the shiny, DT, and shinythemes packages.
[0032] Compared with the prior art, this application has at least the following beneficial effects:
[0033] 1. This application designs a shingled long-fragment amplification primer set as shown in SEQ ID NO. 1~12, targeting the characteristics of the porcine mitochondrial circular genome. By dividing the primers into A and B dual-primer pools for partitioned multiplex PCR amplification, adjacent amplified fragments are seamlessly connected end-to-end, which can completely cover the entire porcine mitochondrial circular genome, effectively avoiding the problems of primer non-specific fusion amplification and uneven amplification coverage. Only 0.1 ng of nucleic acid template is required to achieve 100% full genome coverage sequencing, with a sequencing depth of over 30×.
[0034] 2. The technical solution of this application is not only applicable to raw samples such as fresh pig tissue and muscle, but can also effectively cover a variety of common and complex matrix samples at ports, such as meat products, processed pig by-products, and fecal swabs. It can efficiently enrich the whole mitochondrial genome from complex samples and avoid interference from non-specific stray bands.
[0035] 3. This application integrates public databases and measured sample sequences to construct a multi-regional pig mitochondrial whole genome dataset containing 271 sequences. Abandoning the limitations of single multi-classification models, a four-layer progressive hierarchical training method using random forest is adopted. First, coarse classification is performed by continent, followed by refined discrimination of products from the United States, Russia, and China. All valid SNP loci (1867) are retained as model input throughout the process, fully preserving the genetic differentiation characteristics of pig populations from different geographical locations. This effectively reduces the interference of SNP feature aliasing from closely related populations and maintains good stability even with small sample sizes and high noise levels.
[0036] 4. Relying on the mature R language toolchain, sequence alignment, SNP extraction, data cleaning, model training, accuracy verification, and visualization analysis are completed. An interactive traceability analysis system is also built to realize an integrated technical approach for the entire process from gene sequencing and site mining to intelligent origin identification, providing an efficient and stable molecular detection and intelligent identification solution for the traceability of imported pig products at ports. Attached Figure Description
[0037] Figure 1 This is a schematic diagram of a primer amplification strategy for targeted amplification and enrichment of the entire porcine mitochondrial genome using a shingled long-fragment primer set provided in an embodiment of this application.
[0038] Figure 2 The technical route for using primer sets provided in this application for whole genome sequencing of pig mitochondria and origin tracing on a nanopore sequencing platform.
[0039] Figure 3 This is an agarose gel electrophoresis image of the PCR product in Example 1 of this application.
[0040] Figure 4 This is a schematic diagram of the assembly of the whole pig mitochondrial genome provided in Example 1 of this application.
[0041] Figure 5 This is an agarose gel electrophoresis image of the PCR product in Example 2 of this application.
[0042] Figure 6 Accuracy chart of the development of a limited country origin traceability model provided for embodiments of this application.
[0043] Figure 7 The results of the traceability test for known sample origins provided in the embodiments of this application are shown. Detailed Implementation
[0044] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0045] The materials used in the following embodiments are not limited to those listed below, and other similar materials may be used instead. Unless otherwise specified, the instruments shall be used under conventional conditions or as recommended by the manufacturer. Those skilled in the art should have relevant knowledge of the use of conventional materials and instruments.
[0046] In this application, unless the context clearly indicates otherwise, the terms “including,” “comprising,” “containing,” “having,” etc., shall be understood as open-ended and mean “including but not limited to.”
[0047] In this application, several different terms are used to describe primer sets, such as “primer pools A and B” and “Swap primer set”. Those skilled in the art should understand that these terms are synonymous and refer to all or part of the primers SEQ ID NO.1~12 provided in this application.
[0048] To better understand this teaching and without limiting its scope, all figures and other numerical values used in the specification and claims to express quantities, percentages, or proportions should, in all cases, be understood to be modified by the term "about." Therefore, unless otherwise stated, the numerical parameters set forth in the following specification and appended claims are approximate values that may be adjusted according to the desired performance. At a minimum, each numerical parameter should be interpreted based on the reported significant figures and by applying common rounding techniques.
[0049] To address the problem of amplifying the entire porcine mitochondrial genome in the aforementioned prior art, the inventors of this application designed a 3.5kb amplification fragment in a shingled and end-to-end primer set by comparing the porcine mitochondrial genome sequence. This ensures that the primer set can cover the complete circular genome of porcine mitochondria in different geographical regions. There is an overlap region of about 100 bp between adjacent primer pairs. Two different primer pools were set up for each adjacent primer pair. To ensure sensitivity, the concentration of each primer pair was optimized separately. To ensure the accuracy of SNP sites, a high-fidelity Taq enzyme with 3'→5' activity was used. Then, the two different primer pools were used for amplification and the amplification products were combined. This can effectively avoid non-specific primer fusion amplification and improve amplification efficiency.
[0050] Based on this, embodiments of this application provide a primer set for amplifying the whole genome of porcine mitochondria, which includes nucleotide sequences as shown in SEQ ID NO.1~12.
[0051] In the above primer design rules, the primer length is 22-35 bases, the GC content is 50%-60%, there are no secondary or repetitive structures within the primers, there are no complementary sequences between or within the primers, and the melting temperature (Tm value) difference between the primers is less than 5 ℃.
[0052] In practical applications, the 12 primer sequences are divided into two groups, primer pools A and B, as shown in Table 1, covering the porcine mitochondrial circular genome. The inventors used these two primer groups to amplify the entire porcine mitochondrial genome, thus achieving targeted and efficient enrichment of the entire porcine mitochondrial genome and subsequent origin tracing analysis. The sequences and grouping information of the 12 primers are shown in Table 1 below:
[0053] Table 1. sHEV whole-genome targeted capture enrichment primer set
[0054]
[0055] In the primer sequences above, A represents adenine, G represents guanine, C represents cytosine, and T represents thymine; the degenerate base R represents A / G, and K represents G / T.
[0056] Based on this, this application provides a kit for amplifying the whole genome of porcine mitochondria, which includes a primer set of the nucleotide sequences shown in SEQ ID NO. 1~12 above.
[0057] In some embodiments, the kit further includes multiplex PCR reaction solution, high-fidelity enzyme, and nuclease-free water. These reagents are necessary to ensure the enrichment of the porcine mitochondrial whole-genome retargeting PCR amplification reaction; the multiplex PCR reaction solution may include PCR buffer, MgCl2, dNTPs, and high-fidelity Taq DNA polymerase.
[0058] Based on this, this application also provides a PCR method for enriching and amplifying the whole mitochondrial genome of pigs, which includes using the aforementioned primer set as PCR amplification primers.
[0059] In the above PCR method, except for using the primer set of nucleotide sequences shown in SEQ ID NO.1~12 provided in this application as amplification primers, no specific settings are made for other steps in PCR amplification, so as to achieve the amplification purpose.
[0060] In some embodiments, the specific steps of the PCR method are as follows: collect samples and extract nucleic acids from the samples; use the extracted nucleic acids as templates and perform multiplex PCR amplification reactions using the aforementioned primer pools A and B respectively to obtain two sets of amplification products; mix the two sets of amplification products in equal volumes to obtain the final PCR mixed amplification product of the whole pig mitochondrial genome.
[0061] In some embodiments, the method for extracting nucleic acids from samples is the magnetic bead method or the column method.
[0062] In some embodiments, the sample is derived from at least one of pig tissue, blood, meat products, processed pig by-products, and fecal swabs.
[0063] In some embodiments, the PCR reaction system is as follows:
[0064] 5×PrimeSTAR Buffer 10 μL, dNTP Mixture (2.5 mM each) 4 μL, PrimeSTAR HSDNA Polymerase (2.5 U / μL) 0.5 μL, Primer Pool A or B 0.2–1.4 μmol / L (final concentration of each primer), nucleic acid template 5 μL, add nuclease-free water to 50 μL;
[0065] The reaction conditions for the PCR were as follows:
[0066] 95°C / 2min;
[0067] 98 ℃ / 10 s, 55~50.5 ℃ / 15 s (0.5℃ drop per cycle), 68 ℃ / 3 min 40 s, 10 cycles;
[0068] 98 ℃ / 10 s, 50 ℃ / 15 s, 68 ℃ / 3 min 40 s, 22 cycles;
[0069] 72 ℃ / 5 min.
[0070] Based on this, this application also provides a method for amplifying the entire porcine mitochondrial genome and for limited country-of-origin tracing, including:
[0071] (1) The porcine mitochondrial genome was targeted and amplified using the aforementioned PCR method to obtain mixed amplification products;
[0072] (2) Construct a sequencing library from the amplification products of (1);
[0073] (3) Perform nanopore sequencing on the sequencing library from (2);
[0074] (4) Based on the sequencing results of (3), perform limited country origin traceability analysis.
[0075] Figure 2 This diagram illustrates the technical roadmap for amplifying and sequencing the entire porcine mitochondrial genome using primer pools A and B on a nanopore sequencing platform. It shows the process from nucleic acid extraction, through PCR amplification, library construction and sequencing, to data analysis, ultimately achieving the detection and analysis of the entire porcine mitochondrial genome. The following section combines... Figure 2 The above-mentioned methods for tracing the origin of products will be explained in detail below:
[0076] (1) Pig genomic DNA extraction: Nucleic acid is isolated from the sample to prepare for subsequent testing;
[0077] (2) Construction of multiplex PCR reaction system: The extracted nucleic acid is mixed with "multiplex PCR Mix" (containing enzymes, primers, etc.) and primer pool (different primers such as AB), and the target gene fragment of the pig mitochondrial genome is specifically amplified by "multiplex PCR target capture" instrument;
[0078] (3) Processing of amplification products: After mixing, the amplification products are purified and quantified.
[0079] (4) Library construction and sequencing: The library is constructed by adding native barcodes and native adapters, and finally the sequence is determined by a sequencer.
[0080] (5) Sequence assembly: The pig mitochondrial reference genome was decomposed into 6 corresponding fragments as assembly reference sequences. Sequences were extracted from the sequencing results. The 6 contigs of each sample were removed from the primers and then assembled again to obtain the complete full-length sequence of pig mitochondria.
[0081] (6) Construct datasets. 226 known mitochondrial whole genome sequences from different countries were screened from GenBank. The above sequencing methods were used to determine 45 mitochondrial whole genome sequences from different countries, forming 271 mitochondrial whole genome datasets.
[0082] (7) SNP extraction: After aligning the above sequences using Maft software, a total of 1867 SNP sites were extracted using R language and Biostrings package and written to an Excel file, with the default sites replaced by NA.
[0083] (8) Model training: Due to data skewness and application purpose, a single multi-class model was not adopted. Instead, a four-layer progressive hierarchical model was designed using the random forest method. First, the continents were coarsely classified, and then the regions were subdivided, gradually narrowing the discrimination range and reducing the interference of multi-class feature aliasing. In the first step, the data of Asia and Europe and Oceania were merged respectively. All effective mitochondrial SNP sites were used as the quantitative features of the model input without feature reduction, and global species genetic differentiation information was retained for binary classification training. On this basis, the sequences of the United States and Russia in the Europe and Oceania data were further classified and trained, and the sequences of China in the Asian data were classified and trained. Dataset cleaning and feature preprocessing were performed using R language and the dplyr package. The reading and output of SNP genotype data were performed using the readxl and writexl packages. The randomForest package was used for model training, prediction and accuracy evaluation. The model performance visualization plot was performed using ggplot2. The shiny, DT and shinythemes packages were used to build the interactive source tracing analysis system.
[0084] The following are specific examples:
[0085] Example 1: Establishment of a pig mitochondrial whole-genome targeted nanopore sequencing method
[0086] 1. Primer and probe composition for targeted nanopore sequencing of the entire porcine mitochondrial genome.
[0087] In this embodiment, the entire genome of porcine hepatitis E virus was selected as the target region. Based on multiple sequence alignment, a swathenic primer composition was designed. A schematic diagram of the amplified fragment location is shown below. Figure 1The primers are 22–35 bases in length, with a GC content of 50%–60%. They contain no secondary structures or repetitions, no complementary sequences between or within primers, and the melting temperatures (Tm values) between primers differ by less than 5 °C. The designed primer set comprises single-stranded DNA molecules as shown in SEQ ID NO. 1–12. The 12 primers are divided into two groups, A and B. The nucleotide sequences and grouping information of the primers are shown in Table 1 above.
[0088] 2. Methods for targeted amplification and enrichment of the entire porcine mitochondrial genome
[0089] The method for using multiplex PCR for the whole genome of porcine mitochondria is as follows: Sample pretreatment and nucleic acid extraction; using the extracted nucleic acid as a template, primer pools A and B were prepared for the aforementioned two primer sets, and multiplex PCR amplification was performed separately. The PCR products were then mixed, purified, and used for nanopore sequencing. The specific steps include:
[0090] 2.1 Sample Pretreatment and Nucleic Acid Extraction
[0091] For the samples to be tested, pretreatment and nucleic acid extraction should be performed as follows:
[0092] For liquid samples such as fecal swabs and blood, genomic DNA can be extracted directly from the samples. For samples such as pig tissue, meat products, and processed pig by-products, after thorough grinding, prepare a 20% tissue suspension with PBS and bring the volume to 1 mL. Add 50 μL of proteinase K solution (1 mg / mL), then add 50 μL of 10% SDS solution, gently mix by pipetting, and digest in a 55 ℃ water bath for 2 h. Centrifuge at 3000 rpm for 2 min, and extract DNA from the supernatant. This step can be performed using effective DNA extraction methods, or various commercial DNA extraction kits or fully automated nucleic acid extractors and their matching reagents.
[0093] In constructing the detection method, this application embodiment used pork samples from Denmark and extracted genomic DNA using the "Trace Sample Genomic DNA Extraction Kit" (DP316) according to the instructions.
[0094] 2.2 Primer Concentration Optimization
[0095] Using the genomic DNA obtained in step 2.1 as a template, PCR amplification was performed on each primer pair. The primer concentration was dynamically optimized between 0.2 and 1.6 μmol / L, while the concentrations of other reactants remained constant. The reaction system is shown in Table 2 below:
[0096] Table 2 Primer Concentration Optimization
[0097]
[0098] 5×PrimeSTAR Buffer, dNTP Mixture, and PrimeSTAR HS DNA Polymerase were purchased from TaKaRa; primers were synthesized by Shanghai Sangon Biotech Co., Ltd.; nuclease-free water was prepared by double distillation of purified water, followed by the addition of DEPC to a final concentration of 0.1%, and then treated with stirring at 37 °C for 12 h, followed by autoclaving at 1.034 × 10⁵ Pa for 15 minutes.
[0099] The PCR reaction procedure is as follows:
[0100] 95 °C / 2 min;
[0101] 98 ℃ / 10 s, 55~50.5℃ / 15 s (0.5℃ drop per cycle), 68 ℃ / 3 min 40 s, 10 cycles;
[0102] 98 ℃ / 10 s, 50 ℃ / 15 s (0.5 ℃ drop per cycle), 68 ℃ / 3 min 40 s, 22 cycles;
[0103] 72 ℃ / 5 min.
[0104] After multiple rounds of optimization, the optimal concentrations of each primer pair were finally determined. The electrophoresis results of each PCR product after primer optimization are shown in the figure. Figure 2 As can be seen from the figure, each primer pair can efficiently amplify DNA extracted from Danish pork samples. Primer pools A and B were then prepared according to their optimal concentrations, as shown in Table 3.
[0105] Table 3. Preparation methods of primer pools A and B
[0106]
[0107] 2.3 Targeted Multiplex PCR Amplification
[0108] After preparing primer pools A and B according to Table 3, use primer pools A and B to prepare multiplex PCR systems according to Table 4 to obtain two sets of PCR products.
[0109] Table 4 Multiplex PCR amplification reaction system
[0110]
[0111] The PCR reaction procedure is the same as in 2.2.
[0112] See the electrophoresis diagram of the PCR products. Figure 3 As can be seen from the figure, primer pools A and B can effectively amplify the DNA extracted from Danish pork samples. The products were purified with 0.8× magnetic beads and then used for later use.
[0113] 3. Nanopore sequencing of PCR products
[0114] Follow the instructions for nanopore sequencing kits (including the barcode ligation sequencing kit SQK-NBD114.24; the R10 sequencing chip, FLO-MIN114, NANOPORE; and the ligation sequencing multi-sample DNA library preparation kit, BAIYI) and the specific steps are as follows.
[0115] 3.1 Purification of PCR products:
[0116] ① Equilibrate the AMPure XP magnetic beads from the kit to room temperature, vortex to resuspend, add 40 μL of magnetic beads to 50 μL of the mixed RT-PCR product, mix by gently tapping, and incubate at room temperature for 5 minutes; ② After a brief centrifugation, transfer the liquid to a new 1.5 mL centrifuge tube, place it on a magnetic rack, and wait for the liquid to clarify; ③ Remove the supernatant on the magnetic rack, add 180 μL of freshly prepared 80% ethanol along the tube wall to gently wash the magnetic beads, discard the liquid, and repeat the washing once more; ④ After a brief centrifugation, place the centrifuge tube back on the magnetic rack and remove any residual liquid; allow it to air dry for 30 seconds; ⑤ Add 27 μL of nuclease-free water to resuspend the magnetic beads for nucleic acid elution, incubate at room temperature for 2 minutes, briefly centrifuge, and place it on a magnetic rack until the liquid clarifies; take 1 μL and measure the concentration using the QubitdsDNA HS Assay Kit, then take another 25 μL for the next step of end-repair.
[0117] 3.2 End-of-trim
[0118] ① Add 25 μL of the purified RT-PCR product to a 0.2 mL PCR tube; ② Add the corresponding reagent components of the auxiliary library preparation kit according to Table 5 below, gently mix and briefly centrifuge; ③ React at 20 °C for 5 min on a PCR instrument; then at 65 °C for 5 min, remove; ④ Purify using an equal volume of AMPure XP magnetic beads according to step 3.1. After washing the magnetic beads with ethanol in the last step, elute with 15 μL of nuclease-free water; take 1 μL and measure the concentration using the Qubit dsDNA HS Assay Kit, then aspirate 13 μL for the next step of adding the native barcode.
[0119] Table 5. End-of-phase remediation reaction system
[0120]
[0121] 3.3 Add native barcode
[0122] ① Add the corresponding reagent components of the kit according to Table 6 below, gently tap to mix and briefly centrifuge; ② Gently tap to mix about 10 times, briefly centrifuge, and incubate at 18~25 °C for 10 min; ③ Purify using an equal volume of AMPure XP magnetic beads according to step 3.1. In the last step, wash the magnetic beads with ethanol and elute with 15 μL of nuclease-free water; take 1 μL and measure the concentration using the Qubit dsDNA HSAssay Kit, then take 13 μL for the next step of library construction.
[0123] Table 6. Native Barcode System
[0124]
[0125] 3.4 Library Construction
[0126] ① Add purified, barcode-linked DNA to a final volume of 51 μL with nuclease-free water and transfer to a new 1.5 mL centrifuge tube. Follow the preparation method in Table 7 below, ligate the sequencing adapter, and construct the sequencing library. ② Gently tap the tube approximately 10 times to mix, briefly centrifuge, and incubate at 18–25 °C for 10 min. ③ Add 40 μL of AMPure XP magnetic beads, gently tap the tube wall to mix, and incubate at room temperature for 5 min. ④ After brief centrifugation, place the tube on a magnetic rack and allow it to stand until the liquid is clear; then discard the supernatant. ⑤ Remove the centrifuge tube, gently wash the magnetic beads with 150 μL of SFB, briefly centrifuge, and allow it to stand on a magnetic rack until the liquid is clear; then repeat the washing process once more. ⑥ Briefly centrifuge and return the centrifuge tube to the magnetic rack, discard any remaining liquid, and allow it to air dry for 30 seconds. Add 14 μL of EB elution buffer, remove the centrifuge tube from the magnetic rack, gently tap to mix, and incubate at room temperature for 10 minutes. min; place the centrifuge tube on a magnetic rack until the eluent is clear and colorless; transfer the eluent to a new 1.5 mL centrifuge tube, and measure the concentration of 1 μL using the Qubit dsDNA HS Assay Kit.
[0127] Table 7. Configuration of the Library Construction System
[0128]
[0129] 3.5 Third-generation nanopore sequencing
[0130] ① Melt the Sequencing Buffer (SB), Library Beads (LIB), Flow Cell Tether (FCT), and Flow Cell Flush (FCF) in the kit at room temperature, vortex to mix, briefly centrifuge, and keep on ice; ② Add 30 μL of FCT to 1170 μL of FCF and vortex to mix; ③ After equilibrating the sequencing chip at room temperature for 10 min, rotate 90° clockwise to open the Priming port. Take a 1 mL pipette, adjust it to 200 μL, and hold it vertically to the Priming port orifice. Slowly rotate to 230 μL to aspirate air bubbles from the tubing. A small amount of liquid at the tip of the pipette indicates that air bubbles have been expelled; ④ Use a 1 mL pipette to draw 800 μL of FCF containing FCT and push it into the Priming port orifice at a constant speed, ensuring a small amount of liquid remains at the tip. Re-equilibrate the chip at room temperature for 5 min; ⑤ After chip equilibration, gently open the SpotON injection port by turning it outwards; use a 1 mL pipette to draw 200 μL of FCF containing FCT. ⑥ Add μL of FCT-added FCF to the tubing by rotating the pipette clockwise through the Priming port well to rinse the SpotON well. ⑦ Add the reagents according to Table 8 below to prepare the sequencing loading system. ⑧ Mix the loading system by pipetting and adding it dropwise to the SpotON well. ⑨ Close the SpotON well first, then close the Priming port, remove the waste liquid through the waste liquid well, cover with the light-shielding sticker, and load the chip into the GridION for sequencing.
[0131] Table 8. Preparation of Sequencing Sample Loading System
[0132]
[0133] 3.6 Sequencing Data Assembly and Analysis
[0134] The generated sequence data was filtered to remove sequencing adapters and barcodes, retaining full-length sequence fragments longer than 1000 bp to ensure data integrity and applicability. Then, the published porcine mitochondrial reference genome sequence (GenBank ID: NC_000845.1) was split into six fragments as reference assembly sequences according to the designed amplification products. Contigs for assembling the PCR products of the tested samples were extracted from the sequenced sequences, and the sequencing depth and coverage of each fragment were determined. After manually removing primers from the six contigs of the tested samples, the sixth fragment (including the first and last fragments of the published linear genome sequence) was split into two fragments according to the published reference genome sequence (GenBank ID: NC_000845.1). The full-length mitochondrial genome was then assembled using SeqMan software. An assembly reference diagram is shown below. Figure 4 .
[0135] Example 2: Sensitivity evaluation of the porcine mitochondrial whole-genome targeted nanopore sequencing method
[0136] Using the aforementioned Danish pork DNA as the sample to be tested, a sensitivity test was conducted on the porcine mitochondrial targeted nanopore sequencing method. The specific method is as follows:
[0137] Pork DNA was extracted using the "Micro Sample Genomic DNA Extraction Kit" (DP316) according to the manufacturer's instructions. The specific pretreatment process for pork was the same as in Example 1. After extraction, the DNA concentration was determined using a Qubit analyzer, and the result was 18 ng / μL. The genomic DNA was serially diluted to 0.06 ng / μL. Then, using the aforementioned primer pools A and B, targeted PCR amplification was performed on seven groups of DNA templates at different concentrations, with PBS as a blank control. After purification of the targeted PCR mixtures, the nucleic acid concentration was determined using a Qubit analyzer, and sequencing was performed according to the nanopore sequencing method in Example 1. The results are shown in Table 9.
[0138] Table 9 Sequencing sensitivity evaluation
[0139]
[0140] As shown in Table 9, when the DNA concentration was 0.11 ng / μL, multiplex PCR showed obvious amplification bands (see Table 9). Figure 5 All six fragment sequences had 100% 30× coverage. However, when the concentration was reduced to 0.06 ng / μL, the 30× coverage of fragments 1 and 4 did not reach 100%. Therefore, the sensitivity of the porcine mitochondrial targeted genome nanopore sequencing method established in this application was determined to be 0.11 ng / μL.
[0141] Example 3: Targeted sequencing of porcine mitochondrial genomes from imported pork products and swabs
[0142] Whole-genome targeted nanopore sequencing was performed on imported pork products and pigs using the sequencing method described in Example 1. The specific method is as follows:
[0143] 1. Sample collection and DNA extraction
[0144] The samples included a total of 45 samples, including pork imported through ports, intercepted pork sausages, pork products, and pig feces swabs imported through different ports. Detailed information on the samples is shown in Table 10. After pretreatment according to Example 1, pork DNA was extracted using the "Trace Sample Genomic DNA Extraction Kit" (DP316) in accordance with the instructions.
[0145] Table 10. Names and sources of 45 samples
[0146]
[0147] 2. Targeted PCR amplification was performed using the genomic DNA extracted in step 1 as a template.
[0148] Using the aforementioned primer sets A and B, corresponding primer pools were prepared according to Example 1. Multiplex PCR was performed using genomic DNA extracted from the samples in Table 10 as templates, with PBS as a blank control.
[0149] 3. Nanopore sequencing
[0150] The PCR products were used for library construction and nanopore sequencing according to Example 1.
[0151] 4. Results Analysis
[0152] Two hours after nanopore sequencing, data analysis was performed. The generated sequence data was filtered to remove sequencing adapters and barcodes, retaining full-length sequence fragments longer than 1000 bp to ensure data integrity and applicability. Then, the sequences were assembled according to the method described in 3.6 of Example 1 to obtain the mitochondrial whole genome sequences of 45 samples.
[0153] Example 4: Extraction of SNPs from the whole genome of porcine mitochondria, model training and evaluation, and limited country-of-origin traceability analysis.
[0154] 1. Dataset Construction
[0155] We screened 226 known mitochondrial whole genome sequences from different countries from GenBank, and used the above sequencing methods to determine 45 mitochondrial whole genome sequences from different countries, forming 271 mitochondrial whole genome datasets, which were stored in FASTA format.
[0156] 2. Sequence alignment
[0157] Sequence alignment was performed using MAFFT software v7.526 (https: / / mafft.cbrc.jp / alignment / software / ). To ensure that the order of newly added aligned sequences remains unchanged and to facilitate SNP extraction and prediction of new sequences, the following batch file was added to the MAFFT target folder:
[0158] @echo off
[0159] chcp 65001>nul
[0160] echo==============================================
[0161] echo MAFFT preserves the original sequence order during alignment
[0162] echo==============================================
[0163] echo.
[0164] Echo is comparing data, please wait...
[0165] "%~dp0mafft.bat"--inputorder"%~1">"%~dpn1_aligned.fasta"
[0166] echo.
[0167] echo Alignment complete! Output file: %~n1_aligned.fasta
[0168] echo Sequence order = Input order
[0169] pause
[0170] The aligned sequences are saved in FASTA format.
[0171] 3. SNP extraction
[0172] Using an R programming language program with the Biostrings and openxlsx packages installed, code was written to extract all SNP sites from the aligned sequences. First, the Biostrings package was used to read the aligned sequence FASTA file, which was then converted to a matrix. Basic R language statements such as loops, conditional statements, subset extraction, and uniqueness counting were used to extract SNP sites. Then, the openxlsx package was used to write the extracted SNP sites to an Excel file, replacing default sites with NA. A total of 1867 SNP sites were extracted through this process. This set of SNP sites covers the entire coding and control regions of the porcine mitochondrial genome, fully reflecting the genetic differentiation characteristics among different geographical populations. Specific site information, genotype matrices, and location annotation information are retained by the applicant for future reference and are not submitted with this patent application.
[0173] 4. Model training and evaluation
[0174] In the last column containing SNP loci, source country information is supplemented based on known information. Given the data skewness and the current need for origin prediction only for products from China, Russia, and the United States, to fully utilize existing data and improve prediction accuracy, a four-layer progressive hierarchical model is not adopted. Instead, a random forest method is used to design a model that first coarsely classifies by continent, then subdivides by region, gradually narrowing the discrimination range and reducing interference from overlapping features across multiple categories. The first step involves merging data from Asia and Europe / Oceania, using all valid mitochondrial SNP loci as model input without feature reduction, preserving global species genetic differentiation information, and performing binary classification training. Based on this, further classification training is conducted on sequences from the United States and Russia in the Europe / Oceania data, and on sequences from China in the Asian data.
[0175] Dataset cleaning and feature preprocessing were performed using R language and the dplyr package; SNP genotype data reading and result output were performed using the readxl and writexl packages; model training, prediction, and accuracy evaluation were performed using the randomForest package; and model performance visualization and plotting were performed using ggplot2. The interactive origin tracing analysis system was built using the shiny, DT, and shinythemes packages. After model training, origin tracing model files integrating continental binary classification and further subdivision of Chinese, Russian, and American products were generated. An interactive file based on shiny was also written, which can import models for accuracy evaluation. The accuracy evaluation results of the four models are shown below. Figure 6 .
[0176] 5. Predictive Analysis
[0177] A predictive analysis module was added to the Shiny-based interactive file. By importing an Excel file containing all SNP sites, the module predicts the origin of samples listed as "Unknown" from the three known origins, and provides the confidence level. Prediction results for the seven known samples are shown below. Figure 7 .
[0178] The present application has been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of the present application. The descriptions of the embodiments above are only for the purpose of helping to understand the present application and its core ideas. It should be noted that those skilled in the art can make several improvements and modifications to the present application without departing from the principles of the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A primer set for amplifying the whole genome of porcine mitochondria, comprising nucleotide sequences as shown in SEQ ID NO. 1~12.
2. The primer set according to claim 1, which is divided into two primer pools as follows: Primer pool A consists of the sequences shown in SEQ ID NO.1, SEQ ID NO.2, SEQ ID NO.5, SEQ ID NO.6, SEQ ID NO.9, and SEQ ID NO.10; Primer pool B consists of the sequences shown in SEQ ID NO.3, SEQ ID NO.4, SEQ ID NO.7, SEQ ID NO.8, SEQ ID NO.11, and SEQ ID NO.
12.
3. A kit for amplifying the whole genome of porcine mitochondria, comprising the primer set as described in claim 1 or 2.
4. The kit according to claim 3 further includes multiplex PCR reaction solution, high-fidelity DNA polymerase, and nuclease-free water.
5. A PCR method for amplifying the entire porcine mitochondrial genome, comprising the following steps: Collect samples and extract nucleic acids from them; Using the extracted nucleic acid as a template, multiplex PCR amplification reactions were performed using primer pool A and primer pool B as described in claim 2, respectively, to obtain two sets of amplification products; The two sets of amplification products were mixed to obtain a mixed amplification product of the entire porcine mitochondrial genome.
6. The PCR method according to claim 5, wherein the reaction system for the multiplex PCR amplification is: 5×PrimeSTAR Buffer 10 μL, dNTP Mixture 4 μL, PrimeSTAR HS DNA Polymerase 0.5 μL, Primer Pool A or B 0.2–1.4 μmol / L, Nucleic Acid Template 5 μL, Add nuclease-free water to 50 μL; The reaction conditions for the multiplex PCR amplification are as follows: 95 °C for 2 min; 98 °C for 10 s, 55 °C to 50.5 °C, drop 0.5 °C per cycle and hold for 15 s, 68 °C for 3 min 40 s, 10 cycles; 98 °C for 10 s, 50 °C for 15 s, 68 °C for 3 min 40 s, 22 cycles; 72 °C for 5 min.
7. The PCR method according to claim 5, wherein the sample is derived from at least one of pig tissue, blood, meat products, processed pig by-products, or fecal swabs.
8. The PCR method according to claim 5, wherein the method for extracting nucleic acids from the sample is a magnetic bead method or a column method.
9. A targeted sequencing method for the entire porcine mitochondrial genome, comprising the following steps: (1) The whole genome of porcine mitochondria is targeted amplified using the PCR method described in any one of claims 5-8 to obtain mixed amplification products; (2) Construct sequencing libraries from the amplification products of step (1); (3) Perform nanopore sequencing on the sequencing library from step (2); (4) Sequence assembly of sequencing results: The pig mitochondrial reference genome is split into corresponding fragments according to the primer amplification fragments as assembly reference sequences. Each fragment contig is extracted from the sequencing results, and after removing the primers, it is assembled again to obtain the complete full-length sequence of pig mitochondria.
10. A method for constructing a porcine mitochondrial genome-wide SNP-based origin tracing model, comprising the following steps: (1) Dataset construction: 226 known mitochondrial whole genome sequences from different countries were screened from GenBank, and 45 mitochondrial whole genome sequences from different countries were determined using the targeted sequencing method described in claim 9, forming 271 mitochondrial whole genome datasets; (2) SNP extraction: After aligning the sequence from step (1) using the Maft software, a total of 1867 SNP sites were extracted using R language and the Biostrings package, and the default sites were replaced with NA; the SNP sites cover the entire coding and control regions of the porcine mitochondrial genome; (3) Model training: The random forest method is used to design a four-layer progressive hierarchical modeling, first coarsely classifying continents, then subdividing regions, gradually narrowing the discrimination range and reducing the interference of multi-class feature overlap; Specifically, the process includes: First, merging data from Asia and Europe / Oceania respectively, using all valid mitochondrial SNP sites as model input without feature reduction, preserving global species genetic differentiation information, and performing binary classification training; based on this, further classification training is performed on sequences from the United States and Russia in the Europe / Oceania data, and on sequences from China in the Asian data. The dataset cleaning and feature preprocessing were performed using R language and the dplyr package; the SNP genotype data reading and result output were performed using the readxl and writexl packages; the model training, prediction, and accuracy evaluation were performed using the randomForest package; and the model performance visualization plot was performed using ggplot2. The interactive source tracing analysis system was built using the shiny, DT, and shinythemes packages.