Microsatellite panel and its application in identification of parentage of largemouth bass
By constructing a microsatellite panel for multiplex PCR targeted sequencing, the problems of low efficiency in traditional SSR analysis methods and high cost in SNP technology have been solved, enabling high-throughput, low-cost, and accurate genetic analysis of the phylogenetic relationships of largemouth bass and promoting the scientific breeding of aquatic species.
Patent Information
- Application Number
- CN202510102339.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-01-22
AI Technical Summary
Traditional SSR analysis methods have low automation and efficiency, making it difficult to achieve high throughput. Existing SNP technology is expensive and cannot meet the needs of largemouth bass farming populations for high-throughput kinship analysis, thus hindering the scientific and efficient development of the largemouth bass farming industry.
We constructed a microsatellite panel for multiplex PCR targeted sequencing, containing multiple primer pairs targeting microsatellite loci. We employed an optimized primer set and bioinformatics algorithms, combined with CERVUS software, to perform rigorous paternity testing analysis, achieving high throughput and automation.
It improves operational efficiency, reduces costs, ensures the accuracy and reliability of analytical results, is suitable for large-scale genetic research, especially for the identification of kinship in largemouth bass, and provides a high-throughput genetic analysis tool.
Smart Images

Figure CN119662854B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of aquatic genetic breeding, and particularly relates to a microsatellite panel and application thereof in parentage identification of Micropterus salmoides. BACKGROUND
[0002] Microsatellite (SSR) markers have unique advantages of high polymorphism and co-dominance, and play an important role in the field of genetic research, especially in genetic diversity analysis and parentage identification. Traditional SSR analysis methods rely on manual operation, tedious electrophoresis detection and manual interpretation, which not only has a complex process, but also has a high time cost. Each experiment from sample processing to result acquisition often takes several days, which is difficult to meet the timeliness requirement of large-scale genetic research. Moreover, the traditional typing method relies on the length of the amplified sequence, which limits the number of multiplex PCR, resulting in a throughput of only 3-4. This limitation restricts the rapid analysis of genetic characteristics of large-scale populations and limits the application in aspects such as species evolution research and large-scale germplasm screening. In recent years, SNP technology has emerged, which has advantages of high throughput and automation in some aspects, but the construction and analysis cost is very high. It needs to invest a lot of funds for chip design, reaction system construction and professional data analysis software, which is a heavy economic burden for many scientific research institutions and enterprises. At the same time, under the same number of markers, SSR markers have higher polymorphism than SNP markers, and SNP markers cannot completely replace SSR markers in certain genetic analysis scenarios. Focusing on Micropterus salmoides, an important aquaculture species, there is currently a lack of efficient SSR panel based on multiplex PCR targeted sequencing technology. The large-scale aquaculture industry of Micropterus salmoides is expanding, and high-throughput parentage analysis can help achieve precise breeding, improve breeding efficiency and reduce the problem of genetic degradation caused by inbreeding. However, the existing analysis technology cannot meet the urgent needs of high-throughput parentage analysis of Micropterus salmoides breeding populations, and cannot provide strong technical support for the development of the industry, which seriously hinders the scientific and efficient development of the Micropterus salmoides aquaculture industry. SUMMARY
[0003] The present application provides a microsatellite panel and application thereof in parentage identification of Micropterus salmoides, which solves the problem of low degree of automation, low efficiency and difficulty in achieving high throughput of traditional SSR analysis methods.
[0004] In order to achieve the above-mentioned application purposes, the present application provides the following technical solutions:
[0005] The present application provides a microsatellite panel, which contains a plurality of primer pairs for microsatellite sites.
[0006] The multiple primer pairs targeting microsatellite sites include primers with nucleotide sequences as shown in SEQ ID NO. 1 to 80.
[0007] The present invention provides a microsatellite panel containing multiple primer pairs targeting microsatellite loci;
[0008] The plurality of primer pairs targeting microsatellite loci include:
[0009] The nucleotide sequences are as shown in the primer pairs in SEQ ID NO.27 to SEQ ID NO.28;
[0010] Nucleotide sequences are shown in the primer pairs as indicated in SEQ ID NO.33 to SEQ ID NO.34;
[0011] Nucleotide sequences are shown in primer pairs as indicated in SEQ ID NO.35 to SEQ ID NO.36;
[0012] Nucleotide sequences are shown in the primer pairs as indicated in SEQ ID NO.37 to SEQ ID NO.38;
[0013] The nucleotide sequences are as shown in the primer pairs in SEQ ID NO.41 to SEQ ID NO.42;
[0014] Nucleotide sequences are shown in primer pairs as indicated in SEQ ID NO.47 to SEQ ID NO.48;
[0015] Nucleotide sequences are shown in primer pairs as indicated in SEQ ID NO. 53 to SEQ ID NO. 54;
[0016] Nucleotide sequences are shown in primer pairs as indicated in SEQ ID NO. 55 to SEQ ID NO. 56;
[0017] Nucleotide sequences are shown in the primer pairs as indicated in SEQ ID NO. 57 to SEQ ID NO. 58;
[0018] Nucleotide sequences are shown in primer pairs as indicated in SEQ ID NO. 59 to SEQ ID NO. 60;
[0019] Nucleotide sequences are shown in the primer pairs as indicated by SEQ ID NO. 61 to SEQ ID NO. 62;
[0020] Nucleotide sequences are shown in the primer pairs as indicated in SEQ ID NO. 63 to SEQ ID NO. 64;
[0021] Nucleotide sequences are shown in primer pairs as indicated in SEQ ID NO. 69 to SEQ ID NO. 70;
[0022] The nucleotide sequences are as shown in the primer pairs in SEQ ID NO.71 to SEQ ID NO.72;
[0023] The nucleotide sequences are as shown in the primer pairs in SEQ ID NO.75 to SEQ ID NO.76.
[0024] This invention provides a primer set consisting of multiple primer pairs targeting microsatellite loci;
[0025] The multiple primer pairs targeting microsatellite sites include primers with nucleotide sequences as shown in SEQ ID NO. 1 to 80.
[0026] This invention provides a primer set consisting of multiple primer pairs targeting microsatellite loci;
[0027] The plurality of primer pairs targeting microsatellite loci include:
[0028] The nucleotide sequences are as shown in the primer pairs in SEQ ID NO.27 to SEQ ID NO.28;
[0029] Nucleotide sequences are shown in the primer pairs as indicated in SEQ ID NO.33 to SEQ ID NO.34;
[0030] Nucleotide sequences are shown in primer pairs as indicated in SEQ ID NO.35 to SEQ ID NO.36;
[0031] Nucleotide sequences are shown in the primer pairs as indicated in SEQ ID NO.37 to SEQ ID NO.38;
[0032] The nucleotide sequences are as shown in the primer pairs in SEQ ID NO.41 to SEQ ID NO.42;
[0033] Nucleotide sequences are shown in primer pairs as indicated in SEQ ID NO.47 to SEQ ID NO.48;
[0034] Nucleotide sequences are shown in primer pairs as indicated in SEQ ID NO. 53 to SEQ ID NO. 54;
[0035] Nucleotide sequences are shown in primer pairs as indicated in SEQ ID NO. 55 to SEQ ID NO. 56;
[0036] Nucleotide sequences are shown in the primer pairs as indicated in SEQ ID NO. 57 to SEQ ID NO. 58;
[0037] Nucleotide sequences are shown in primer pairs as indicated in SEQ ID NO. 59 to SEQ ID NO. 60;
[0038] Nucleotide sequences are shown in the primer pairs as indicated by SEQ ID NO. 61 to SEQ ID NO. 62;
[0039] Nucleotide sequences are shown in the primer pairs as indicated in SEQ ID NO. 63 to SEQ ID NO. 64;
[0040] Nucleotide sequences are shown in primer pairs as indicated in SEQ ID NO. 69 to SEQ ID NO. 70;
[0041] The nucleotide sequences are as shown in the primer pairs in SEQ ID NO.71 to SEQ ID NO.72;
[0042] The nucleotide sequences are as shown in the primer pairs in SEQ ID NO.75 to SEQ ID NO.76.
[0043] This invention provides a multi-microsatellite identification system, which includes microsatellite panel A, microsatellite panel B, microsatellite panel C and microsatellite panel D;
[0044] The microsatellite panel A contains multiple primer pairs targeting microsatellite sites, including primers with nucleotide sequences as shown in SEQ ID NO. 1 to 20;
[0045] The microsatellite panel B contains multiple primer pairs targeting microsatellite sites, including primers with nucleotide sequences as shown in SEQ ID NO.21-40;
[0046] The microsatellite panel C contains multiple primer pairs targeting microsatellite sites, including primers with nucleotide sequences as shown in SEQ ID NO. 41-60;
[0047] The microsatellite panel D contains multiple primer pairs targeting microsatellite sites, including primers with nucleotide sequences as shown in SEQ ID NO. 61-80.
[0048] This invention provides a multi-microsatellite identification system, which includes microsatellite panel 1, microsatellite panel 2 and microsatellite panel 3;
[0049] The microsatellite panel 1 contains:
[0050] The nucleotide sequences are as shown in the primer pairs in SEQ ID NO.27 to SEQ ID NO.28;
[0051] Nucleotide sequences are shown in the primer pairs as indicated in SEQ ID NO.33 to SEQ ID NO.34;
[0052] Nucleotide sequences are shown in primer pairs as indicated in SEQ ID NO.35 to SEQ ID NO.36;
[0053] Nucleotide sequences are shown in the primer pairs as indicated in SEQ ID NO.37 to SEQ ID NO.38;
[0054] The nucleotide sequences are as shown in the primer pairs in SEQ ID NO.41 to SEQ ID NO.42;
[0055] The microsatellite panel 2 contains:
[0056] Nucleotide sequences are shown in primer pairs as indicated in SEQ ID NO.47 to SEQ ID NO.48;
[0057] Nucleotide sequences are shown in primer pairs as indicated in SEQ ID NO. 53 to SEQ ID NO. 54;
[0058] Nucleotide sequences are shown in primer pairs as indicated in SEQ ID NO. 55 to SEQ ID NO. 56;
[0059] Nucleotide sequences are shown in the primer pairs as indicated in SEQ ID NO. 57 to SEQ ID NO. 58;
[0060] Nucleotide sequences are shown in primer pairs as indicated in SEQ ID NO. 59 to SEQ ID NO. 60;
[0061] The microsatellite panel 3 contains:
[0062] Nucleotide sequences are shown in the primer pairs as indicated by SEQ ID NO. 61 to SEQ ID NO. 62;
[0063] Nucleotide sequences are shown in the primer pairs as indicated in SEQ ID NO. 63 to SEQ ID NO. 64;
[0064] Nucleotide sequences are shown in primer pairs as indicated in SEQ ID NO. 69 to SEQ ID NO. 70;
[0065] The nucleotide sequences are as shown in the primer pairs in SEQ ID NO.71 to SEQ ID NO.72;
[0066] The nucleotide sequences are as shown in the primer pairs in SEQ ID NO.75 to SEQ ID NO.76.
[0067] This invention provides a multiplex PCR detection kit for identifying the kinship of largemouth bass, characterized by comprising the above-mentioned primer set.
[0068] This invention provides the application of the above-mentioned microsatellite panels, primer sets, multiplex microsatellite identification systems, or multiplex PCR detection kits in individual identification, parentage identification, or population genetic analysis of largemouth bass.
[0069] The beneficial effects of this invention are:
[0070] This invention focuses on the identification of kinship among largemouth bass and constructs an SSR panel based on multiplex PCR targeted sequencing. It has significant advantages in terms of operational efficiency, cost control, result accuracy, and application scope, and has powerfully promoted genetic research on largemouth bass and related aquatic species.
[0071] This invention utilizes multiplex PCR targeted sequencing technology to achieve a high throughput and automation level comparable to SNP panels. Compared with traditional SSR analysis methods, it significantly improves operational efficiency, overcomes the problems of low automation and difficulty in achieving high throughput in traditional methods, and can meet the needs of large-scale genetic research. It also greatly reduces human intervention and lowers costs. Compared with traditional SSR analysis methods and SNP technology, this invention has a significant cost advantage in the construction and analysis process, alleviating the problem of high construction and analysis costs of SNP technology, making large-scale kinship analysis more feasible.
[0072] This invention utilizes optimized primer sets and bioinformatics algorithms to ensure the accuracy of kinship analysis at multiple stages. When screening SSR loci, various software and algorithms are employed to construct primer matrices and calculate scores to select appropriate primer sets. In the data analysis phase, a rigorous paternity testing process is conducted using CERVUS software, with thresholds set through simulations to ensure reliable results. Paternity exclusion probability assessments of four sets of deca-fold SSR panels show a high cumulative exclusion probability under different known parentage conditions. Furthermore, validation with different panel combinations and optimized primer sets significantly improves paternity assignment accuracy. The SSR panels constructed in this invention are not only applicable to largemouth bass but can also be extended to kinship and genetic analysis of other aquatic species, providing a universal and effective tool and method for genetic research in aquatic species. Attached Figure Description
[0073] Figure 1 The chromosomal distribution of the loci in the SSR marker panel. These loci are distributed on 15 chromosomes (Chr1-Chr23), with no loci on chromosomes 2, 3, 5, 7, 11, 13, 19, and 22;
[0074] Figure 2 Heatmap representation of genotype sequencing reads for four deca-fold SSR panels. Figures (A) to (D) show the read distribution for each locus and sample in each SSR panel, with color intensity (from yellow to blue) representing the number of reads;
[0075] Figure 3 The cumulative exclusion probabilities for paternity testing of four deca-sigma SSR marker combinations are shown in Figures (A) to (D). Figures (A) to (D) illustrate the exclusion probabilities for each locus in three scenarios: unknown parentage (CE-1P), known parentage (CE-2P), and known parentage (CE-PP). The cumulative exclusion probabilities range from 0 to 1, where 0 indicates no exclusion capability and 1 indicates complete exclusion capability.
[0076] Figure 4 To illustrate the cumulative exclusion probability and paternity testing accuracy for the optimized three quintuple SSR marker combinations, different loci and sample sizes are presented. (A) The figure shows the cumulative exclusion probability for single parent (CE-1P), two parents (CE-2P), and known parent pair (CE-PP) as the loci increase sequentially. The exclusion probability ranges from 0 to 1, where 0 indicates no exclusion capability and 1 indicates complete exclusion capability. (B) The figure shows the paternity testing accuracy under different conditions (maternal only, paternal only, and known sex parent pair), evaluated as a function of sample size. The accuracy range is also from 0 to 1, where 0 indicates a completely incorrect identification result and 1 indicates a completely accurate identification result based on complete consistency with the known parent pair. Detailed Implementation
[0077] The technical solutions provided by the present invention will be described in detail below with reference to the embodiments, but they should not be construed as limiting the scope of protection of the present invention.
[0078] Example
[0079] 1. Sample collection and DNA extraction:
[0080] The largemouth bass experimental samples studied in this study were obtained from four families with known pedigrees established in 2023 by the Zhejiang Provincial Aquatic Technology Extension Station in Hangzhou, Zhejiang Province, China. Four single-pair families (each family containing one paternal parent, one maternal parent, and ten offspring) were used for paternity testing. Pectoral fin or muscle tissue from all individuals was preserved in anhydrous ethanol at -20°C. Genomic DNA was extracted using a magnetic bead-based tissue genomic DNA extraction kit (Wuhan Nanomagnetic Biotechnology Co., Ltd.). The concentration of each DNA sample was determined using an ND5000 spectrophotometer (Beijing Baitai Scientific Instruments Co., Ltd.). All DNA samples were diluted to a working concentration of 100 ng / μL and then diluted with deionized water before being stored at -20°C until subsequent use.
[0081] 2. SSR screening and primer design:
[0082] To conduct simple sequence repeat (SSR) screening and develop a multiplex PCR system, this study selected six largemouth bass individuals from three different populations. The samples were obtained from Bolongkeng Reservoir in Quzhou City, Zhejiang Province, China; Anhui Zhanglin Fishery Co., Ltd. in Zhejiang Province, China; and Laibei Fishery Co., Ltd. in Panzhihua City, Sichuan Province, China, with two samples taken from each population. Genomic DNA was extracted from pectoral fin tissue collected and preserved in anhydrous ethanol using a magnetic bead-based tissue genomic DNA extraction kit (Wuhan Nanomagnetic Biotechnology Co., Ltd.).
[0083] Library preparation and sequencing were performed according to Illumina's whole-genome next-generation sequencing (NGS) protocol. Sequencing was conducted by Wuhan Shafei Bioinformatics Co., Ltd., China. Raw reads were quality controlled using Fastp v0.23.1 with the parameter set to -cD to ensure the generation of high-quality, clean reads. Clean reads were aligned with the largemouth bass reference genome GCA_022435785.1 using bwa-memev 1.0.5 to generate SAM files. These SAM files were converted to BAM format using samtools v1.6, and duplicate reads were removed using sambamba v1.0.
[0084] Single nucleotide polymorphisms (SNPs) were identified using the `mpileup` and `call` commands in the `bcftools` v1.8 software package, generating a variant call format (VCF) file containing the SNPs. For simple sequence repeats (SSRs), the VCF file was generated using `lobSTR` v4.0.6 software. Based on these SNP and SSR VCF files, this study used `MultiplexSSR` software, setting the maximum amplified fragment length to 230 bp and the annealing temperature to 60 °C, to construct an SSR multiplex primer matrix. This matrix was generated using `MultiPLX` v1.4 software. `MultiplexSSR` can mask SNPs and SSRs in the target region, thereby improving the effectiveness of primer design. Subsequently, the `SADDLE` algorithm was applied to process the matrix, calculating the cumulative badness score for each primer pair and arranging them in ascending order of score. This process enabled the construction of four deca-base primer pairs, with the badness score for each primer pair calculated using Equation 1.
[0085]
[0086] When defining Badness, the summation process considers all inverse complementary subsequences between primers pa and pb, which must have at least four consecutive complementary bases. The parameters are defined as follows: len represents the length of the complementary subsequence; d1 and d2 represent the distances from the 3' ends of primers pa and pb to the start positions of the complementary subsequences, respectively; and numGC refers to the number of G / C bases in the subsequence. These parameters together constitute the Badness scoring system, which assesses the likelihood of unwanted primer interactions based on the complementary regions of the primers, particularly near the 3' ends, as such interactions can affect PCR efficiency. A lower Badness score means a lower probability of dimer formation between the primer pairs, thus less interference with multiplex PCR. Conversely, the formation of numerous dimers inhibits multiplex PCR reactions and reduces the amplification efficiency of the target region.
[0087] We utilized a small amount (typically 2-6) of low-depth whole-genome resequencing data and employed the MultiplexSSR method, setting the maximum amplified fragment length to 230 bp and the annealing temperature to 60℃. Using MultiPLX v1.4 software, we generated an SSR multiplex primer matrix containing 40 primer sets, each set containing 10-11 primers. Based on this multiplex primer matrix, we calculated all possible deca-primer combinations. The Badness score for each primer set was calculated using the SADDLE algorithm. Based on the ascending order of Badness scores, the four deca-primer sets with the lowest scores were selected to construct four SSR panels. The chromosomal locations of the corresponding loci are shown in the image. Figure 1 These loci are presented in the image. They cover 15 of the 23 chromosomes, with the exception of chromosomes 2, 3, 5, 7, 11, 13, 19, and 22.
[0088] Table 1. Characteristics of four deca-score SSR panels for largemouth bass.
[0089]
[0090]
[0091]
[0092] The Badness score refers to the cumulative sum of Badness values for a deca-particle PCR primer. This score is calculated for both the forward and reverse primers after adding the bridging sequence “ggagtgagtacggtgtgc (SEQ ID NO.81)” to the forward primer and the bridging sequence “gagttggatgctggatgg (SEQ ID NO.82)” to the reverse primer.
[0093] 3. Library Construction and Expansion:
[0094] Based on four sets of deca-primer sets, bridging sequences were added to each primer pair to meet the requirements of Hi-TOM amplicon library preparation: "ggagtgagtacggtgtgc" for the forward primer and "gagttggat gctggatgg" for the reverse primer. The purpose of introducing the bridging sequences was to ligate them to the Hi-TOM sequencing adapters for amplicon sequencing. Subsequently, the cumulative Badness score for each primer set was recalculated (the Badness score in Table 1 refers to the score obtained after introducing the bridging sequences). PCR amplification was performed in a 50 μL reaction volume, including 20 μL of each primer set (2 μL per primer pair, 10 μmol), 25 μL of Gold Multiplex PCR Mix (Jiangsu Kewen Biotechnology Co., Ltd.), and 5 μL of sample DNA. The PCR program consisted of initial denaturation at 95°C for 10 minutes, followed by 10 cycles, each consisting of 30 seconds at 95°C, 30 seconds of annealing, and 30 seconds at 72°C. This was followed by an additional 20 cycles, each consisting of 30 seconds at 95°C, 30 seconds at 50°C, and 30 seconds at 72°C. Finally, a final extension step was performed at 72°C for 5 minutes. The initial annealing temperature was set at 60°C and decreased by 1°C per cycle. Equal amounts of the PCR products were submitted to the Hi-TOM platform for next-generation sequencing (NGS), located at the State Key Laboratory of Rice Biology, National Rice Research Institute, Chinese Academy of Agricultural Sciences, Hangzhou, China.
[0095] Raw paired-end reads obtained from Hi-TOM sequencing were quality-controlled using Fastp v0.23.1 with parameters set to -c and -D to ensure the generation of high-quality, clean reads for subsequent analysis. These clean reads were used as input for SSRseq for SSR genotyping. SSRseq first used FLASH v1.2.11 to merge overlapping clean reads; only the merged overlapping reads were retained for downstream SSR genotyping. Then, SSRseq used BLAST v2.12 to detect SSRs and applied its internal algorithm to mitigate the PCR slippage effect during SSR genotyping. For each sample, the slippage ratio of the SSR sequence at the locus was calculated using the following formula:
[0096]
[0097] At a specific site, the slip ratio (SRn) of the SSR sequence is related to the number of repetitions n, where Fn-1 represents the read frequency of the stutter peak with a repetition count of n-1, and Fn represents the read frequency of the SSR sequence. The slip ratio is modeled as a variation with the number of repetitions at the site, and its mathematical expression is shown below:
[0098] SR n =an 2 +bn
[0099] a and b represent site-specific coefficients. The coefficients a and b for each SSR locus were estimated using regression analysis between the mean observed slip ratio and the number of SSR alleles for each locus of repeats. For SSR sequences lacking observed slip ratios, the expected slip ratio was calculated using an equation. These slip ratios, estimated based on high-quality samples, were then used to correct for slip artifacts in all samples. The final output is an Excel file containing SSR genotyping results with slip artifacts removed.
[0100] Based on genotype data, allele frequencies were analyzed using CERVUS v3.0.7 software. The calculated genetic parameters included the number of alleles (Na), observed heterozygosity (Ho), expected heterozygosity (He), polymorphism information content (PIC), Hardy-Weinberg equilibrium (HWE), and zero allele frequency (Fnull). HWE was assessed using a minimum expected frequency of 3 as a criterion, and significance was tested using the Bonferroni correction method. The comprehensive exclusion probability (CEP) was calculated as follows:
[0101] CEP=1-NEP1×NEP2×NEP3×...×NEPn
[0102] Where NEPn represents the non-exclusion probability of the nth SSR locus.
[0103] Genotyping was performed on four single-pair families (each family containing one father, one mother, and ten offspring) using SSRseq technology, involving four 10-fold PCR SSR panels. SSRseq first used FLASH to merge paired end reads, followed by BLAST screening for merged reads containing SSRs for genotyping. When using SSRseq for genotyping, it is essential to ensure that the SSRs are perfect SSRs, meaning the repetitive sequence must be an integer multiple of the motif and must not contain SNPs, insertions, or deletions. Therefore, SSRs with multiple motifs at the same location were split. The number before the underscore ("_") represents the SSR locus number, while the number after the underscore indicates the motif number at that locus.
[0104] In panel A, the average percentage of pooled reads was 92.33%, of which 77.38% were available for genotyping. Panel B had an average percentage of pooled reads of 96.17%, with 84.94% available for genotyping. The corresponding percentages for panel C were 95.90% and 80.58%, respectively, while panel D had an average percentage of pooled reads of 95.79%, with 79.40% available for genotyping. We then quantified the number of reads available for genotyping at each locus in each SSR panel and presented the results. Figure 2 Notably, we observed significant differences in read counts at each locus across samples within each SSR panel. Panel B had the lowest coefficient of variation for read counts, averaging 0.347, while panel D had the highest, at 0.749. Despite this variability, genotyping was successfully performed at all loci across all panels except locus 440 in panel D. Since SSRseq requires a minimum of 30 reads for reliable genotyping, this locus failed to genotype in ten samples due to insufficient reads. All other loci exceeded this threshold, with minimum read counts far exceeding the limit: 34 for locus 362, 93 for locus 870, 279 for locus 164, 367 for locus 304, 503 for locus 881, 520 for locus 614, 747 for locus 760, and so on. Figure 2 As shown.
[0105] Polymorphism statistical analysis was performed on amplicon sequencing data using CERVUS v3.0.7 software. As shown in Table 2, the sample size for genotyping ranged from 38 to 48, and the genotyping success rate for all loci reached 98.2%. The number of alleles (Na) ranged from 2 to 7, with an average of 4.06. The average Na was highest in group C (4.36) and lowest in group D (3.69). Observed heterozygosity (Ho) ranged from 0.104 to 1.000, with an average of 0.929; while expected heterozygosity (He) ranged from 0.101 to 0.754, with an average of 0.589. Polymorphism information content (PIC) ranged from 0.098 to 0.713, with an average of 0.507. In addition, we found that most sites had high Ho values, usually close to or equal to 1, but a few sites (such as 614_3, 42 and 730) had low Ho values, which may be the reason for the deviation from Hardy-Weinberg equilibrium.
[0106] Table 2. Polymorphism statistics of four deca-fold SSR marker panels for largemouth bass
[0107]
[0108]
[0109] Notes: NS: No significant difference (P>0.05); ***: Extremely significant difference (P<0.001); **: Highly significant difference (P<0.01); *: Significant difference (P<0.05); ND: No significant difference analysis performed; HWE: Hardy-Weinberg equilibrium; F(Null): Null allele frequency.
[0110] 4. Data analysis and kinship testing:
[0111] Paternity testing was performed using CERVUS v3.0.7 software in a three-stage process. First, the allele frequency module was used to calculate the allele frequency for each SSR locus, as described previously. Second, the simulation module was used to estimate the threshold log-likelihood ratio (LOD) score for parental pairing, assuming sex was known. In this stage, 10,000 offspring were simulated, with four candidate mothers and four candidate fathers. Based on allele frequency analysis, genotyping was performed on the locus proportions, with a minimum number of genotyping loci set to 5. The simulation process used a default genotype error rate of 1%, and set 95% strict confidence levels and 90% relaxed confidence levels to calculate the LOD score. Finally, the parental attribution module assigned parental pairs to each offspring based on the genotype data and the simulated LOD threshold. The accuracy of the paternity testing was verified by comparing the results with known pedigrees, ensuring the reliability and consistency of the results.
[0112] The parentage exclusion probability was assessed for four sets of deca-fold SSR panels. When both parents were unknown (E-1P), the exclusion probability for a single locus ranged from 0.005 to 0.353; when one parent was known (E-2P), the exclusion probability ranged from 0.051 to 0.536; and when both parents were known (E-PP), the exclusion probability ranged from 0.096 to 0.728. Figure 3 The results show that when both parents are unknown, the cumulative exclusion probability of all panels exceeds 0.9, with panel A reaching a high of 0.964. If one parent is known, the cumulative exclusion probabilities of panels A, B, C, and D are 0.998, 0.987, 0.997, and 0.992, respectively. When both parents are known, the cumulative exclusion probabilities of each panel approach 1.0: panel A is 0.99997, panel B is 0.99928, panel C is 0.99995, and panel D is 0.99961.
[0113] Simulated parental pairing allocation was performed across four panels, assuming parental information was known. Parental pairing was assigned to each offspring based on the simulated LOD threshold for each panel. In panel A, the simulated and actual parental pairing rates reached 100% in all categories (single mother, single father, or parental pairing) regardless of whether the confidence level was strict or relaxed. In panel B, the simulated and actual parental pairing rates reached 100% at the relaxed confidence level. However, at the strict confidence level, the simulated parental pairing rates were 98%, 95%, and 100%, respectively, while the actual parental pairing rates were 60%, 57%, and 100%, respectively. In panel C, the simulated parental pairing rate was 100% at both confidence levels, while the actual parental pairing rates were 93%, 100%, and 95%, respectively. In panel D, the simulated and actual parental pairing rates reached 100% at both strict and relaxed confidence levels.
[0114] To verify the accuracy of the four SSR panels in paternity testing, four single-pair families (each family containing one father, one mother, and ten offspring) were used. As shown in Table 3, although each SSR panel achieved deca-pair PCR capability and contained more than 10 loci, the accuracy of any single panel was low. With known sex, the highest accuracy for identifying parental pairs was only 0.5, while the accuracy slightly improved when assessing the mother or father individually. By testing various combinations of the four SSR panels, the B+C+D panel combination was found to have the highest accuracy: 0.775 for a single mother, 0.975 for a single father, and 0.750 for parental pairs with known sex. Based on the B+C+D combination, locus selection was further optimized by removing loci that might interfere with paternity testing. This yielded an optimal primer set containing 15 primer pairs (368, 555, 565, 760, 12, 317, 715, 768, 795, 861, 172, 42, 700, 703, 777), covering a total of 22 loci. This optimized set improved accuracy to 0.90 for single-mother pairs, 0.95 for single-father pairs, and 0.85 for parental pairs with known sex.
[0115] Table 3. Accuracy of Paternity Testing for Largemouth Bass Using Four Sets of Decade-SSR Marker Panels
[0116]
[0117]
[0118] Note: To verify the accuracy of paternity testing using four deca-sigma SSR marker combinations (A, B, C, and D), we used 48 individuals (40 offspring and 8 parents). A set of 15 optimal primer pairs (368, 555, 565, 760, 12, 317, 715, 768, 795, 861, 172, 42, 700, 703, and 777) covering a total of 22 loci were ultimately selected. Accuracy values range from 0 to 1, where 0 represents a completely incorrect result and 1 represents a completely accurate result based on comparison with the actual parents.
[0119] In the initial design of 15 primer pairs, the cumulative Badness score reached as high as 8372.0163, indicating potential problems such as primer dimer formation in multiplex PCR reactions, which could reduce the amplification efficiency of multiplex PCR. To mitigate these problems, this invention reorganized the primers into three quintuplet SSR panels: Panel 1 (368,555,565,760,12), Panel 2 (317,715,768,795,861), and Panel 3 (172,42,700,703,777). The Badness scores of these panels were 45.8104, 40.0378, and 11.8396, respectively, effectively reducing interference between primers. The optimized primer set was validated using 104 progeny samples to evaluate its accuracy in parentage allocation. Figure 4 As shown, the optimized primer set achieved an accuracy of 0.9423 in single-mother assignments, 0.9712 in single-father assignments, and 0.9135 in parent-child pair assignments with known sex.
[0120] As can be seen from the above embodiments, this invention addresses the shortcomings of traditional SSR analysis methods and existing technologies in the identification of kinship in largemouth bass. It constructs an SSR panel based on multiplex PCR targeted sequencing for largemouth bass kinship identification through a series of steps. This method offers advantages such as high efficiency, low cost, high accuracy, and wide applicability.
[0121] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A multiple microsatellite identification system for the phylogenetic relationships of largemouth bass, characterized in that, The multi-microsatellite identification system includes microsatellite panel A, microsatellite panel B, microsatellite panel C, and microsatellite panel D; The microsatellite panel A contains multiple primer pairs targeting microsatellite sites, and the nucleotide sequences of the primer pairs are shown in SEQ ID NO.1~20; The microsatellite panel B contains multiple primer pairs targeting microsatellite sites, and the nucleotide sequences of the primer pairs are shown in SEQ ID NO.21~40; The microsatellite panel C contains multiple primer pairs targeting microsatellite sites, and the nucleotide sequences of the primer pairs are shown in SEQ ID NO.41~60; The microsatellite panel D contains multiple primer pairs targeting microsatellite sites, and the nucleotide sequences of the primer pairs are shown in SEQ ID NO. 61~80.
2. A multiple microsatellite identification system for the phylogenetic relationship of largemouth bass, characterized in that, The multi-microsatellite identification system includes microsatellite panel 1, microsatellite panel 2, and microsatellite panel 3; The microsatellite panel 1 contains: The nucleotide sequences are as shown in the primer pairs SEQ ID NO.27~SEQ ID NO.28; Nucleotide sequences are shown in the primer pairs as indicated in SEQ ID NO.33~SEQ ID NO.34; Nucleotide sequences are shown in primer pairs as indicated in SEQ ID NO.35~SEQ ID NO.36; The nucleotide sequences are as shown in the primer pairs SEQ ID NO.37~SEQ ID NO.38; The nucleotide sequences are as shown in the primer pairs in SEQ ID NO.41~SEQ ID NO.42; The microsatellite panel 2 contains: The nucleotide sequences are as shown in the primer pairs SEQ ID NO.47~SEQ ID NO.48; The nucleotide sequences are as shown in the primer pairs in SEQ ID NO. 53~SEQ ID NO. 54; Nucleotide sequences are shown in primer pairs as indicated by SEQ ID NO. 55 to SEQ ID NO. 56; The nucleotide sequences are as shown in the primer pairs SEQ ID NO.57~SEQ ID NO.58; Nucleotide sequences are shown in primer pairs as indicated by SEQ ID NO. 59 to SEQ ID NO. 60; The microsatellite panel 3 contains: Nucleotide sequences are shown in the primer pairs as indicated by SEQ ID NO. 61 to SEQ ID NO. 62; Nucleotide sequences are shown in the primer pairs as indicated by SEQ ID NO. 63~SEQ ID NO. 64; Nucleotide sequences are shown in primer pairs as indicated by SEQ ID NO. 69 to SEQ ID NO. 70; The nucleotide sequences are as shown in the primer pairs in SEQ ID NO.71~SEQ ID NO.72; The nucleotide sequences are primer pairs as shown in SEQ ID NO.75~SEQ ID NO.
76.
3. A multiplex PCR detection kit for identifying the kinship of largemouth bass, characterized in that, Including the identification system as described in claim 1.
4. A multiplex PCR detection kit for identifying the kinship of largemouth bass, characterized in that, Including the identification system as described in claim 2.
5. The application of the multiplex microsatellite identification system according to claim 1 or 2 or the multiplex PCR detection kit according to claim 3 or 4 in individual identification, parentage identification or population genetic analysis of largemouth bass.
Citation Information
Patent Citations
Microsatellite marker-based primer and method for paternity test of lateolabrax maculatus
CN113774152A