A molecular marker primer for a str of a sandfish and application thereof
By developing a primer combination of STR molecular markers for sea cucumber and a gradient enhancement model, the problem of rapid and low-cost detection for identifying the marine origin of sea cucumber has been solved, achieving efficient and accurate marine traceability, which is suitable for large-scale trade detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SCIENCE & TECHNOLOGY RESEARCH CENTER OF CHINA CUSTOMS
- Filing Date
- 2026-06-08
- Publication Date
- 2026-07-24
AI Technical Summary
Existing technologies are insufficient to efficiently distinguish sandfish from different sea areas. Traditional morphological identification methods are ineffective, and DNA molecular marker methods are costly and complex to use, making it difficult to meet the needs of rapid and low-cost detection in large-scale trade scenarios.
We developed three pairs of STR molecular marker primer combinations for sand gudgeon, combined with multiplex PCR and capillary electrophoresis or high-throughput sequencing, to construct a gradient boosting model for efficient identification of marine origin.
It achieves high-throughput, low-cost, and rapid identification of seaweed origin, with a detection accuracy of 100%. It is suitable for large-scale trade testing, reduces sample usage and reagent costs, and supports multi-platform data compatibility.
Smart Images

Figure CN122445809A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of molecular biology marker technology, specifically to a STR molecular marker primer for sand eel and its application. Background Technology
[0002] The sand gudgeon (Lepturacanthus savala) is an important economic fish species distributed in tropical and subtropical waters, accounting for a significant share of my country's and international seafood trade. Statistics show that in 2025, my country imported approximately 148,100 tons of sand gudgeon from India, primarily from the northwestern Indian Ocean (FAO 57.1) and adjacent areas of the Arabian Sea (FAO 51.4). The differences in environmental conditions and population genetic structure across different sea areas provide a theoretical basis for identifying the marine origin of sand gudgeon. However, sand gudgeon loses its morphological characteristics after processing such as freezing, decapitation, and segmentation, making traditional morphological identification methods insufficient for the needs of trade traceability.
[0003] Currently, DNA molecular marker-based species identification and origin tracing technologies are widely used. While mitochondrial DNA markers (such as COI and Cyt b genes) can be used for species differentiation, their degree of variation within populations is limited, making it difficult to effectively distinguish sand eels from different sea areas. Methods based on whole-genome or SNP markers, although possessing high resolution, suffer from high detection costs, complex data analysis, and long detection cycles, making it difficult to meet the rapid, low-cost detection needs of large-scale trade scenarios.
[0004] Short tandem repeats (STRs, also known as microsatellite DNA) are a class of highly polymorphic nuclear gene markers that have been widely used in forensic medicine, animal genetics and breeding, and aquatic product traceability. In Lepturacanthus savala, 28 polymorphic microsatellite loci have been developed using SLAF-seq technology, and their application in the genetic diversity of coastal populations in China has been evaluated (Development of microsatellite markers for Lepturacanthus savala based on SLAF-seq technology and their universality in closely related species, Genomics and Applied Biology, 2018, 37(8):3331-3338). This study provides a basic tool for population genetics of sand gudgeon, but its technical solution has the following shortcomings: (1) The purpose of this study is to investigate the genetic resources of the population, and it does not involve the application of origin tracing in different sea areas (especially the sea areas from which cross-border trade origins); (2) The detection method is conventional PCR combined with gel electrophoresis, which has low throughput and limited resolution, making it difficult to achieve multiple amplification and automated interpretation; (3) No predictive model for sea area origin identification has been constructed, and it is impossible to output quantifiable traceability results and confidence levels.
[0005] Chinese patents CN106947816A and CN103757113A disclose technical solutions for identifying other fish populations or determining parentage using microsatellite fluorescence multiplex PCR systems. There are also a few reports of applying machine learning models to the geographic origin tracing of genetic markers. However, these existing technologies all have the following drawbacks: (1) The target species is not gudgeon, and its primer sequences and amplification conditions are not universal; (2) Some solutions only use a single pair of primers or a single detection, which cannot achieve precise identification of multiple sea areas; (3) Most solutions do not integrate the entire chain of technologies including multiplex PCR, dual-platform detection (CE / NGS), and machine learning prediction, which is difficult to meet the needs of high-throughput, high-accuracy, and low-cost sea area origin identification in the international trade of gudgeon.
[0006] In addition, existing studies (such as “Development of microsatellite markers for sea cucumber based on SLAF-seq technology and their universality in closely related species”, 2018) have developed 28 polymorphic microsatellite loci for sea cucumber, but their detection method is conventional PCR combined with gel electrophoresis, which has low throughput and long detection time. Moreover, the developed markers are mainly used for population genetic diversity analysis, without involving the identification of individuals from different sea areas, and do not provide standardized feature construction methods and machine learning prediction models, which cannot meet the industrial application needs of sea area origin tracing in sea cucumber trade.
[0007] Therefore, there is an urgent need to develop a STR molecular marker and traceability method for sand ribbonfish that can efficiently distinguish the origins of different sea areas, is easy to operate, and is applicable to actual trade detection. Summary of the Invention
[0008] Therefore, this invention provides a STR molecular marker primer for sand eel and its application to solve the above-mentioned problems.
[0009] To achieve the above objectives, the present invention provides the following technical solution: According to a first aspect of the present invention, a molecular marker primer for sand eel STR is provided, wherein the primer combination consists of the following three pairs of primers: (1) LS4:CA8, its primer sequence is: SEQ ID NO.1: Forward primer: CAACAGCAGCAACAGGTTCC; SEQ ID NO.2: Reverse primer: GAGTGAGTGAGTGCGTGTGA; (2) LS6:AC16, its primer sequence is: SEQ ID NO.3: Forward primer: CAGCAAACAAAGAGTCGCGT; SEQ ID NO.4: Reverse primer: TGTTCCCTCCATCTGTCCCA; (3) LS8:AC16, its primer sequence is: SEQ ID NO.5: Forward primer: TTCGGCCCTTTCATGTGAGA; SEQ ID NO.6: Reverse primer: ACGCCACCATTTTCCAGACT.
[0010] In the primer combination, the 5' end of the forward primer of each primer pair is labeled with a fluorescent group; according to the length range of the amplified product of each primer pair, the three fluorescent groups FAM, HEX and ROX are used in rotation in ascending order to avoid signal interference when the product lengths overlap.
[0011] The second aspect of the present invention provides the application of the above-described STR molecular marker primers for sea cucumbers in the detection of genetic diversity, construction of STR fingerprints, and / or identification of the marine origin of sea cucumbers.
[0012] According to a third aspect of the present invention, a method for identifying the marine origin of *Gymnocypris spp.* based on the aforementioned *Gymnocypris spp.* STR molecular marker primers includes the following steps: (1) Extract genomic DNA from the sand ribbonfish to be tested; (2) Using the DNA obtained in step (1) as a template, a primer mixture is prepared using the three pairs of primers in claim 1, and a PCR reaction system is prepared for multiplex PCR amplification to obtain the amplification product; (3) Capillary electrophoresis was used to detect the amplification products to obtain raw capillary electrophoresis data, and / or second-generation high-throughput sequencing was used to sequence the amplification products to obtain sequencing data; (4) Construct corresponding sample STR feature vectors based on capillary electrophoresis data or next-generation sequencing data respectively; (5) Using the pre-trained model, the marine origin of the sand ribbonfish is predicted and determined, and the marine origin result and prediction confidence are output.
[0013] Furthermore, in step (2), the total PCR system is 20 μL: 10 μL of 2×KAPA KK2602 PCR Master Mix, 2.0 μL of premixed primer mixture, 1.5–2.0 μL of template DNA (total DNA 20–30 ng), and the remaining ddH2O is added to bring the total to 20 μL; the final concentration of each primer pair in the system is 0.1–0.2 μM, and the Mg in the Master Mix is [missing information]. 2+ Final concentration 1.5 mM, final concentration of each dNTP 200 μM.
[0014] Furthermore, the PCR amplification program for step (2) is as follows: pre-denaturation at 95℃ for 3 min; denaturation at 95℃ for 20 s, annealing at 59℃ for 15 s, extension at 72℃ for 30 s, for 30 cycles; final extension at 72℃ for 1 min, and storage at 4℃.
[0015] Furthermore, the steps for constructing STR feature vectors using capillary electrophoresis data are as follows: read the raw capillary electrophoresis signal → baseline calibration → LIZ internal standard calibration → peak identification → STR locus allocation → allele determination → generate sample STR feature vectors; when there is no effective amplification signal at a locus, the corresponding feature is assigned a value of -1 and a missing marker variable is added.
[0016] Furthermore, the steps for constructing STR feature vectors using second-generation sequencing data are as follows: quality control of raw sequencing data to remove contaminating sequences, low-quality sequences, and adapter sequences → assembly of paired-end sequencing sequences → statistical analysis of the distribution of different genotypes at each STR locus → calculation of the weighted average of the number of genotype repeats at each STR locus as the feature value of a single locus → combination of all locus feature values to obtain the STR feature vector; when there is no valid sequencing data for a locus, the corresponding feature is assigned a value of -1 and a missing marker variable is added.
[0017] The marine origin classification model mentioned in step (4) is a gradient boosting model; the model is trained and evaluated using leave-one-out cross-validation to distinguish sand gudgeon samples from the Indian Ocean FAO 57.1 marine area from the Arabian Sea FAO 51.4 marine area.
[0018] Furthermore, the primers used for capillary electrophoresis detection carry fluorescent labels at their 5' ends; according to the order of amplified product fragment length from shortest to longest, three fluorescent labels, FAM, HEX, and ROX, are used in rotation, and primers with overlapping fragment lengths are distinguished by different fluorescent labels.
[0019] According to a fourth aspect of the present invention, a marine origin identification system for sand gudgeon is provided, which is equipped with the identification method described above and includes a data input module, a STR data parsing module, a feature vector construction module, a model prediction module, and a result output module. The system imports the raw CE or NGS data of the sample to be tested, automatically completes feature construction and origin prediction, and outputs a marine origin and confidence report for sand gudgeon.
[0020] The machine learning model used in step (5) is gradient boosting.
[0021] As a preferred technical solution, the present invention also establishes a marine origin tracing database for sand gudgeon. The database contains the repetition frequency and phenotypic information of sand gudgeon samples from different marine areas at multiple STR sites, which is used to construct a standard phenotypic reference set and provide a data foundation for subsequent marine origin prediction and determination.
[0022] This invention also provides a method for constructing a marine source prediction model for sand gudgeon, comprising: Collect samples of sand gudgeon from known sea areas and obtain their STR spectral data; Provide the number of repetitions, fragment length, or electrophoretic signal characteristics of each STR locus as feature values; Using the aforementioned feature values as input and the source of the sample sea area as the label, a training dataset is constructed; A sea area source prediction model is obtained by training the training dataset using machine learning algorithms; The model was validated and optimized to improve prediction accuracy and stability.
[0023] Preferably, the machine learning algorithm includes one or more of the following: random forest model, gradient boosting model, or neural network model; more preferably, the machine learning algorithm adopts gradient boosting model.
[0024] This invention also provides a method for determining the marine origin of sand eels based on capillary electrophoresis (CE) data, comprising the following steps: Obtain the raw capillary electrophoresis data file (fsa file) of the sample to be tested; The fsa file is parsed to extract the fragment length and electrophoretic signal features of each STR site; The occurrence position and distribution of peak values at each point are detected according to the preset bp length range, and the occurrence position of the dominant peak is used as a feature to construct the STR feature vector; During feature construction, for STR sites where no effective amplification signal was detected, the position of their dominant peak was recorded as a preset placeholder value (e.g., -1), and a missing marker variable was introduced to characterize the signal loss at that site. The feature vector is input into a pre-trained machine learning model for calculation; Output the sea area origin prediction results and corresponding confidence information for the sample.
[0025] This invention also provides a method for predicting the marine origin of sea cucumbers based on high-throughput sequencing (NGS) data, comprising the following steps: Obtain STR sequencing data from the sample to be tested; Quality control and sequence analysis are performed on sequencing data to extract repeat count information or corresponding fragment length information for each STR locus; Collect and organize the information of each point, and construct the corresponding STR feature vector based on the repetition frequency or dominant length feature; For STR loci for which no valid sequencing results were obtained, the feature value was set to a preset placeholder value (e.g., -1), and a missing marker variable was introduced for identification. The feature vector is input into a pre-trained machine learning model for prediction; Output the sea area origin prediction results and corresponding confidence information for the sample.
[0026] The present invention has the following advantages: This invention is the first to develop a STR molecular marker primer set specifically for marine origin identification of Lepturacanthus savala. Compared with existing microsatellite markers used only for basic population genetics research of Lepturacanthus, the primer set of this invention has undergone multi-level screening and optimization, and can effectively distinguish Lepturacanthus samples originating from the Indian Ocean FAO 57.1 area and the Arabian Sea FAO 51.4 area, filling a technological gap in the field of marine origin tracing for cross-border trade of Lepturacanthus.
[0027] This invention constructs a multiplex PCR system that simultaneously amplifies three pairs of primers. By optimizing primer concentration (0.1~0.2 μM per pair), annealing temperature (59℃), and cycling parameters, simultaneous and efficient amplification of multiple STR loci is achieved. Compared with conventional single-pair PCR or gel electrophoresis detection in existing technologies, this invention significantly improves detection throughput, reduces sample volume (only 20~30 ng DNA) and reagent costs, and is suitable for rapid screening of large-scale trade samples.
[0028] This invention supports both capillary electrophoresis (CE) and next-generation sequencing (NGS) platforms, allowing users to choose according to their specific needs. The CE platform is suitable for rapid, low-cost routine testing, while the NGS platform provides higher resolution and can obtain precise repeat count information for STR loci. Data from both platforms can be uniformly converted into standardized feature vectors, achieving cross-platform data compatibility and model sharing, thus enhancing the applicability and flexibility of the technology.
[0029] This invention addresses the issue of potential overlap in the lengths of products amplified by three primer pairs. It employs three fluorescent groups—FAM, HEX, and ROX—to label the products in ascending order of length, ensuring that products with overlapping lengths can be distinguished by different fluorescence signals. This strategy effectively solves the signal confusion problem caused by product overlap in multiplex PCR, improving the accuracy and reliability of CE detection.
[0030] This invention designs standardized STR feature construction processes for both CE and NGS platforms: CE data generates feature vectors based on peak identification and allele determination; NGS data uses the weighted average of genotype repeat counts as feature values. Specifically, for STR loci where no valid signal was detected, this invention assigns a feature value of -1 and introduces a missing data marker variable, avoiding interference from missing data on model predictions and improving the robustness of the method.
[0031] This invention employs a gradient boosting machine learning model for marine origin prediction, and uses leave-one-out cross-validation (LOOCV) for training and evaluation. Validation results show that the model achieves 100% accuracy in distinguishing between sea cucumber samples originating from the Indian Ocean FAO57.1 and the Arabian Sea FAO 51.4. The model not only outputs the marine origin category but also the prediction confidence level, providing a quantifiable reliability indicator for the source tracing results.
[0032] This invention also provides a marine origin identification system for sand eels, comprising a data input module, a STR data parsing module, a feature vector construction module, a model prediction module, and a result output module. This system can directly import raw CE or NGS data, automatically complete feature construction and origin prediction, and output a marine origin and confidence report. It achieves full-chain automation from sample testing to result output, significantly reducing the barrier to manual operation and facilitating its widespread application in customs, quality inspection, and aquatic product trading enterprises.
[0033] The STR molecular marker primers of this invention can be used not only for marine origin identification of gudgeon, but also for genetic diversity detection and STR fingerprinting of gudgeon. Furthermore, the technical concept of this solution can be extended to the geographical tracing and population identification of other closely related fish species, demonstrating promising industrialization prospects and commercial value. Attached Figure Description
[0034] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary, and those skilled in the art can derive other embodiments based on the provided drawings without creative effort.
[0035] The structures, proportions, sizes, etc. illustrated in this specification are only for the purpose of assisting those skilled in the art in understanding and reading the content disclosed herein, and are not intended to limit the conditions under which the present invention can be implemented. Therefore, they have no substantial technical significance. Any modifications to the structure, changes in the proportions, or adjustments to the size, without affecting the effects and objectives that the present invention can produce, should still fall within the scope of the technical content disclosed in the present invention.
[0036] Figure 1 This is a flowchart of the method for predicting the source of sand gudgeon provided in Embodiment 1 of the present invention; Figure 2 This is a diagram showing the capillary electrophoresis (CE) detection results of STR sites provided in Embodiment 1 of the present invention; Figure 3The training sample gel image provided in Embodiment 1 of the present invention; Figure 4 This is a schematic diagram of the STR spectral capillary electrophoresis (CE) feature construction provided in Embodiment 1 of the present invention; Figure 5 This is a schematic diagram of STR phenotypic next-generation sequencing (NGS) feature construction provided in Embodiment 2 of the present invention; Figure 6 Principal component analysis (PCA) distribution of STR features of the NGS sample provided in Embodiment 3 of the present invention; Figure 7 Principal component analysis (PCA) distribution of STR features for the CE sample provided in Embodiment 3 of the present invention. Detailed Implementation
[0037] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0038] The TIANamp Genomic DNA Kit was purchased from Tiangen Biotech (Beijing) Co., Ltd.; 2×KAPA KK2602PCR Master Mix was purchased from Roche Kapa's product line; capillary electrophoresis was performed using an ABI 3730xl gene analyzer (Applied Biosystems, USA).
[0039] the term: STR short tandem repeat; CE capillary electrophoresis. NGS high-throughput sequencing (Next-Generation Sequencing). LOOCV (Leave-One-Out Cross Validation) PCA (Principal Component Analysis).
[0040] Example 1
[0041] The flowchart of the method for predicting the source of sand ribbonfish is as follows: Figure 1 As shown.
[0042] 1. Extraction of genomic DNA from sand ribbonfish The samples used in this embodiment were obtained from individual sand gudgeons from different sea areas, including the East Arabian Sea (FAO 51.4, Pakistani waters, LA-4258-250923-013) and the Northwest Indian Ocean (FAO 57.1, Indian waters, LS-4258-251027-024).
[0043] Muscle tissue samples were collected from various sea areas of *Gastrodon scutellarioides*. 10–50 mg of tail muscle tissue or pelvic fin samples were taken and processed under low-temperature conditions. Genomic DNA was extracted using an animal tissue genomic DNA extraction kit (such as the TIANamp Genomic DNA Kit or an equivalent kit) according to the manufacturer's instructions. The extraction steps included tissue lysis, proteinase K digestion, removal of proteins and other impurities, DNA purification and elution, ultimately yielding a genomic DNA solution.
[0044] The concentration and purity of the extracted DNA samples were determined using a NanoDrop micro-spectrophotometer. An OD260 / OD280 ratio between 1.8 and 2.0 indicated good DNA purity. The extracted genomic DNA was diluted to a working concentration (approximately 10–20 ng / μL) and stored at -20°C for later use.
[0045] To more accurately identify STR loci, sequencing data from multiple samples were integrated. Specifically, cleaned data from each sample were assembled from scratch to construct a population reference sequence. This leveraged the shared information from multiple samples to improve the coverage depth and assembly integrity of repetitive regions, reducing the risk of missing STR-related regions during assembly. The resulting population reference sequence was used for subsequent STR locus identification and primer design.
[0046] 2. STR locus screening and primer design A whole-genome scan of the *Gastrodon scutellario* genome was performed using SLAF-seq technology to identify tandem repeat regions composed of short repeat units (2–6 bp). A full-sequence scan of the reference sequence was then performed, and preliminary screening was conducted based on repeat unit length and repeat count, resulting in 208,168 candidate STR sites.
[0047] For each candidate site, target fragments containing the STR core region and its flanking sequences were extracted, and primers were designed. Primer design comprehensively considered factors such as sequence specificity, amplified product length, GC content, and primer interactions. After primer design analysis, effective primers corresponding to 46,789 sites were obtained.
[0048] Further cross-sample detectability screening (eliminating sites that could not obtain effective amplification information in all samples) yielded 12,030 sites. Subsequently, specificity screening and 3' end structural stability screening were performed (excluding primer pairs prone to 3' end dimerization and poor amplification stability), leaving 108 candidate sites. Based on capillary electrophoresis resolution requirements and amplified fragment length differences, 10 primer pairs were selected for multiplex PCR system construction. Finally, after system compatibility and amplification efficiency verification, the following three pairs of STR primers with high polymorphism and strong specificity were obtained; their specific sequence information is shown in Table 1.
[0049] Table 1. STR primer nucleotide sequences For capillary electrophoresis detection, the 5' end of the forward primer of each primer pair is labeled with a fluorescent group. Based on the length range of the amplified products of each primer pair (LS4:CA8 approximately 130~160 bp, LS6:AC16 approximately 110~140 bp, LS8:AC16 approximately 150~200 bp), the three fluorescent groups FAM, HEX, and ROX are used in rotation in ascending order to ensure that even if there is a possibility of overlap in the product length range, they can still be effectively distinguished by different fluorescence signals.
[0050] After the above multi-level screening (from the 208,168 candidate STR sites obtained in the initial screening, the screening was carried out based on cross-sample detectability, specificity, 3' end structural stability, amplified fragment length and CE resolution, system compatibility, etc.), the three pairs of STR primer combinations with high polymorphism and high specificity shown in Table 1 were finally obtained.
[0051] 3. Construction and optimization of multiplex PCR amplification system Using the genomic DNA extracted in step 1 as a template, multiplex PCR amplification was performed using the three primer pairs listed in Table 1. The total PCR volume was 20 μL, and the specific components and final concentrations are shown in Table 2.
[0052] Table 2 PCR amplification reaction system The 2×KAPA KK2602 PCR Master Mix already contains Mg. 2+ And dNTPs, Mg in the final reaction system 2 + The final concentration was 1.5 mM, and the final concentration of each dNTP was 200 μM. The PCR reaction procedure is as follows: Pre-denaturation at 95℃ for 3 minutes; Depolymerize at 95℃ for 20 seconds, anneal at 59℃ for 15 seconds, extend at 72℃ for 30 seconds, repeat 30 cycles; Finally, extend the heat to 72°C for 1 minute, then store at 4°C.
[0053] The optimized multiplex PCR system enables simultaneous and efficient amplification of three STR loci in the same reaction system. The amplification products were verified by 2% agarose gel electrophoresis, yielding clear amplification bands within the expected fragment size range, without non-specific amplification or primer dimer interference (e.g., Figure 3 (As shown).
[0054] Furthermore, the relative abundance of different primer pairs can be fine-tuned based on the differences in amplification efficiency exhibited by different primer pairs in the actual PCR reaction system, in order to reduce the systematic bias caused by multiplex PCR competition introduced in the experimental process.
[0055] 4. STR detection and feature construction based on capillary electrophoresis (CE) 1) Fluorescent labeling strategy: The amplification product lengths of the three primer pairs in this invention are as follows: LS4:CA8 approximately 130~160 bp, LS6:AC16 approximately 110~140 bp, and LS8:AC16 approximately 150~200 bp. Since the product length ranges of LS4:CA8 and LS6:AC16 may overlap (both are in the 130~140 bp range), without differentiation, the site assignment of the amplified product cannot be accurately determined during capillary electrophoresis.
[0056] To address this issue, this invention employs a rotating labeling strategy using three fluorescent groups—FAM, HEX, and ROX—according to the amplified product fragment lengths from shortest to longest: LS6:AC16 (shortest product) is labeled with FAM, LS4:CA8 (medium-length product) with HEX, and LS8:AC16 (longest product) with ROX. Even if product lengths overlap, the different wavelengths of fluorescence signals from different sites allow for effective differentiation through fluorescence color. This strategy effectively solves the signal confusion problem caused by product overlap in multiplex PCR, improving the accuracy and reliability of CE detection.
[0057] 2) Perform capillary electrophoresis on the PCR amplification products obtained in step 3: Take 1-2 μL of PCR amplification product, mix it with 0.3 μL of fluorescently labeled LIZ internal standard and 9.5 μL of Hi-Di formamide, denature at 95℃ for 3 minutes, and then cool on ice. Use an ABI 3730xl gene analyzer for capillary electrophoresis separation and detection. Obtain fragment length information and peak data corresponding to each STR site by detecting fluorescence signals. An example of capillary electrophoresis detection results is shown below. Figure 2 As shown.
[0058] 3) Feature Construction: The process for constructing STR feature vectors based on capillary electrophoresis data is as follows: Read the raw capillary electrophoresis signal → Calibrate the baseline → Calibrate the LIZ internal standard → Detect peak values → Assign STR loci → Determine alleles → Generate sample STR feature vectors (see...) Figure 4 The amplification results of each sample at the three STR sites are shown in Table 3, including information such as dominant genotype, weighted average repeat count, and fragment length distribution.
[0059] Table 3. Genotyping results of sand gudgeon samples from different sea areas at three STR loci. The results indicate that the dominant genotype at the LS4:CA8 locus was (CA)~14~ or (CA)~15~ in samples from Pakistani waters (FAO 51.4), while the dominant genotype at this locus was (CA)~8~ in samples from Indian waters (FAO 57.1). At the LS6:AC16 locus, the dominant genotype was (AC)~7~ in samples from Pakistani waters and (AC)~13~~(AC)~16~ in samples from Indian waters. At the LS8:AC16 locus, the dominant genotype was (AC)~12~ in samples from Pakistani waters and (AC)~16~~(AC)~20~ in samples from Indian waters. These results demonstrate that significant genotypic differences were observed at all three STR loci in samples from both sea areas, providing a molecular basis for sea area origin identification.
[0060] For STR sites where no effective amplification signal was detected, their eigenvalues were assigned a value of -1 and a missing marker variable was added to ensure the integrity of the data analysis.
[0061] Example 2
[0062] STR detection and feature construction based on high-throughput sequencing (NGS): The PCR amplification products obtained in Example 1 were used to construct a high-throughput sequencing library. Paired-end 150 bp sequencing was performed using the Illumina MiSeq sequencing platform to obtain raw sequencing data. The raw sequencing data underwent quality control: adapter sequences and low-quality sequences were removed using Trimmomatic software, and bases with quality values below Q20 were filtered to obtain high-quality cleaned data.
[0063] The workflow for constructing STR feature vectors based on high-throughput sequencing data is as follows: quality control of raw sequencing data to remove contaminating, low-quality, and adapter sequences → assembly of paired-end sequencing sequences → statistical analysis of the distribution of different genotypes at each STR locus → calculation of the weighted average of the genotype repeat counts at each STR locus as the single-locus feature value → combination of all locus feature values to obtain the STR feature vector (see...). Figure 5 ).
[0064] The method for constructing STR feature vectors based on high-throughput sequencing data is as follows: For each STR locus, the distribution of repeat counts of all sequencing reads at that locus is statistically analyzed. For samples with multiple alleles, the mean repeat count is calculated based on the frequency weights of each allele and used as the feature value for that locus. For STR loci without valid sequencing data, the feature value is assigned as -1.
[0065] The feature values of the three STR sites are combined into a three-dimensional feature vector for subsequent marine origin prediction analysis.
[0066] Example 3
[0067] Based on the STR feature data of gudgeon obtained in Examples 1 and 2, a marine origin prediction model was constructed. The training dataset contains gudgeon samples from two marine origins (FAO 51.4 and FAO 57.1). Feature vectors are composed of feature values of three STR loci for each sample, with marine origin as the category label.
[0068] Three machine learning models (logistic regression, random forest, and gradient boosting) were trained and their performance compared. Due to the high cost of obtaining source-labeled marine samples and the relatively limited scale of training data, leave-one-out cross-validation (LOOCV) was used to evaluate the performance of each model. LOOCV involves retaining only one sample as the test set during each training iteration, using the remaining samples for model training. This maximizes the utilization of limited sample information while verifying the generalization ability of each model individually, resulting in more stable and reliable evaluation results.
[0069] The results are shown in Table 4.
[0070] Table 4. Performance comparison of different models in marine origin classification tasks. As shown in Table 4, under the constructed STR feature conditions, the Gradient Boosting model achieved the best classification performance on this dataset, with a leave-one-out cross-validation accuracy of 100%. The Random Forest model achieved an accuracy of only 50%, and the Logistic Regression model achieved an accuracy of 83%. Considering the model's adaptability to nonlinear feature combinations and the need for subsequent expansion to multi-oceanic area classification tasks, the Gradient Boosting model was ultimately selected as the preferred model for ocean origin tracing analysis. The Logistic Regression model is suitable for linearly separable problems and has good interpretability in binary classification tasks, but its ability to express complex nonlinear feature combinations is limited. Although the Random Forest model has a certain nonlinear fitting ability, it exhibits relatively low stability on this dataset. Therefore, the Gradient Boosting model is more suitable for the multidimensional STR nonlinear feature classification problem constructed in this invention and has better scalability.
[0071] Principal component analysis (PCA) was further used to perform dimensionality reduction and visualization analysis on the STR eigenvectors (see [link]). Figure 6 , Figure 7 The results showed that the sand gudgeon samples from FAO 51.4 (Pakistani waters) and FAO 57.1 (Indian waters) exhibited a clear separation trend in the PCA projection space, indicating that the three STR loci screened in this invention have sufficient resolution to effectively distinguish sand gudgeon populations from different sea areas.
[0072] Example 4
[0073] This invention also provides a marine origin identification system for sand gudgeon, which is equipped with the identification methods described in Examples 1-4 and includes the following functional modules: (1) Data input module: Supports importing raw data files of capillary electrophoresis (CE) (.fsa format) or raw data files of high-throughput sequencing (NGS) (.fastq format); (2) STR data parsing module: Automatically parses CE peak data or NGS sequencing data, extracts fragment length information, peak information or repeat number information of each STR site, and automatically marks sites with no effective signal as -1; (3) Feature vector construction module: Following the process described in Example 1 or Example 2, the original data is automatically converted into a standardized STR feature vector format; (4) Model prediction module: Load the pre-trained gradient boosting model, input the feature vector, and output the predicted sea area origin category and confidence level; (5) Results output module: Generates a complete traceability report containing sample number, STR locus typing results, predicted marine origin, confidence level and test report number, and supports export in PDF / Excel format.
[0074] This system automates the entire process from raw data import to result report output, significantly reducing the barrier to manual operation, improving testing efficiency and result consistency, and facilitating its widespread application in customs, quality inspection agencies, and aquatic product trading companies. The system has a user-friendly interface suitable for non-professional users, and its modular design facilitates subsequent maintenance and functional expansion.
[0075] Example 5
[0076] To support the training and continuous optimization of the marine origin prediction model for *Gymnocypris spp.*, this invention constructs a marine origin tracing database for *Gymnocypris spp.*. This database contains phenotypic information of *Gymnocypris spp.* samples from known marine origins (FAO 51.4 and FAO 57.1) at three STR loci (LS4:CA8, LS6:AC16, LS8:AC16), including: Each sample's serial number, marine origin, and collection information; Dominant genotypes, weighted average number of repeats, and fragment length range for each STR locus; Raw peak data from capillary electrophoresis (fsa file) or raw data from high-throughput sequencing (fastq file).
[0077] The database uses a relational database (such as MySQL) for storage and management, and implements data creation, deletion, modification, and querying through a standard interface. This database serves two purposes: firstly, it constructs a standard spectral reference set, providing benchmark data for subsequent marine origin prediction; secondly, it serves as a training dataset, periodically retraining and optimizing the prediction model to continuously improve its accuracy and generalization ability. For newly determined sand eel samples, their STR spectral data can be included in the database, enabling dynamic data expansion and updates.
[0078] Specifically, the algorithm flow is as follows: enter: Test sample data (CE or NGS) The trained classification model M Output: Marine source prediction result Y Confidence level P I. STR locus screening and feature system construction (offline stage) Step 1: Construction of reference sequence Collect sequencing data from multiple known samples Perform quality control and noise reduction A population reference sequence Ref is constructed using an ab initio assembly method. Step 2: STR locus mining Scan the cascaded repeating region in Ref. Filter regions with repeating unit lengths of 2~6bp Obtain the initial STR candidate set S0 Step 3: Primer Design Flanking sequences were extracted from each STR site in S0. Design primer pairs Obtain the expandable STR set S1 Step 4: Multi-level filtering 1) Cross-sample detectability screening → S2 2) Specific screening → S3 3) 3' end structural stability screening → S4 4) Screening based on amplified fragment length and CE resolution → S5 5) System compatibility screening → S6 The final core STR marker set S_final (e.g., 3 primer pairs) is obtained. II. STR Feature Extraction and Marine Source Tracing Prediction (Online Stage) Step 5: Sample STR testing CE or NGS testing of the sample to be tested Obtain raw STR data D Step 6: Feature Construction For each STR locus: If it is NGS data: Extracting the number of repetitions or the length of the fragment If it is CE data: Extract the peak value corresponding to the bp position Constructing the feature vector X If a certain site is missing: Assign a value of -1 and record the missing flag. Step 7: Model Prediction Input the feature vector X into the model M Output the predicted category Y and probability P Return Y, P Although the present invention has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of the present invention fall within the scope of protection claimed by the present invention.
Claims
1. A molecular marker primer for STR in sand eels, characterized in that, The primer combination consists of the following 3 pairs of primers: (1) LS4:CA8, its primer sequence is: SEQ ID NO.1: Forward primer: CAACAGCAGCAACAGGTTCC; SEQ ID NO.2: Reverse primer: GAGTGAGTGAGTGCGTGTGA; (2) LS6:AC16, its primer sequence is: SEQ ID NO.3: Forward primer: CAGCAAACAAAGAGTCGCGT; SEQ ID NO.4: Reverse primer: TGTTCCCTCCATCTGTCCCA; (3) LS8:AC16, its primer sequence is: SEQ ID NO.5: Forward primer: TTCGGCCCTTTCATGTGAGA; SEQ ID NO.6: Reverse primer: ACGCCACCATTTTCCAGACT.
2. The application of the STR molecular marker primer for sea cucumber as described in claim 1 in the detection of genetic diversity of sea cucumber, construction of STR fingerprint profiles and / or identification of sea cucumber origin.
3. A method for identifying the marine origin of *Gymnocypris spp.* based on the STR molecular marker primers for *Gymnocypris spp.* as described in claim 1, characterized in that... Includes the following steps: (1) Extract genomic DNA from the sand ribbonfish to be tested; (2) Using the DNA obtained in step (1) as a template, a primer mixture is prepared using the three pairs of primers in claim 1, and a PCR reaction system is prepared for multiplex PCR amplification to obtain the amplification product; (3) Capillary electrophoresis was used to detect the amplification products to obtain raw capillary electrophoresis data, and / or second-generation high-throughput sequencing was used to sequence the amplification products to obtain sequencing data; (4) Construct corresponding sample STR feature vectors based on capillary electrophoresis data or next-generation sequencing data respectively; (5) Using the pre-trained model, the marine origin of the sand ribbonfish is predicted and determined, and the marine origin result and prediction confidence are output.
4. The method for identifying the marine origin of sand ribbonfish according to claim 3, characterized in that, In step (2), the total PCR system is 20 μL: 10 μL of 2×KAPA KK2602 PCR Master Mix, 2.0 μL of premixed primer mixture, 1.5–2.0 μL of template DNA (total DNA 20–30 ng), and the remaining ddH2O is added to bring the total to 20 μL; the final concentration of each primer pair in the system is 0.1–0.2 μM, and the Mg in the Master Mix is [missing information]. 2+ Final concentration 1.5 mM, final concentration of each dNTP 200 μM.
5. The method for identifying the marine origin of the sand ribbonfish according to claim 3, characterized in that, The PCR amplification program for step (2) is as follows: pre-denaturation at 95℃ for 3 min; denaturation at 95℃ for 20 s, annealing at 59℃ for 15 s, extension at 72℃ for 30 s, for 30 cycles; final extension at 72℃ for 1 min, and storage at 4℃.
6. The method for identifying the marine origin of the sand ribbonfish according to claim 3, characterized in that, The steps for constructing STR feature vectors using capillary electrophoresis data are as follows: read the raw capillary electrophoresis signal → baseline calibration → LIZ internal standard calibration → peak identification → STR locus allocation → allele determination → generate sample STR feature vectors; when there is no effective amplification signal at a locus, the corresponding feature is assigned a value of -1 and a missing marker variable is added.
7. The method for identifying the marine origin of sand ribbonfish according to claim 3, characterized in that, The steps for constructing STR feature vectors using second-generation sequencing data are as follows: quality control of raw sequencing data to remove contaminating sequences, low-quality sequences, and adapter sequences → assembly of paired-end sequencing sequences → statistical analysis of the distribution of different genotypes at each STR locus → calculation of the weighted average of the number of genotype repeats at each STR locus as the feature value of a single locus → combination of all locus feature values to obtain the STR feature vector; when there is no valid sequencing data for a locus, the corresponding feature is assigned a value of -1 and a missing marker variable is added.
8. The method for identifying the marine origin of sand ribbonfish according to claim 3, characterized in that, In step (3), the primers used for capillary electrophoresis detection carry fluorescent labels at their 5' ends; according to the order of amplified product fragment length from shortest to longest, three fluorescent labels, FAM, HEX and ROX, are used in rotation, and primers with overlapping fragment lengths are distinguished by different fluorescent labels.
9. A marine origin identification system for sand ribbonfish, characterized in that, The system is equipped with the identification method described in any one of claims 3-8, including a data input module, a STR data parsing module, a feature vector construction module, a model prediction module, and a result output module; the system imports the raw CE or NGS data of the sample to be tested, automatically completes feature construction and origin prediction, and outputs a report on the marine origin and confidence level of the sand ribbonfish.