Bird identification universal primer design method based on machine learning optimization and application

By using a primer design method optimized by machine learning, combined with Shannon entropy to determine conserved bases and a multi-dimensional scoring system, the problems of inaccurate determination of conserved bases and low screening efficiency in general bird primer design are solved, achieving efficient and accurate bird identification, which is applicable to a wide range of bird species identification.

CN121709031APending Publication Date: 2026-03-20TIANJIN INST OF CRIMINAL SCI & TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511915707.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-18
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing universal primer designs for birds suffer from inaccurate conserved base determination, low screening efficiency, and poor overall performance, resulting in low amplification efficiency and poor identification accuracy, making it difficult to meet the screening needs of a massive number of candidate primers.

Method used

A primer design method based on machine learning optimization was adopted. Conserved bases were determined by Shannon entropy, a multi-dimensional scoring system was established, and candidate primers were screened using machine learning models to construct primer combinations and optimize primer performance.

Benefits of technology

It improves the universality and specificity of primers, amplification stability, and enhances the efficiency and accuracy of bird identification. It is suitable for the detection of trace and degraded samples and covers a wide range of bird species.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121709031A_ABST
    Figure CN121709031A_ABST
Patent Text Reader

Abstract

The invention discloses a design method and application of a universal primer for bird identification based on machine learning optimization, the method is based on a DNA bar code technology, and the method judges conservative bases through multiple comparison in combination with Shannon entropy, so that the accuracy of obtaining conservative sequences is improved; establishing a multi-dimensional primer scoring system covering a coverage degree, a Tm value, an interaction degree, a degenerate base number and a poly structure; training a random forest machine learning model by using known primer sequences and score data, and optimizing model parameters through grid search and cross validation; and finally, quickly evaluating and screening a large number of candidate primers by using the trained model to obtain a high-performance primer combination. Traditional manual screening is replaced with a machine learning algorithm, the primer design efficiency and screening precision are remarkably improved, the designed universal primer is wide in coverage, high in specificity and high in amplification stability and can be widely applied to bird species identification, and reliable technical support is provided for ecological monitoring, species protection and law enforcement detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of bird species identification technology and involves the intersection of molecular biology and artificial intelligence. Specifically, it is a method and application of a universal primer design for bird identification based on machine learning optimization, and the application of primer combinations designed by this method in bird species identification, ecological monitoring, biodiversity assessment, species protection and law enforcement detection. Background Technology

[0002] Birds are a key component of ecosystems, and their species identification is of great importance in ecological monitoring and surveys, biodiversity conservation, border quarantine, combating illegal hunting, and investigating bird strikes on aircraft. As a very large group, bird species identification is challenging.

[0003] Traditional methods for species differentiation rely on morphology, which are easily limited by sample condition and subjective judgment. With the development of molecular biology, species differentiation based on DNA sequence has gradually become mainstream, with DNA barcoding being a relatively mature method. A DNA barcode is a short, standardized gene sequence that is sufficiently conserved within a species but varies considerably between species, thus allowing for species differentiation. Commonly used DNA barcode markers for birds include the COXI gene (cytochrome oxidase I gene), the CYTB gene (cytochrome b gene), the 12S rRNA gene, and the 16S rRNA gene. The design of universal primers is crucial for the application of DNA barcoding technology; their performance directly determines amplification efficiency, species coverage, and identification accuracy.

[0004] In the existing technology, there are many shortcomings of universal primers for bird species identification: (1) The method of conserved base determination is unreasonable. Traditional methods screen conserved bases by using fixed allele frequency thresholds, which is easily affected by the number of samples, leading to deviations in the identification of conserved sequences, and thus affecting the universality and specificity of primers; (2) Primer screening relies on manual operation. Researchers usually only manually screen a limited number of primers and verify them one by one using multiple evaluation indicators, which is inefficient and cannot fully encompass all candidate primers, nor can it meet the screening needs of a large number of candidate primers; (3) Universal primers generally amplify longer fragments and have poor tolerance to degraded samples, but at the same time, shorter amplification regions will reduce the identification accuracy. It is necessary to balance the length of amplicon and the number of universal primers.

[0005] Although existing studies have attempted to optimize primer design processes, a complete scheme combining objective conserved base identification, multi-dimensional comprehensive evaluation, and efficient intelligent screening has yet to be established. Therefore, developing a design method capable of accurately identifying conserved sequences, systematically evaluating primer performance, and rapidly screening for optimal primers is of great significance for improving the practicality and efficiency of avian DNA barcoding identification technology. Summary of the Invention

[0006] Technical problem to be solved: In order to overcome the shortcomings of existing technologies and solve the problems of inaccurate determination of conserved bases, low screening efficiency and poor overall performance in existing bird universal primer designs, this invention provides a method and application for bird identification universal primer design based on machine learning optimization; so as to achieve accurate identification of bird conserved sequences, systematic evaluation of primer performance and efficient intelligent screening, and obtain bird identification universal primers with wide coverage, high specificity and stable amplification.

[0007] Technical solution: A machine learning-optimized universal primer design method for bird identification, comprising the following steps: S1. Download avian COXI gene sequences from public databases and construct a local database of avian target sequences; S2. Determine conserved bases to identify conserved sequences. Use multiple sequence alignment software (such as MEGA) to perform multiple alignments on sequences in the local database and generate alignment files. Conserved bases are determined based on the following principles to obtain conserved sequences: Calculate the Shannon entropy at each base position. When the Shannon entropy is ≤0.5, the conserved base is the base with the highest allele frequency. When the Shannon entropy is >0.5, the base with an allele frequency ≥0.2 is a conserved base. When two or more bases meet the conditions, the corresponding degenerate base representation is used. The representation of degenerate bases follows the IUPAC rules. The formula for calculating Shannon entropy is as follows: ; Where pi is the allele frequency of the i-th base, and k is the number of base types at that site; S3. Establish a multi-dimensional primer scoring system with a total score of 100 points, including coverage score, Tm score, primer interaction score, degenerate base number score, and sequence poly structure score. The weights of the above scores are 40%, 20%, 20%, 10%, and 10%, respectively. S4. Machine learning model training: Primer dataset is formed by using public databases, literature, or self-designed primers. The primer dataset is expanded by primer reverse complementation. The score of each primer is labeled according to the scoring system in S3. The model is trained using the random forest algorithm of the sklearn module. The optimal parameters are obtained through grid search and cross-validation and the model is saved. S5. Candidate primer screening and combination: Based on the conserved sequences in S2, a large number of candidate primers are generated. After evaluation by the model trained in S4, primers with high scores and amplicon lengths of 100-300bp are selected to construct primer combinations.

[0008] Preferably, in S1, the publicly available databases include NCBI (https: / / www.ncbi.nlm.nih.gov / ) and BOLD (https: / / boldsystems.org / ). The downloaded sequences are cleaned and processed as necessary, including: removing low-quality sequences with more than 25% of ambiguous base N; performing BLAST alignment with the reference genome GCF_016699485.2 (species: chicken) to remove sequences with similarity less than 50% (removing erroneous sequences that are not avian or are not the target).

[0009] Preferably, in S3, the scores for each item in the scoring system are calculated using the following method: Coverage score refers to the proportion of primers that match the species sequences in the local database. Matching here means that the sequence is exactly matched without mismatches or insertions / deletions. Primers containing degenerate bases are broken down into different bases and matched separately to calculate the total matching proportion. The score is calculated by multiplying the matching proportion by 40. Tm score refers to scoring based on the primer Tm value. Primers containing degenerate bases are broken down into different bases and the average Tm value of the primers is calculated. 57℃<Tm<63℃ gets 20 points, 55℃≤Tm≤57℃ or 63℃≤Tm<65℃ gets 15 points, 53℃≤Tm<55℃ or 65℃≤Tm≤67℃ gets 10 points, and Tm<53℃ or Tm>67℃ gets 0 points. Primer interaction scoring: Primer interaction includes primers forming dimers or forming hairpin structures. Scoring is based on the average Tm value of the interacting sequences. For primers containing degenerate bases, the average Tm value of the interacting sequences is calculated by breaking them down into different bases. Tm < 45℃ gets 20 points, 45℃ ≤ Tm ≤ 50℃ gets 15 points, 50℃ < Tm ≤ 55℃ gets 10 points, and Tm > 55℃ gets 0 points. Degenerate base count score: The primer is scored based on the number of degenerate bases it contains. 0 bases get 10 points, 1 base gets 7 points, 2 bases get 4 points, and more than 2 bases get 0 points. Sequence poly structure scoring: Scoring is based on the number of consecutive bases contained in the primer. No points are deducted for ≤5 consecutive bases, 5 points are deducted for 6 consecutive bases, 10 points are deducted for 7 or more consecutive bases, and points are deducted cumulatively for containing 2 or more consecutive bases, until the score reaches 0.

[0010] The general primer combination for bird identification designed by any of the methods described above.

[0011] Preferably, the primer combination includes the following primer sequences: Cox1bird1, sequence SEQ ID NO:1-2; Cox1bird2, sequence SEQ ID NO:3-4; Cox1bird3, sequence SEQ ID NO:5-6; Cox1bird4, sequence SEQ ID NO:7-8.

[0012] Furthermore, the primer sequences 5'-3' are specifically as follows: SEQ ID NO:1 is TTCTCAACCAACCACAAAGA, SEQ ID NO:2 is GKGGGAATGCTATGTCKGGG; SEQ ID NO:3 is TTCTGATTCTTYGGMCACCC, SEQ ID NO:4 is ACTGTRAAYATGTGGTGGGC; SEQ ID NO:5 is TTYACCCACTGATTCCCMCT, SEQ ID NO:6 is GKCCTAGGAAGTGTTGKGGG; SEQ ID NO:7 is TRCCACGACGATACTCAGAC; SEQ ID NO:8 is GCAGCCGTGGATTCATTCRA.

[0013] Preferably, the primer combination is amplified according to the following PCR procedure: 95℃ for 5 min; 95℃ for 30 s, 58℃ for 40 s, 72℃ for 30 s, 32 cycles; 72℃ for 5 min; and 4℃ at a constant temperature.

[0014] The application of the primer combinations described above in the preparation of bird species identification kits.

[0015] Preferably, the kit includes: primer combination, DNA extraction reagent, PCR amplification reagent and sequencing reagent.

[0016] Preferably, the kit is used for bird species identification, bird ecological monitoring, bird biodiversity assessment, bird species protection, or law enforcement detection.

[0017] Beneficial effects: (1) This application provides a method for designing universal primers for bird identification based on machine learning optimization. This method combines Shannon entropy and base frequency to determine conserved bases. Traditional methods use fixed allele frequency thresholds to screen conserved bases, which are easily affected by the number of samples, leading to deviations in the identification of conserved sequences, and thus affecting the universality and specificity of primers. This design method introduces Shannon entropy to determine conserved bases, which can effectively reduce the impact of the number of samples, improve the accuracy of conserved sequence identification, and thus enhance the universality of primers.

[0018] (2) This application provides a machine learning-optimized general primer design method for bird identification. This method effectively avoids the problems of traditional primer design methods, which rely on manual selection and require researchers to verify each primer based on multiple evaluation indicators, resulting in low efficiency, inability to comprehensively cover all candidate primers, and difficulty in handling the screening needs of a large number of candidate primers. This method uses machine learning algorithms to construct a primer evaluation model, allowing researchers to consider as many candidate primers as possible, quickly evaluate primers, avoid verifying each multiple indicator one by one, and significantly improve primer design efficiency.

[0019] (3) This application provides a method for designing universal primers for bird identification based on machine learning optimization. This method establishes a multi-dimensional primer scoring system, comprehensively considers key indicators such as coverage, Tm value, and interaction degree, and combines the accurate prediction of the machine learning model to select primers with high universality, strong specificity and stable amplification efficiency.

[0020] (4) The second aspect of this application provides a universal primer for bird identification based on machine learning optimization. The primers cover a wide range of bird species and are universally applicable in bird identification. The primers amplify fragments smaller than 300 bp, are highly tolerant to the detection of trace samples and degraded samples, and the amplified products can be adapted to first-generation and second-generation sequencing technologies for detection. The primers amplify multiple target fragments, improve the amplification success rate, and effectively improve the accuracy of bird species identification. Attached Figure Description

[0021] Figure 1 This is a comparison between the primer score predictions and the actual values ​​obtained from training the COXI target model in Example 1. Figure 2 This is a comparison of the coefficient of determination (R²) values ​​under different parameter combinations during the training of the COXI target model in Example 1.

[0023] COXI sequences of birds were downloaded from public databases. Bird COXI sequences were downloaded in batches using the BOLD API (https: / / boldsystems.org / data / api / ) and NCBI Datasets CLI (https: / / www.ncbi.nlm.nih.gov / datasets / ). Low-quality sequences with more than 25% ambiguous bases were removed; BLAST alignment with the reference genome (GCF_016699485.2, chicken species) was performed, and sequences with less than 50% similarity were removed (non-avian or non-target erroneous sequences were removed). A total of 2563 high-quality COXI sequences from 949 bird species were obtained.

[0024] MEGA was used for multiple sequence alignment, and a Python script was used to determine conserved bases. Since the determination of conserved bases is based on base frequency, a fixed threshold is typically used, which is easily affected by sample size. Using Shannon entropy can effectively reduce this influence. The Shannon entropy of each base is calculated. If the Shannon entropy is ≤0.5, the conserved base is the base with the highest allele frequency. If the Shannon entropy is >0.5, then the base with an allele frequency ≥0.2 is considered a conserved base. When multiple bases meet the criteria, degenerate bases are used, and the degenerate base representation follows IUPAC rules. The Shannon entropy calculation formula is as follows: , where pi is the allele frequency of the i-th base and k is the number of base types at that site.

[0025] A multi-dimensional primer scoring system was established, with a total score (100 points) = Coverage score (40% weight) + Tm score (20% weight) + Primer interaction score (20% weight) + Degenerate base number score (10% weight) + Sequence poly (10% weight). Each score was calculated according to the following rules: (1) Coverage score: refers to the proportion of primers that match the species sequences in the local database. Matching here means exact sequence matching, without mismatches or insertions / deletions. Primers containing degenerate bases need to be broken down into different bases and matched separately to calculate the total matching ratio. The score is calculated by multiplying the matching ratio by 40. (2) Tm score: The score is based on the primer Tm value. Primers containing degenerate bases are broken down into different bases and the average Tm value of the primers is calculated. 57℃<Tm<63℃ gets 20 points, 55℃≤Tm≤57℃ or 63℃≤Tm<65℃ gets 15 points, 53℃≤Tm<55℃ or 65℃≤Tm≤67℃ gets 10 points, and Tm<53℃ or Tm>67℃ gets 0 points. (3) Primer interaction score: Primer interaction includes primers forming dimers or forming hairpin structures. The score is based on the average Tm value of the interaction sequence. For primers containing degenerate bases, the average Tm value of the interaction sequence is calculated by breaking them down into different bases. Tm < 45℃ gets 20 points, 45℃ ≤ Tm ≤ 50℃ gets 15 points, 50℃ < Tm ≤ 55℃ gets 10 points, and Tm > 55℃ gets 0 points. (4) Degenerate base number score: The primer is scored according to the number of degenerate bases it contains. 0 bases get 10 points, 1 base gets 7 points, 2 bases get 4 points, and more than 2 bases get 0 points. (5) Sequence poly structure scoring: Scoring is based on the number of consecutive bases contained in the primer. No points are deducted for ≤5 consecutive bases, 5 points are deducted for 6 consecutive bases, 10 points are deducted for 7 or more consecutive bases, and the points are deducted cumulatively for multiple consecutive bases, with a minimum of 0 points.

[0026] After determining the multi-dimensional scoring system for primers, a machine learning model was constructed and trained. Primer datasets were collected using public databases (such as the BOLD database), literature, and several self-designed primers. The dataset was expanded by reverse complementation of primers, resulting in a final training dataset for the COXI target containing 400 primers. Each primer was scored according to the aforementioned multi-dimensional scoring system, and features were extracted from the primer sequences. Each base was mapped to a 4-dimensional vector; for example, A bases were mapped to "1, 0, 0, 0", T bases to "0, 1, 0, 0", and degenerate R bases to "0.5, 0, 0, 0.5". The base mapping rules are listed in Table 1. Table 1. Base mapping rules for model training The training data was divided into an 80% training set and a 20% test set. Using Python's sklearn module, a Random Forest algorithm model was built and trained. Optimal parameters were obtained through grid search and cross-validation, and the model was saved. The Random Forest algorithm consists of multiple independent decision trees. The results of individual decision trees are integrated using a "voting / averaging" method to solve classification or regression problems. Its core advantages are resistance to overfitting, strong generalization ability, and the ability to handle high-dimensional feature data, such as primer sequences themselves. Multiple permutations and combinations of key parameters for model construction (such as "tree depth" and "number of trees") were performed. Five-fold cross-validation combined with grid search was used to traverse all parameter combinations. The coefficient of determination (R²) was used as the evaluation metric to select the optimal parameters, and finally, the best model was saved.

[0027] The results of the primer evaluation training model for COXI targets are shown below. Figure 1 and Figure 2 . Figure 1 The scatter plot shows the true and predicted values ​​of the optimal model on the test set. The model's coefficient of determination (R²) is 0.3338, and its mean squared error (MSE) is 99.4494, indicating that the model can explain approximately 33.38% of the score variation. The average error of a single prediction is approximately 9.97 points (√(2&99.4494)≈99.97). The model can initially capture the basic trend of primer scores, with larger prediction errors for high-scoring primers, which is related to the lack of high-scoring primers in the dataset. Figure 2 The results show the 5-fold cross-validation R² scores of the top 20 optimal parameter combinations in the grid search. The optimal model has the parameters "max_depth=None, min_samples_leaf=2, min_samples_split=2, n_estimators=200".

[0028] A large number of candidate primers of suitable length are generated based on conserved sequences. For example, 20 bp primers can be selected starting from the first base position, 20 bp primers starting from the second base position, and then reverse complementation can be performed. This process can yield thousands of candidate primers (including forward and reverse primers). Using the primer sequences as model input, a pre-trained target primer scoring model is used to quickly output the comprehensive score of the primers. Primers with high scores are selected to construct primer pairs, while also considering amplicon length of 100-300 bp and low interaction between primers, thus constructing primer combinations. The selected primer combinations are as follows: Table 2. Universal primers for bird species identification

[0029] The performance of the universal primers for bird identification in Table 2 of Example 1 of this invention was evaluated. These primers received high scores in the machine learning model. Their actual Tm values ​​were calculated to check whether the Tm values ​​met the general primer amplification requirements. For primers containing degenerate bases, the average Tm value corresponding to different bases was calculated.

[0030] Simultaneously, the NCBI Primer-Blast web-based tool was used to verify the amplification coverage of bird species by the primers. Since NCBI Primer-Blast does not recognize degenerate bases, degenerate bases were replaced with any one of their corresponding bases for testing; mismatches were allowed during blotting. In the NCBI Primer-Blast web-based tool settings, "Database" was set to "nt," "Organism" was left blank (meaning it would be compared with all species in the database), "Max number of sequences returned by Blast" was set to 50,000, and "Max targetamplicon size" was set to 400. The default settings were used for base mismatches, allowing a maximum of two mismatches. The number of amplifiable bird species returned by NCBI Primer-Blast was counted.

[0031] The statistical results are listed in Table 3. The universal primers selected based on model scoring have good Tm values ​​and may cover a wide range of bird species.

[0032] Table 3. Model scores, actual Tm values, and number of amplifiable bird species for universal avian primers.

[0033] DNA was extracted from common bird species in the Beijing-Tianjin-Hebei region. A total of 8 bird species were identified and are listed in Table 4.

[0034] Prepare reagents according to the following system: total reaction volume 50ul; Taq DNA polymerase (Takara): 25ul; forward primer (5uM), 4ul; reverse primer (5uM), 4ul; DNA template (10-20bg / ul), 2ul; add sterile water to 50ul.

[0035] Amplification was performed according to the following procedure: 95℃ for 5 min; 95℃ for 30 s, 58℃ for 40 s, 72℃ for 30 s, 32 cycles; 72℃ for 5 min; and isothermal at 4℃.

[0036] The amplified products were sent to a sequencing company for next-generation sequencing (NGS), using a single-end sequencing strategy with 400 reads. The sequencing data underwent quality control, filtering, and primer removal. CD-hit clustering was used to obtain the major sequences of the samples. Since the products were amplicones, sequences with 100% similarity were clustered together, the major sequence depth was calculated, and BLAST alignment was performed on the major sequences. The results showed that all four pairs of universal primers yielded high amplicon depths, and the obtained major sequences successfully identified the corresponding species. However, the universal primer pair Cox1bird4 showed poor amplification for the White-naped Crane, and the universal primer pair Cox1bird1 showed poor amplification for the Great Bustard. Nevertheless, the sequences from the other three pairs of universal primers were still sufficient to identify the relevant species.

[0037] The results show that the universal primers for bird identification provided by this invention can cover and accurately identify eight common bird species in the Beijing-Tianjin-Hebei region. This complementary multi-primer design avoids identification failures caused by single primer amplification failure, significantly improves the method's resistance to interference and result stability, and reduces the risk of missed detections due to species-specific gene variations. The universal primers for bird identification of this invention have practical applications and can supplement existing universal primers for birds. The primer design concept of this invention can provide a reference for the development of other universal identification primers.

[0038] Table 4. Sequencing depth of amplification products from eight bird species using universal primers. birds Cox1bird1 Cox1bird2 Cox1bird3 Cox1bird4 Blas species similarity Painted-faced duck 2105 5638 8926 2103 100% Spot-billed duck 2100 3526 5620 895 100% Mallard 3526 5631 8610 1205 100% White-naped Crane 1520 4502 5620 3 99.95 Wild Geese 850 2635 5681 562 99.95% Great Bustard 0 6543 5860 1456 99.90 Oriental White Stork 1520 2503 1502 6520 100% nightingale 2653 1560 3564 1520 99.95%

Claims

1. A method for designing universal primers for bird identification based on machine learning optimization, characterized in that, Includes the following steps: S1. Download avian COXI gene sequences from public databases and construct a local database of avian target sequences; S2. Determine conserved bases to identify conserved sequences. Conserved bases are determined based on the following principles to obtain conserved sequences: Calculate the Shannon entropy at each base position. When the Shannon entropy is ≤0.5, the conserved base is the base with the highest allele frequency. When the Shannon entropy is >0.5, the base with an allele frequency ≥0.2 is a conserved base. When two or more bases meet the conditions, the corresponding degenerate base is used. The representation of degenerate bases follows the IUPAC rules. The formula for calculating Shannon entropy is as follows: ; Where pi is the allele frequency of the i-th base, and k is the number of base types at that site; S3. Establish a multi-dimensional primer scoring system with a total score of 100 points, including coverage score, Tm score, primer interaction score, degenerate base number score, and sequence poly structure score. The weights of the above scores are 40%, 20%, 20%, 10%, and 10%, respectively. S4. Machine learning model training: Primer dataset is formed by using public databases, literature, or self-designed primers. The primer dataset is expanded by primer reverse complementation. The score of each primer is labeled according to the scoring system in S3. The model is trained using the random forest algorithm of the sklearn module. The optimal parameters are obtained through grid search and cross-validation and the model is saved. S5. Candidate primer screening and combination: Based on the conserved sequences in S2, a large number of candidate primers are generated. After evaluation by the model trained in S4, primers with high scores and amplicon lengths of 100-300bp are selected to construct primer combinations.

2. The method for designing universal primers for bird identification based on machine learning optimization according to claim 1, characterized in that, In S1, publicly available databases include NCBI and BOLD. The downloaded sequences are cleaned and processed as necessary, including: removing low-quality sequences with more than 25% of ambiguous bases; and performing BLAST alignment with the reference genome GCF_016699485.2 to remove sequences with less than 50% similarity.

3. The method for designing universal primers for bird identification based on machine learning optimization according to claim 1, characterized in that, In S3, the scores for each item in the scoring system are calculated using the following method: Coverage score refers to the proportion of primers that match the species sequences in the local database. Matching here means that the sequence is exactly matched without mismatches or insertions / deletions. Primers containing degenerate bases are broken down into different bases and matched separately to calculate the total matching proportion. The score is calculated by multiplying the matching proportion by 40. Tm score refers to scoring based on the primer Tm value. Primers containing degenerate bases are broken down into different bases and the average Tm value of the primers is calculated. 57℃<Tm<63℃ gets 20 points, 55℃≤Tm≤57℃ or 63℃≤Tm<65℃ gets 15 points, 53℃≤Tm<55℃ or 65℃≤Tm≤67℃ gets 10 points, and Tm<53℃ or Tm>67℃ gets 0 points. Primer interaction scoring: Primer interaction includes primers forming dimers or forming hairpin structures. Scoring is based on the average Tm value of the interacting sequences. For primers containing degenerate bases, the average Tm value of the interacting sequences is calculated by breaking them down into different bases. Tm < 45℃ gets 20 points, 45℃ ≤ Tm ≤ 50℃ gets 15 points, 50℃ < Tm ≤ 55℃ gets 10 points, and Tm > 55℃ gets 0 points. Degenerate base count score: The primer is scored based on the number of degenerate bases it contains. 0 bases get 10 points, 1 base gets 7 points, 2 bases get 4 points, and more than 2 bases get 0 points. Sequence poly structure scoring: Scoring is based on the number of consecutive bases contained in the primer. No points are deducted for ≤5 consecutive bases, 5 points are deducted for 6 consecutive bases, 10 points are deducted for 7 or more consecutive bases, and points are deducted cumulatively for containing 2 or more consecutive bases, until the score reaches 0.

4. A universal primer combination for bird identification designed by any one of the methods described in claims 1-3.

5. The universal primer combination for bird identification according to claim 4, characterized in that, The primer sequences include the following: Cox1bird1, sequence SEQ ID NO:1-2; Cox1bird2, sequence SEQ ID NO:3-4; Cox1bird3, sequence SEQ ID NO:5-6; Cox1bird4, sequence SEQ ID NO:7-8.

6. The universal primer combination for bird identification according to claim 5, characterized in that, The primer sequences 5'-3' are specifically as follows: SEQ ID NO:1 is TTCTCAACCAACCACAAAGA, SEQ ID NO:2 is GKGGGAATGCTATGTCKGGG; SEQ ID NO:3 is TTCTGATTCTTYGGMCACCC, SEQ ID NO:4 is ACTGTRAAYATGTGGTGGGC; SEQ ID NO:5 is TTYACCCACTGATTCCCMCT, SEQ ID NO:6 is GKCCTAGGAAGTGTTGKGGG; SEQ ID NO:7 is TRCCACGACGATACTCAGAC; SEQ ID NO:8 is GCAGCCGTGGATTCATTCRA.

7. The universal primer combination for bird identification according to claim 5, characterized in that, The primer combination was amplified according to the following PCR procedure: 95℃ for 5 min; 95℃ for 30 s, 58℃ for 40 s, 72℃ for 30 s, 32 cycles; 72℃ for 5 min; and constant temperature at 4℃.

8. The use of the primer combination of claim 4 in the preparation of a kit for identifying bird species.

9. The application according to claim 8, characterized in that, The kit includes: primer sets, DNA extraction reagents, PCR amplification reagents, and sequencing reagents.

10. The application according to claim 8, characterized in that, The kit is used for bird species identification, bird ecological monitoring, bird biodiversity assessment, bird species protection, or law enforcement detection.