Method and system for screening specific sites of homologous species and electronic equipment
By screening homologous species-specific sites or sequences through multi-genome alignment and multi-dimensional scoring models, the non-specificity and classification errors in the distinction of homologous species in existing technologies are solved, and homologous species identification with high accuracy and reliability is achieved.
Patent Information
- Application Number
- CN202510901585.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-09-30
AI Technical Summary
In existing technologies, the distinction between homologous species mainly relies on the analysis of a single site or a few gene regions, which is prone to non-specific misjudgments and classification errors. It lacks systematic multi-dimensional analysis and is difficult to comprehensively evaluate the specificity and reliability of candidate sites.
Multi-genome alignment and variation detection are used to construct a multidimensional feature matrix. Through a scoring model of specificity score, coverage score, distinguishability score and primerability score, homologous species-specific sites or sequences are screened out, and candidate nucleic acid molecules that can identify homologous species are designed.
The accuracy of evaluating specific sites or sequences between homologous species is improved, the misjudgment rate is reduced, the reliability and adaptability of the method are enhanced, and it can effectively identify specific gene sites or sequences between homologous species.
Smart Images

Figure CN120727092A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bioinformatics, and in particular to a method, system and electronic equipment for screening homologous species-specific sites. Background Art
[0002] Nucleic acid sequences are highly consistent between many pathogens or between different subtypes of pathogens. For example, Aspergillus flavus is one of the main pathogens causing pulmonary aspergillosis. Aspergillus flavus and Aspergillus oryzae are very similar in morphology, but Aspergillus flavus can produce aflatoxin, which is highly toxic and carcinogenic, while Aspergillus oryzae is domesticated and used for fermenting food and does not produce aflatoxin. For example, there is a high degree of sequence consistency between Escherichia coli and Shigella. For example, it is difficult to distinguish between different subtypes of influenza virus. In existing technologies, the distinction between highly homologous species mainly relies on the analysis of single sites or a few gene regions. However, these methods have the following problems: (1) Due to the high similarity between species, a single site or region may be non-specific, leading to misjudgment; (2) Traditional methods cannot effectively solve the identification difficulties caused by the accumulation of classification errors in genome databases; (3) There is a lack of systematic multi-dimensional analysis, making it difficult to comprehensively evaluate the specificity and reliability of candidate sites; (4) Not all genomes are fully assembled, so some sequences may be missing or have merged bases.
[0003] In view of this, the present invention is proposed. Summary of the Invention
[0004] The purpose of the present invention is to provide a method, system and electronic device for screening homologous species-specific sites to solve the problem of insufficient specificity due to high sequence similarity between homologous species.
[0005] The present invention is achieved in that: In a first aspect, the present invention provides a method for screening homologous species-specific sites and / or specific sequences, comprising the following steps: (1) Multi-genome alignment and variation detection: Input the sequenced genome sequences of each homologous species, align them to the reference genome, and generate an alignment file; use genome variation data processing tools to generate a file containing variation information; (2) Multi-dimensional data extraction and integration: Extract sequencing depth information and reference allele depth information for each sample from the file containing variant information; construct a multidimensional feature matrix containing at least two of the following features: coverage of the site in the target species, coverage of the site in non-target species, location of the variant site on the chromosome, variant identifier, variant base and type; (3) Scoring model construction: Based on the extracted multidimensional features, the specificity score, coverage score, distinguishability score and primerability score of each site were calculated using the following formula: Specificity score (Specificity_Score) = number of samples with qualified coverage of the site in non-target species / number of samples in all non-target species; Coverage score (Coverage_Score) = number of samples with qualified coverage of the site in the target species / number of samples in the total target species; Amplification_refractory_Score = (sum_mut - sum_gap) / 100bp; sum_mut represents the sum of the mutated bases within the 100-bp reference genome base range starting from the current mutation; sum_gap represents the sum of the non-mutated bases within the 100-bp reference genome base range starting from the current mutation; Primer_Score = max(ref,mut) / Tm60; max(ref,mut) is the larger of the numeric values of the variant base, Tm60 is the current mutation as the 3' end, and the position of consecutive bases > 60°C bases is calculated towards the 5' end; The above scoring indicators are used to establish the following scoring model:
[0006] β0: is the intercept term of the scoring model; set β1 to 0.4, β2 to 0.3, β3 to 0.2, and β4 to 0.1; (4) Screening and verification of candidate sites or sequences: Based on the P value of the scoring model and the preset threshold, candidate sites or sequences that are specific between homologous species are screened.
[0007] In a second aspect, the present invention provides a method for screening nucleic acid molecules that specifically identify homologous species, which comprises the following steps: based on the specific candidate sites or sequences obtained by the above-mentioned method for screening homologous species-specific sites and / or specific sequences, designing candidate nucleic acid molecules that can identify homologous species through primer design rules; and verifying the identification efficacy.
[0008] In a third aspect, the present invention provides a system for screening homologous species-specific sites and / or specific sequences, which comprises: a multi-genome alignment and variation detection unit, a multi-dimensional data extraction and integration unit, a scoring model construction unit, and a candidate site or sequence screening and verification unit; Multi-genome alignment and variation detection unit: used to input the sequenced genome sequences of each homologous species, align them to the reference genome, and generate an alignment file; through the genome variation data processing tool, a file containing variation information is generated; Multi-dimensional data extraction and integration unit: used to extract sequencing depth information and reference allele depth information of each sample from the file containing variation information output by the multi-genome alignment and variation detection unit; construct a multi-dimensional feature matrix containing at least two of the following features: coverage of the site in the target species, coverage of the site in non-target species, location of the variant site on the chromosome, variant identifier, variant base and type, variant quality value, filtering status and global annotation information of the variant; Scoring model construction unit: used to calculate the specificity score, coverage score, distinguishability score and primerability score of each site based on the multidimensional feature matrix output by the multidimensional data extraction and integration unit. The calculation formula is: Specificity score (Specificity_Score) = number of samples with qualified coverage of the site in non-target species / number of samples in all non-target species; Coverage score (Coverage_Score) = number of samples with qualified coverage of the site in the target species / number of samples in the total target species; Amplification_refractory_Score = (sum_mut - sum_gap) / 100bp; sum_mut represents the sum of the mutated bases within the 100-bp reference genome base range starting from the current mutation; sum_gap represents the sum of the non-mutated bases within the 100-bp reference genome base range starting from the current mutation; Primer_Score = max(ref,mut) / Tm60; max(ref,mut) is the larger of the numeric values of the variant base, Tm60 is the current mutation as the 3' end, and the position of consecutive bases > 60°C bases is calculated towards the 5' end; The above scoring indicators are used to establish the following scoring model:
[0009] β0: is the intercept term of the scoring model; set β1 to 0.4, β2 to 0.3, β3 to 0.2, and β4 to 0.1; The candidate site or sequence screening and verification unit is used to screen out candidate sites or regions that are specific between homologous species based on the P value of the scoring model and a preset threshold.
[0010] In a fourth aspect, the present invention provides a system for screening species-specific sites and / or specific sequences for use in screening specific gene sites and / or specific sequences between homologous species.
[0011] In a fifth aspect, the present invention provides an electronic device comprising: a processor and a memory; the processor and the memory are connected, wherein the memory is used to store a computer program, and the processor is used to call the computer program to execute the above-mentioned method for screening homologous species-specific sites and / or specific sequences, or the above-mentioned method for screening nucleic acid molecules that specifically identify homologous species.
[0012] In a sixth aspect, the present invention provides a computer-readable storage medium comprising instructions which, when executed on a computer, enable the computer to execute the above-mentioned method for screening homologous species-specific sites and / or specific sequences, or the above-mentioned method for screening nucleic acid molecules that specifically identify homologous species.
[0013] The present invention has the following beneficial effects: The present invention provides a method and system for accurately identifying or screening for species-specific loci or sequences between homologous species. This method and system address the lack of specificity in existing technologies. By integrating multiple dimensions of genomic features, including coverage, the accuracy of locus-specific assessment is significantly improved. The resulting species-specific loci or sequences can reduce the false positive rate and mitigate or avoid interference from database classification errors.
[0014] By analyzing population data composed of multiple samples, the method provided by this invention effectively reduces the interference of individual differences on the results and enhances the reliability of the method. The preset threshold of the scoring model provided by this invention is dynamically adjustable, and the screening threshold can be dynamically adjusted according to the similarity between different species, thereby improving the adaptability and universality of the method.
[0015] The method and system provided by the present invention can be widely used in fields such as biodiversity research, such as homologous species identification, vaccine strain identification, virus subtype identification, and influenza virus strain identification, and have good market prospects and application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0017] Figure 1It is a simplified flow chart of screening for homologous species-specific sites and / or specific sequences; Figure 2 The results of species identification of Aspergillus flavus, Aspergillus oryzae, Aspergillus fumigatus, Aspergillus nidulans and Aspergillus niger using primers designed based on sites 1, 2 and 3, respectively; Figure 3 The results of species identification of H3N2, H1N1, and H5N1 and H7N9 mixed quality control products using primers designed based on sites 1, 2, and 3, respectively. DETAILED DESCRIPTION
[0018] Reference will now be made in detail to embodiments of the present invention, one or more examples of which are described below. Each example is provided to illustrate, not to limit, the present invention. Indeed, it will be apparent to those skilled in the art that various modifications and variations may be made to the present invention without departing from the scope or spirit of the invention. For example, features illustrated or described as part of one embodiment may be used in another embodiment to produce further embodiments.
[0019] Glossary: Homologous species are groups of organisms that share a common ancestor in evolutionary history, whose similarities stem from genetic inheritance. Homologous species include, but are not limited to, orthologous, paralogous, heterologous, morphologically homologous, different species within the same genus, and different strains of the same species.
[0020] "Sequencing depth" refers to the average number of times a site is read by a sequencing instrument (for example, 30X depth means that each base is read 30 times on average).
[0021] Reference allele depth refers to the number of sequencing reads that support the reference genome sequence at a specific locus in the genome. It is a key metric for assessing sequencing data quality, verifying mutation reliability, and filtering false positives.
[0022] "Coverage" refers to the proportion of sample individuals (or sequence reads) in which a specific gene site (such as a SNP, an insertion / deletion site, or a specific sequence) is successfully detected (i.e., valid data is obtained) to all sequenced / typed sample individuals (or total reads) after a large-scale sequencing or genotyping experiment on the target species.
[0023] "Locus coverage in the target species" measures the proportion of valid data obtained for a specific genomic position within the target species' study sample set. It reflects the completeness of the data at that locus and its prevalence across the sample set, and is a key metric for assessing data quality, analytical reliability, and technical platform performance. High coverage is crucial for ensuring accurate and reliable results in subsequent genetic analysis.
[0024] The discriminability score can also be called the allele-specific PCR amplification score.
[0025] β0: is the intercept term of the model.
[0026] The "preset threshold of the scoring model" is a fundamental metric for screening homologous species-specific sites and / or sequences. In practical applications, the preset threshold can be dynamically (i.e., adaptively) adjusted as needed. For example, the threshold can be adjusted based on species similarities, such as animals, bacteria, and viruses, which have different interspecies similarities.
[0027] In a first aspect, the present invention provides a method for screening homologous species-specific sites and / or specific sequences, comprising the following steps: (1) Multi-genome alignment and variation detection: Input the sequenced genome sequences of each homologous species, align them to the reference genome, and generate an alignment file; use genome variation data processing tools to generate a file containing variation information; (2) Multi-dimensional data extraction and integration: Extract sequencing depth information and reference allele depth information for each sample from the file containing variant information; construct a multidimensional feature matrix containing at least two of the following features: coverage of the site in the target species, coverage of the site in non-target species, location of the variant site on the chromosome, variant identifier, variant base and type; (3) Scoring model construction: Based on the extracted multidimensional features, the specificity score, coverage score, distinguishability score and primerability score of each site were calculated using the following formula: Specificity score (Specificity_Score) = number of samples with qualified coverage of the site in non-target species / number of samples in all non-target species; the value range of specificity score is 90%~100%; Coverage score (Coverage_Score) = number of samples with qualified coverage of the site in the target species / number of samples in the total target species; the coverage score value range is 90%~100%; Amplification_refractory_Score = (sum_mut - sum_gap) / 100bp; sum_mut represents the sum of the mutated bases within the 100bp reference genome base range starting from the current mutation; sum_gap represents the sum of the non-mutated bases within the 100bp reference genome base range starting from the current mutation; the amplification_refractory_score ranges from 1% to 100%. Primer_Score = max(ref,mut) / Tm60; max(ref,mut) is the larger of the numeric values of the variant base, Tm60 is the current mutation as the 3' end, and the position of consecutive bases > 60°C bases is calculated towards the 5' end; the value range of the primer score is 1% to 100%; The above scoring indicators are used to establish the following scoring model:
[0028] β0: is the intercept term of the scoring model; the value range of β0 is [-1,1], for example, it can be 0, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, -1. The specific range can be adjusted as needed; Set β1 to 0.4, β2 to 0.3, β3 to 0.2, and β4 to 0.1; (4) Screening and verification of candidate sites or sequences: Based on the P value of the scoring model and the preset threshold, candidate sites or sequences that are specific between homologous species are screened.
[0029] The above-mentioned construction of a multidimensional feature matrix containing multiple features is to take the coverage of each site in the target species, the coverage in the non-target species and other features as a matrix.
[0030] This method integrates multiple genomic features, including coverage, and uses a scoring model to assign specificity, coverage, discriminability, and primerability scores to each locus. Based on a logistic regression formula, this method can screen for highly specific candidate loci or sequences, significantly improving the accuracy of locus-specificity assessments. The resulting selection of interspecies-specific gene loci or sequences can reduce the rate of false positives and mitigate or avoid interference from database classification errors.
[0031] By analyzing population data composed of multiple samples, the method provided by this invention effectively reduces the interference of individual differences on the results and enhances the reliability of the method. The preset threshold of the scoring model provided by this invention is dynamically adjustable, and the screening threshold can be dynamically adjusted according to the similarity between different species, thereby improving the adaptability and universality of the method.
[0032] The method and system provided by the present invention can be widely used in fields such as biodiversity research, such as homologous species identification, vaccine strain identification, virus subtype identification, and influenza virus strain identification, and have good market prospects and application value.
[0033] In step (2), multidimensional data extraction and integration relies on the following tools to extract the sequencing depth information and reference allele depth information of each sample from the files containing variant information: minimap2, samtools, bcftools, pandas, dask, and numpy.
[0034] In a preferred embodiment of the present invention, the screening and verification of candidate sites or sequences further includes: cross-validation of the candidate sites to ensure their specificity and stability on independent sample sets.
[0035] In a preferred embodiment of the present invention, the candidate sites or sequences that meet the preset threshold of the scoring model and have the maximum P value of the scoring model are selected as specific candidate sites or sequences.
[0036] The preset threshold is, for example, to select the top 3 sites with the highest P values, the top 5 sites with the highest P values, the top 8 sites with the highest P values, or the top 10 sites with the highest P values as specific candidate sites or sequences.
[0037] In a second aspect, the present invention provides a method for screening nucleic acid molecules that specifically identify homologous species, which comprises the following steps: based on the specific candidate sites or sequences obtained by the above-mentioned method for screening homologous species-specific sites and / or specific sequences, designing candidate nucleic acid molecules that can identify homologous species through primer design rules; and verifying the identification efficacy.
[0038] Primers can be designed through online primer design websites or primer design software, by inputting candidate sites or sequences and setting primer requirements (such as GC content, Tm value, etc.) to obtain candidate nucleic acid molecules that can identify homologous species; further, candidate nucleic acid molecules that can identify homologous species are synthesized, and the identification efficiency of the nucleic acid molecules is verified through experiments.
[0039] In a third aspect, the present invention provides a system for screening homologous species-specific sites and / or specific sequences, which comprises: a multi-genome alignment and variation detection unit, a multi-dimensional data extraction and integration unit, a scoring model construction unit, and a candidate site or sequence screening and verification unit; Multi-genome alignment and variation detection unit: used to input the sequenced genome sequences of each homologous species, align them to the reference genome, and generate an alignment file; through the genome variation data processing tool, a file containing variation information is generated; Multi-dimensional data extraction and integration unit: used to extract sequencing depth information and reference allele depth information for each sample from the file containing variation information output by the multi-genome alignment and variation detection unit; construct a multi-dimensional feature matrix containing at least two of the following features: coverage of the site in the target species, coverage of the site in non-target species, location of the variant site on the chromosome, variant identifier, variant base and type; Scoring model construction unit: used to calculate the specificity score, coverage score, distinguishability score and primerability score of each site based on the multidimensional feature matrix output by the multidimensional data extraction and integration unit. The calculation formula is: Specificity score (Specificity_Score) = number of samples with qualified coverage of the site in non-target species / number of samples in all non-target species; Coverage score (Coverage_Score) = number of samples with qualified coverage of the site in the target species / number of samples in the total target species; Amplification_refractory_Score = (sum_mut - sum_gap) / 100bp; sum_mut represents the sum of the mutated bases within the 100-bp reference genome base range starting from the current mutation; sum_gap represents the sum of the non-mutated bases within the 100-bp reference genome base range starting from the current mutation; Primer_Score = max(ref,mut) / Tm60; max(ref,mut) is the larger of the numeric values of the variant base, Tm60 is the current mutation as the 3' end, and the position of consecutive bases > 60°C bases is calculated towards the 5' end; The above scoring indicators are used to establish the following scoring model:
[0040] β0: is the intercept term of the scoring model; set β1 to 0.4, β2 to 0.3, β3 to 0.2, and β4 to 0.1; The candidate site or sequence screening and verification unit is used to screen out candidate sites or regions that are specific between homologous species based on the P value of the scoring model and a preset threshold.
[0041] In a preferred embodiment of the present invention, the screening and verification of candidate sites or sequences further comprises: cross-validation of the candidate sites.
[0042] In a preferred embodiment of the present invention, the candidate sites or sequences that meet the preset threshold of the scoring model and have the maximum P value of the scoring model are selected as specific candidate sites or sequences.
[0043] In a fourth aspect, the present invention provides a system for screening species-specific sites and / or specific sequences for use in screening specific gene sites and / or specific sequences between homologous species.
[0044] In a fifth aspect, the present invention provides an electronic device comprising: a processor and a memory; the processor and the memory are connected, wherein the memory is used to store a computer program, and the processor is used to call the computer program to execute the above-mentioned method for screening homologous species-specific sites and / or specific sequences, or the above-mentioned method for screening nucleic acid molecules that specifically identify homologous species.
[0045] Specifically, the electronic device may include a memory, a processor, a bus, and a communication interface, wherein the memory, processor, and communication interface are electrically connected to each other directly or indirectly to enable data transmission or interaction. For example, these components may be electrically connected to each other via one or more buses or signal lines. The processor may process information and / or data related to target identification to perform one or more functions described in this application.
[0046] The memory may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.
[0047] A processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, including a central processing unit (CPU) or a network processor (NP). It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0048] In a sixth aspect, the present invention provides a computer-readable storage medium comprising instructions which, when executed on a computer, enable the computer to execute the above-mentioned method for screening homologous species-specific sites and / or specific sequences, or the above-mentioned method for screening nucleic acid molecules that specifically identify homologous species.
[0049] To make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are described clearly and completely below. Where specific conditions are not specified in the embodiments, conventional conditions or conditions recommended by the manufacturer are used. Where the manufacturer of the reagents or instruments is not specified, all are conventional products that can be purchased commercially.
[0050] The features and performance of the present invention are further described in detail below with reference to the embodiments.
[0051] Example 1 This embodiment provides a method for screening species-specific sites and / or specific sequences of Aspergillus flavus and Aspergillus oryzae.
[0052] Data preparation: 306 Aspergillus flavus genome sequences and 124 Aspergillus oryzae genome sequences were collected from https: / / www.ncbi.nlm.nih.gov / datasets / genome / ; The reference genome GCA_009017415.1 of Aspergillus flavus was selected as the reference genome.
[0053] Flow chart of the method for screening Aspergillus flavus and Aspergillus oryzae species-specific sites and / or specific sequences Figure 1 As shown, the specific steps include: 1. Multi-genome alignment and variation detection Perform whole-genome analysis of all sequenced genomes of target and non-target species: Use the minimap2 tool to align the sequencing data to the reference genome and generate an alignment file in SAM / BAM format; SNP detection was performed using the bcftools tool to generate a VCF file containing variation information.
[0054] 2. Multi-dimensional data extraction and integration (1) Extract the sequencing depth information and reference allele depth information of each sample from the files containing variant information using the following tools: minimap2, samtools, bcftools, pandas, dask, and numpy.
[0055] A custom script extracted the sequencing depth (DP) and reference allele depth (RefDP) information of each sample from the VCF file for subsequent calculation of the sensitivity and specificity of the generated loci.
[0056] (2) Constructing a multidimensional feature matrix: Constructing a multidimensional feature matrix. This example constructs a multidimensional feature matrix that includes the following features: the coverage of each locus in the Aspergillus flavus sample (i.e., the reference allele frequency of the locus in the Aspergillus flavus sample) and the coverage of the locus in the Aspergillus oryzae sample (i.e., the reference allele frequency of the locus in the Aspergillus oryzae sample). The coverage information of each locus in the Aspergillus flavus sample and the coverage information of the locus in the Aspergillus oryzae sample are integrated into a matrix.
[0057] 3. Construction of specificity scoring model Based on the extracted multidimensional features, the specificity, coverage, distinguishability score, and primerability score of each site were calculated using the following formula: Specificity score (Specificity_Score) = number of samples with qualified coverage of the site in non-target species / number of samples in all non-target species; Coverage score (Coverage_Score) = number of samples with qualified coverage of the site in the target species / number of samples in the total target species; Amplification_refractory_Score = (sum_mut - sum_gap) / 100bp; sum_mut represents the sum of the mutated bases within the 100-bp reference genome base range starting from the current mutation; sum_gap represents the sum of the non-mutated bases within the 100-bp reference genome base range starting from the current mutation; Mutations were grouped and sorted according to chromosomes, and for each mutation, a comprehensive score was assigned based on the number of mutated bases and the mutation position within a 60-120 bp window on the positive and negative strands; Primer_Score = max(ref,mut) / Tm60; max(ref,mut) takes the larger value of the variant base, for example, SNP: G>C takes a value of 1, indel: A>CTC takes a value of 3, and ATC>G also takes a value of 3; Tm60 is the current mutation as the 3' end, and the position of consecutive bases > 60° towards the 5' end is calculated; Primers were scored based on the position of each mutation site in the primer's full length, 10 bp from the 3' end, and 5 bp from the 3' end, as well as the number and position of mutations within three different windows.
[0058] The above scoring indicators are used to establish the following scoring model:
[0059] β0: is the intercept term of the scoring model, which is 0 in this embodiment; set β1 to 0.4, β2 to 0.3, β3 to 0.2, and β4 to 0.1; The above indicators are used to establish a logistic regression formula.
[0060] 4. Screening of candidate sites or regions: Sites with scores higher than a preset threshold are selected as candidate specific sites.
[0061] The preset threshold of the scoring model in this embodiment is to select the sites with the top 3 P values.
[0062] Write the above steps into a python script and run the command as follows: python homologous-oligo.py --ref ref.fna --target targets.csv --nontarget nontargets.csv --target_key A_flavus --nontarget_key A_oryzae --target_count 275 --nontarget_count 112 --debug The highest scoring sites are: Site 1: NC_092404.1 at position 1193927 AACCTGCCATGCAATCCGTTGCGCCA>-; Site 2: NC_092405.1 at position 1883741 AGGCTGA>-; Site 3: NC_092404.1 at position 1845881 GGA T>-.
[0063] 5. Verification of candidate sites or sequences: The selected specific sites were applied to species identification of actual samples. The following primer sequences were designed for the above specific sites 1-3: Site 1: NC_092404.1 at position 1193927 AACCTGCCATGCAATCCGTTGCGCCA>-; Forward primer sequence: AACCTGCCATGCAATCCGT (SEQ ID NO: 1); Reverse primer sequence: GACTTCCGTGCGATGATCCA (SEQ ID NO: 2); Probe sequence: CGCGAGGAGTGAGGTATGGCGCA (SEQ ID NO: 3); Site 2: NC_092405.1 at position 1883741 AGGCTGA>-; Forward primer sequence: AAGATTCGCTATACAGAAGAGGCTGA (SEQ ID NO: 4); Reverse primer sequence: CGGAGGAGCTTGTCCGCTATA (SEQ ID NO: 5); Probe sequence: CACCAGGGCTCAGGGCTGAAATTACC (SEQ ID NO: 6).
[0064] Site 3: NC_092404.1 at position 1845881 GGA T>-; Forward primer sequence: TTACCCAGCTTGGGGTGGAT (SEQ ID NO: 7); Reverse primer sequence: TAGCAAAAGAAACTGCGATTGCA (SEQ ID NO: 8); Probe sequence: ATCGTCAACCCGGCGTCGATCA (SEQ ID NO: 9).
[0065] The amplification system is as follows:
[0066] The amplification procedure is as follows: Step 1: Reverse transcription reaction at 42°C for 5 min; 95°C for 10 sec; 1 cycle; Step 2: PCR reaction: 95℃ for 5 seconds, 60℃ for 20 seconds, 40 cycles.
[0067] The amplified products were subjected to fluorescence quantitative PCR, and the amplification results were referred to Figure 2 As shown in the results, the screened specific sites 1, 2, and 3 can highly specifically distinguish A. flavus from other homologous species. Primers from the literature (DOI: 10.3852 / 08-022) were unable to effectively distinguish A. flavus from other homologous species. Therefore, the specific sites 1, 2, and 3 have been validated as specific sites for A. flavus and A. oryzae.
[0068] Example 2 This example provides a method for screening for specific sites and / or specific sequences of influenza A virus H3N2 strains and influenza A virus non-H3N2 strains (including but not limited to H1N1, H5N1, and H7N9).
[0069] 1. Data Preparation 265 Influenza A virus H3N2 subtype segement 4 sequences and 1562 non-H3N2 subtype segement 4 sequences of Influenza A virus were collected from https: / / www.ncbi.nlm.nih.gov / labs / virus / vssi / # / ; The reference genome NC_007366.1 of the Influenza A virus H3N2 subtype reference genome segment 4 was selected as the reference genome.
[0070] 2. Alignment and variation detection Minimap2 was used to align the sequencing data of each sample to the reference genome; SNP detection was performed using bcftools to generate VCF files.
[0071] 3. Feature extraction and analysis Extract the DP and RefDP information of each site from the VCF file; The positive coverage of each site in H3N2 subtype sequences and in 1562 non-H3N2 sequences of Influenza A virus was calculated.
[0072] 4. Construction of Specificity Scoring Model Based on the extracted multidimensional features, the specificity, coverage, distinguishability score, and primerability score of each site were calculated using the following formula: Specificity score (Specificity_Score) = number of samples with qualified coverage of the site in non-target species / number of samples in all non-target species; Coverage score (Coverage_Score) = number of samples with qualified coverage of the site in the target species / number of samples in the total target species; Amplification_refractory_Score = (sum_mut - sum_gap) / 100bp; sum_mut represents the sum of the mutated bases within the 100-bp reference genome base range starting from the current mutation; sum_gap represents the sum of the non-mutated bases within the 100-bp reference genome base range starting from the current mutation; Primer_Score = max(ref,mut) / Tm60; max(ref,mut) takes the larger value of the variant base, for example, SNP: G>C takes a value of 1, indel: A>CTC takes a value of 3, and ATC>G also takes a value of 3; Tm60 is the current mutation as the 3' end, and the position of consecutive bases > 60° towards the 5' end is calculated; The above scoring indicators are used to establish the following scoring model:
[0073] β0: is the intercept term of the scoring model, which is -1 in this embodiment; set β1 to 0.4, β2 to 0.3, β3 to 0.2, and β4 to 0.1; The above indicators are used to establish a logistic regression formula.
[0074] 5. Specificity Scoring and Site Screening The score of each site was calculated according to the specificity scoring formula; the preset threshold of the scoring model in this embodiment was: the sites with the top three P values.
[0075] Sites with scores higher than a preset threshold were selected as candidate specific sites.
[0076] The python script runs as follows: python homologous_oligo.py --ref REF.fasta --target H3N2.csv --nontargetNot_H3N2.csv --target_key Is_H3N2 --nontarget_key Not_H3N2 --target_count 260 --nontarget_count 1500 --debug The highest scoring sites are: Site 1: The forward primer consists of multiple positions at NC_007366 positions 740: C>., 746: A>., 749: G>., 757: C>., and the reverse primer consists of multiple positions at 851: C>., 855: A>., 858: C>., 863: T>., 875: A>., 876: A>.
[0077] Site 2: The forward primer is composed of multiple positions at positions 974: G>., 975: A>., 977: C>. of NC_007366, and the reverse primer is composed of multiple positions at positions 1052: G>., 1060: C>., 1067: C>.
[0078] Site 3: The forward primer consists of multiple positions at NC_007366 positions 1052: G>., 1060: C>., 1067: C>., and the reverse primer consists of multiple positions at 1124: T>., 1137: A>., 1140: C>.
[0079] 6. Verification of candidate sites or sequences: The selected specific sites were applied to species identification of actual samples. The following primer sequences were designed for the above specific sites 1-3: Site 1: The forward primer is composed of multiple positions at positions 740: C>., 746: A>., 749: G>., 757: C>. of NC_007366, and the reverse primer is composed of multiple positions at positions 851: C>., 855: A>., 858: C>., 863: T>., 875: A>., 876: A>.
[0080] Forward primer sequence: ACCCAGGATAAGGGATGTCCC; Reverse primer sequence: TTGAGCTTTTCCCACTTCGTATTTTG; Probe sequence: AAGTAACCCCGAGGAGCAATTAGATTCCC.
[0081] Site 2: The forward primer is composed of multiple positions at positions 974: G>., 975: A>., 977: C>. of NC_007366, and the reverse primer is composed of multiple positions at positions 1052: G>., 1060: C>., 1067: C>.
[0082] Forward primer sequence: AAACCATTTCAAAATGTAAACAGGATC; Reverse primer sequence: GCCAAATATGCCTCTAGTTTGTTTC; Probe sequence: CTCTGAAATTGGCAACAGGGATGCGA.
[0083] Site 3: The forward primer consists of multiple positions at NC_007366 positions 1052: G>., 1060: C>., 1067: C>., and the reverse primer consists of multiple positions at 1124: T>., 1137: A>., 1140: C>.
[0084] Forward primer sequence: AATGTACCAGAGAAACAAACTAGAGGC; Reverse primer sequence: TGATGCCTGAAACCGTACCA; Probe sequence: CTCCCAACCATTTTCTATGAAACCCGC.
[0085] The amplification system is as follows:
[0086] The amplification procedure is as follows: Step 1: Reverse transcription reaction at 42°C for 5 min; 95°C for 10 sec; 1 cycle; Step 2: PCR reaction: 95℃ for 5 seconds, 60℃ for 20 seconds, 40 cycles.
[0087] The selected specific sites were applied to the species identification of actual clinical isolates and quality control products, and amplified according to the above system and amplification procedure. The amplification results were referred to Figure 3 The results show that the sites 1, 2 and 3 selected in this example can all achieve specific detection of influenza A virus H3N2 strain.
[0088] In summary, the present invention significantly improves the accuracy of distinguishing highly homologous species by integrating multidimensional genomic features and population-level analysis, and has a significant advantage in the accuracy of distinguishing highly homologous species compared to existing single-site analysis methods. The specificity scoring model and dynamic threshold adjustment method proposed in the present invention solve key problems existing in the prior art. The method of the present invention can be widely used in fields such as biodiversity research such as highly homologous species identification, vaccine strain identification, and viral subtype identification, and has good market prospects and application value.
[0089] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A method for screening homologous species-specific sites and / or specific sequences, characterized in that: It includes the following steps: (1) Multi-genome alignment and variation detection: Input the sequenced genome sequences of each homologous species, align them to the reference genome, and generate an alignment file; use genome variation data processing tools to generate a file containing variation information; (2) Multi-dimensional data extraction and integration: Extracting sequencing depth information and reference allele depth information for each sample from the file including the variation information; constructing a multidimensional feature matrix comprising at least two of the following features: coverage of the site in the target species, coverage of the site in the non-target species, location of the variation site on the chromosome, variation identifier, and variation base and type; (3) Scoring model construction: Based on the extracted multidimensional features, the specificity score, coverage score, distinguishability score and primerability score of each site were calculated using the following formula: Specificity score (Specificity_Score) = number of samples with qualified coverage of the site in non-target species / number of samples in all non-target species; Coverage score (Coverage_Score) = number of samples with qualified coverage of the site in the target species / number of samples in the total target species; Amplification_refractory_Score = (sum_mut - sum_gap) / 100bp; sum_mut represents the sum of the mutated bases within the 100-bp reference genome base range starting from the current mutation; sum_gap represents the sum of the non-mutated bases within the 100-bp reference genome base range starting from the current mutation; Primer_Score = max(ref,mut) / Tm60; max(ref,mut) is the larger of the numeric values of the variant base, Tm60 is the current mutation as the 3' end, and the position of consecutive bases > 60°C bases is calculated towards the 5' end; The above scoring indicators are used to establish the following scoring model: β0: is the intercept term of the scoring model; set β1 to 0.4, β2 to 0.3, β3 to 0.2, and β4 to 0.1; (4) Screening of candidate sites or sequences: Based on the P value of the scoring model and the preset threshold, candidate sites or sequences that are specific between homologous species are screened.
2. The method for screening homologous species-specific sites and / or specific sequences according to claim 1, characterized in that: The screening of the candidate sites or sequences also includes verification of the candidate sites or sequences, which includes: cross-validation of the candidate sites.
3. The method for screening homologous species-specific sites and / or specific sequences according to claim 1, characterized in that: The candidate sites or sequences that meet the preset threshold of the scoring model and have the maximum P value of the scoring model are selected as specific candidate sites or sequences.
4. A method for screening nucleic acid molecules for specific identification of homologous species, characterized in that: The method comprises the following steps: based on the specific candidate sites or sequences obtained by the method for screening homologous species-specific sites and / or specific sequences according to any one of claims 1 to 3, designing candidate nucleic acid molecules that can identify homologous species according to primer design rules; and verifying the identification efficacy.
5. A system for screening homologous species-specific sites and / or specific sequences, characterized in that: It includes: Multi-genome alignment and variation detection unit, multi-dimensional data extraction and integration unit, scoring model construction unit, and candidate site or sequence screening unit; The multi-genome alignment and variation detection unit is used to input the sequenced genome sequences of each homologous species, align them to the reference genome, and generate an alignment file; and generate a file including variation information using a genome variation data processing tool; The multidimensional data extraction and integration unit is used to extract the sequencing depth information and reference allele depth information of each sample from the file containing the variation information output by the multi-genome alignment and variation detection unit; construct a multidimensional feature matrix containing at least two of the following features: the coverage of the site in the target species, the coverage of the site in the non-target species, the position of the variation site on the chromosome, the variation identifier, the variation base and type, the variation quality value, the filtering status and the global annotation information of the variation; The scoring model construction unit is used to calculate the specificity score, coverage score, distinguishability score and primerability score of each site based on the multidimensional feature matrix output by the multidimensional data extraction and integration unit. The calculation formula is: Specificity score (Specificity_Score) = number of samples with qualified coverage of the site in non-target species / number of samples in all non-target species; Coverage score (Coverage_Score) = number of samples with qualified coverage of the site in the target species / number of samples in the total target species; Amplification_refractory_Score = (sum_mut - sum_gap) / 100bp; sum_mut represents the sum of the mutated bases within the 100-bp reference genome base range starting from the current mutation; sum_gap represents the sum of the non-mutated bases within the 100-bp reference genome base range starting from the current mutation; Primer_Score = max(ref,mut) / Tm60; max(ref,mut) is the larger of the numeric values of the variant base, Tm60 is the current mutation as the 3' end, and the position of consecutive bases > 60°C bases is calculated towards the 5' end; The above scoring indicators are used to establish the following scoring model: β0: is the intercept term of the scoring model; set β1 to 0.4, β2 to 0.3, β3 to 0.2, and β4 to 0.1; The candidate site or sequence screening unit is used to screen out candidate sites or regions that are specific between homologous species based on the P value of the scoring model and a preset threshold.
6. The system for screening homologous species-specific sites and / or specific sequences according to claim 5, characterized in that: The system further comprises a candidate site or sequence verification unit, which is used for cross-verifying the candidate sites.
7. The system for screening homologous species-specific sites and / or specific sequences according to claim 5, characterized in that: The candidate sites or sequences that meet the preset threshold of the scoring model and have the maximum P value of the scoring model are selected as specific candidate sites or sequences.
8. Use of the system for screening species-specific sites and / or specific sequences according to any one of claims 5 to 7 in screening specific gene sites and / or specific sequences between homologous species.
9. An electronic device, characterized in that: It includes: processor and memory; The processor is connected to the memory, wherein: The memory is used to store a computer program, and the processor is used to call the computer program to execute the method for screening homologous species-specific sites and / or specific sequences as described in any one of claims 1 to 3, or the method for screening nucleic acid molecules for specifically identifying homologous species as described in claim 4.
10. A computer-readable storage medium, characterized in that The invention comprises instructions which, when run on a computer, enable the computer to execute the method for screening homologous species-specific sites and / or specific sequences according to any one of claims 1 to 3, or the method for screening nucleic acid molecules for specifically identifying homologous species according to claim 4.