Species identification reagent, identification method and application

By combining identification and extraction reagents and magnetic bead purification with identification and analysis using deep learning models, the problem of insufficient identification capabilities of existing DNA barcoding technology for special samples has been solved, achieving efficient and accurate species identification.

CN120683307APending Publication Date: 2025-09-23NANCHANG SAGE BIOTECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511129050.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-09-23

Smart Images

  • Figure CN120683307A_ABST
    Figure CN120683307A_ABST
Patent Text Reader

Abstract

The invention discloses a species identification reagent, an identification method and application, belongs to the technical field of species identification, and solves the problems that according to an existing method, DNA bar codes are generated based on repetitive sequences, amplification of the repetitive sequences is easily interfered by genome repetitive areas, and the identification capacity is limited. A sequencing amplification product is purified through an agarose gel electrophoresis recovery method, the sequencing amplification product is purified through a magnetic bead purification reagent, the sequencing amplification product is subjected to bidirectional sequencing, and a sequencing data set is identified and analyzed based on an identification analysis model; according to the species identification method provided by the invention, the to-be-identified sample tissue can be subjected to purification and impurity removal through synergistic cooperation of the identification extraction reagent and the magnetic bead purification reagent, the sequencing quality of a special sample such as a degradation sample or an inhibitor-containing sample is remarkably improved, and the species can be more accurately identified by utilizing the identification analysis model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of species identification, and in particular relates to a species identification reagent, an identification method and an application thereof. Background Art

[0002] Species identification is an important technical link in the fields of forensic medicine, ecological protection, food safety, and biodiversity research. Traditional species identification methods mainly rely on morphological feature analysis. Their limitations are that they are highly dependent on expert experience, are easily affected by sample integrity, and are difficult to effectively identify partially degraded, processed, or mixed samples. With the development of molecular biology technology, identification methods based on DNA sequence analysis have gradually become mainstream. Among them, DNA barcoding technology has been recommended by the International Barcode of Life (ICB) as a standardized technology for species identification due to its advantages such as standardization, high sensitivity, rapidity, and strong traceability.

[0003] However, the application of existing DNA barcoding technology still faces many challenges:

[0004] Traditional methods are sensitive to sample types (such as different tissues such as plant roots, stems, and leaves) and degree of degradation. Some test kits are only suitable for specific sample types, resulting in a high failure rate in identification of samples from complex sources (such as processed foods and corrupted tissues).

[0005] Most existing technical solutions require multiple independent steps (such as DNA extraction, amplification, purification, sequencing, etc.), and the links are not tightly connected, which can easily introduce cross-contamination or operational errors, and has high technical requirements for experimental personnel.

[0006] Chinese patent CN109979536B discloses a method for species identification based on DNA barcodes, specifically relating to the field of species identification methods. The method includes the following three identification methods: Identification method 1: Organizing a database based on the species' living environment and family; Identification method 2: Establishing a tree-like database based on the number of ATCG bases in the DNA sequence; Identification method 3: Generating DNA barcodes based on repeat sequence data in the DNA sequence to establish a database. However, existing methods generate DNA barcodes based on repeat sequences. The amplification of repeat sequences (such as ISSRs) is susceptible to interference from genomic repeat regions, lacks specificity, and requires manual primer design. The experimental process is cumbersome, and the sample requirements are high, making it impossible to optimize sample adaptability. The identification capability for some special samples, such as degraded samples or samples containing inhibitors, is limited. To address these issues, we have proposed species identification reagents, identification methods, and applications. Summary of the Invention

[0007] The purpose of the present invention is to address the shortcomings of the existing technology and provide species identification reagents, identification methods and applications, which solve the problems of the existing methods based on repetitive sequences to generate DNA barcodes, the amplification of repetitive sequences is easily interfered by genomic repetitive regions, lacks specificity, requires manual primer design, has a cumbersome experimental process, has high requirements for samples and cannot optimize the adaptability of samples, and has limited identification capabilities for some special samples such as degraded samples or samples containing inhibitors.

[0008] The present invention is achieved by providing a species identification method, which comprises:

[0009] Obtaining a sample tissue to be identified, pre-treating the sample tissue to be identified with an identification extraction reagent to remove impurities in the sample tissue to be identified;

[0010] The pre-treated sample tissue to be identified is subjected to PCR amplification, the amplified DNA product is extracted, and the sequencing amplification product is purified by agarose gel electrophoresis recovery method to obtain the purified amplified DNA product;

[0011] The purified amplified DNA product is taken and subjected to sequencing reaction amplification to obtain a sequencing amplification product, which is purified by a magnetic bead purification reagent and bidirectionally sequenced to obtain a sequencing data set representing the tissue information of the sample to be identified;

[0012] Capturing publicly available genomic information, pre-building a data comparison library based on the genomic information, and pre-training an identification and analysis model using the genomic information in the data comparison library to obtain an identification and analysis model based on maximum likelihood method and support vector machine;

[0013] Load the sequencing dataset, identify and analyze the sequencing dataset based on the pre-trained identification analysis model, calculate the sequence similarity between the sequencing dataset and the data comparison library, and dynamically determine the similarity threshold based on the sequencing dataset to determine whether the sequence similarity exceeds the similarity threshold. If so, output the species identification result corresponding to the sample tissue to be identified.

[0014] Preferably, the method of pre-treating the sample tissue to be identified using the identification extraction reagent specifically includes:

[0015] Take the sample tissue to be identified, determine the tissue type based on the tissue acquisition source, and pre-prepare the identification extraction reagent. Wherein, the tissue type is plant tissue or animal tissue, and the identification extraction reagent includes 8mL of modified extraction solution, 180μL of mixed enzyme solution, and 400μL of proteinase K;

[0016] If the tissue type is animal tissue, take 0.01-0.02g of animal tissue, cut it into pieces and put it into a 1.5mL centrifuge tube, add 100μL of modified extract solution to the centrifuge tube, mix it for 1 minute, add 5μL of proteinase K to the centrifuge tube, place the centrifuge tube containing the animal tissue sample in a centrifuge and centrifuge for 10-15 seconds, then shake it for 10 seconds, take the shaken centrifuge tube and place it in a metal bath for heating treatment, the heating conditions are 56℃ for 30 minutes and 95℃ for 10 minutes, after heating, place the centrifuge tube containing the animal tissue sample in a centrifuge and centrifuge for 3 minutes, and take the animal tissue supernatant;

[0017] If the tissue type is plant tissue, the collected plant tissue is placed in liquid nitrogen for rapid cold preservation. During the test, the roots and stems of the plant tissue are cut into pieces, and then the plant tissue is ground with liquid nitrogen. During grinding, the plant tissue is placed in a mesh bag for filtration to release the plant tissue cells. 0.03 g of the ground plant tissue is weighed and placed in a centrifuge tube. Pure water is added to immerse the plant tissue. The mixed enzyme is added to the centrifuge tube at a volume ratio of 2:1 between the mixed enzyme and the plant tissue. After mixing evenly, the centrifuge tube is placed in a metal bath and enzymolyzed at 50-55°C for 6-8 hours. The centrifuge tube containing the enzymatic hydrolysis solution after enzymolysis is centrifuged for 1 minute, the residue is discarded, and the supernatant is recovered and placed in a high-speed centrifuge for 5 minutes. The precipitate is recovered and placed in a precipitate recovery tube. 100 μL of modified extract is added to the precipitate recovery tube. After mixing for 1 minute, 5 μL of proteinase K is added to the precipitate recovery tube. The precipitate recovery tube is centrifuged and heated, and the plant tissue supernatant is taken;

[0018] The extracted animal tissue supernatant / plant tissue supernatant is placed in a purification tank, and the animal tissue supernatant / plant tissue supernatant is diluted with a dilution buffer to obtain a tissue dilution solution, which is extracted with chloroform and precipitated to obtain sample tissue DNA.

[0019] Preferably, the dilution buffer comprises the following raw materials in percentage by volume:

[0020] Sodium lauryl sulfate: 20%;

[0021] Proteinase K: 5%;

[0022] Isopropyl alcohol: 40%;

[0023] Tris-EDTA: 30%;

[0024] Polyvinylpyrrolidone: 5%;

[0025] The modified extract comprises the following raw materials in percentage by volume:

[0026] Chelex-100: 80%;

[0027] β-glucanase: 5%;

[0028] Citric acid: 5%;

[0029] Disodium hydrogen phosphate: 10%.

[0030] Preferably, the method for preparing the modified extract comprises:

[0031] Weigh Chelex-100 resin and place it in a mixed solution of citric acid and disodium hydrogen phosphate. Add five times the volume of Tris-HCl solution to the mixed solution of Chelex-100 resin, citric acid, and disodium hydrogen phosphate, and vortex overnight until completely dissolved to obtain a Chelex solution.

[0032] β-glucanase at a concentration of 30 U / μL was dissolved in Tris-HCl buffer, and then the β-glucanase was added dropwise to the Chelex solution. After gently inverting to mix, the solution was allowed to stand at room temperature for 30 minutes to prepare a modified extract.

[0033] Preferably, the magnetic bead purification reagent comprises the following raw materials in parts by weight: 50-60 parts of porous nano-magnetic beads, 100-120 parts of magnetic bead washing solution, 100-120 parts of sodium chloride solution, 5-12 parts of epoxysilane, 10-15 parts of carbodiimide, and 5-15 parts of triethanolamine;

[0034] The method for purifying the sequencing amplification product by using a magnetic bead purification reagent includes:

[0035] The triethanolamine was dissolved in deionized water, and CTAB, sodium salicylate and triethanolamine solution were mixed in a mass ratio of 3:1. After mixing evenly, the mixture was heated to 75-80°C and magnetically stirred for 30-45 minutes. Then, 0.5 times the mass of sodium salicylate was added to tetraethyl orthosilicate, and the mixture was stirred for 2 hours. The product was centrifuged and collected, and washed with anhydrous ethanol for 3-5 times. The prepared product was dispersed in 3 times the volume of concentrated hydrochloric acid, condensed and refluxed at 55-65°C for 5 hours, and the refluxed product was dispersed in 2 times the volume of anhydrous ethanol. Ammonia water and 3-aminopropyltriethoxysilane with a mass of 0.5 times the mass of tetraethyl orthosilicate were added to the mixed solution of the refluxed product and anhydrous ethanol, and stirred at room temperature for 5-6 hours at a stirring speed of 2000-2200 rpm to obtain porous nanomagnetic beads.

[0036] Place porous nanomagnetic beads in a reactor and soak them in 30% sodium hydroxide solution for 15-20 minutes. Then pour the porous nanomagnetic beads into a round-bottom flask and wash them with 70% ethanol to remove non-polar residues on the surface of the porous nanomagnetic beads. Mix the porous nanomagnetic beads with isopropanol (3 times the volume of the porous nanomagnetic beads) and then ultrasonically clean them for 2-5 times before use.

[0037] The porous nanomagnetic beads and epoxysilane were ultrasonically dispersed in anhydrous methanol 4 times the volume of the porous nanomagnetic beads in a mass ratio of 1:1, and then the temperature was raised to 65°C and stirred at a stirring speed of 500-520 rpm for 30-35 minutes. Carbodiimide and 3-aminopropyltriethoxysilane were mixed in MES buffer at a mass ratio of 2:1, and the mixed solution of carbodiimide and 3-aminopropyltriethoxysilane was mixed with the porous nanomagnetic beads at room temperature for 15 minutes to obtain a modified magnetic bead mixture. Triethanolamine was dissolved in MES buffer, and then 0.02 times the volume of triethanolamine was added to streptavidin to obtain a triethanolamine and streptavidin mixture. The triethanolamine and streptavidin mixture and the modified magnetic bead mixture were oscillated and reacted for 3 hours. The mixture was washed with MES buffer 3-5 times, and then the modified magnetic beads were mixed with sodium chloride solution to obtain a magnetic bead-sodium chloride solution.

[0038] Take 0.5 μL of magnetic beads-sodium chloride solution and sequencing amplification product, add the magnetic beads-sodium chloride solution and sequencing amplification product into anhydrous ethanol twice the volume of the sequencing amplification product, and use an oscillator to mix the sample to obtain a purified mixed system. Place the purified mixed system on a magnetic stand for 1 minute, add magnetic bead washing solution to wash impurities, wash 3-5 times, add 50-100 μL of elution buffer ethanol to elute the modified magnetic beads, incubate at room temperature for 5-10 minutes, and obtain the purified sequencing amplification product.

[0039] Preferably, the identification and analysis model uses a support vector machine as the initial model, introduces a convolutional neural network into the initial model, embeds a multiple sequence alignment layer MSA and a Bayesian hybrid model into the convolutional neural network, introduces a maximum likelihood method into the multiple sequence alignment layer MSA, and the multiple sequence alignment layer MSA is used to align the DNA sequences of multiple species in the sequencing data set, and combines the maximum likelihood method to consider the base substitution frequency and inequality to evaluate the support rate of the data alignment library sequence through Bootstrap resampling, introduces a hidden Markov model after the convolutional neural network, and matches the conserved domains through HMMER to identify subspecies or closely related species-specific variations, and combines the k-mer similarity to determine the similarity between the sequencing data set and the data alignment library sequence.

[0040] Preferably, the method for pre-training an identification and analysis model using genomic information in a data comparison library comprises:

[0041] Pre-build a data comparison library, capture the published genome information of various species from public databases, organize and annotate the collected genome information, establish a local data comparison library, pre-process the genome information, annotate the genome sequences in the genome information, and obtain the unique labels corresponding to the genome sequences;

[0042] Traverse the data comparison library, grab the training set, validation set, and test set from the data comparison library, and set the hyperparameters, loss function, and training rounds of the identification analysis model;

[0043] Load the training set, encode it using one-hot encoding, freeze the hidden Markov model and support vector machine layers, pre-train the CNN-MSA architecture using the training set, generate a phylogenetic tree using the multiple sequence alignment layer (MSA), calculate bootstrap support, adjust sequence weights using a Bayesian mixture model, evaluate the SVM classification accuracy using the validation set after each round of training, and dynamically adjust hyperparameters;

[0044] After training, HMMER and k-mer similarity are integrated and the similarity threshold is dynamically determined. To dynamically determine the similarity threshold, the HMMER layer is unfrozen, the test set sequence is input, conserved domains are matched, the k-mer similarity matrix is ​​calculated, and weighted fusion is performed with the BLAST results. The SVM and CNN-MSA architectures are jointly trained. The k-mer similarity constraint is added to the loss function of the CNN-MSA architecture, and the similarity threshold for a single label category is determined based on the loss function of the CNN-MSA architecture.

[0045] Load the validation set, use the validation set to evaluate the pre-trained identification and analysis model, calculate the evaluation index, and check whether the identification and analysis model has overfitting or underfitting. If the evaluation index of the validation set is greater than the preset index threshold and the identification and analysis model does not have overfitting or underfitting, output the converged identification and analysis model.

[0046] Preferably, the method for identifying and analyzing a sequencing dataset based on a pre-trained identification and analysis model comprises:

[0047] Obtain sequencing data sets, use FastQC to remove low-quality reads, convert clean reads to FASTA format, merge overlapping reads, and obtain preprocessed DNA sequences;

[0048] The DNA sequence is converted into a numerical vector through k-mer encoding, and the feature extraction of the DNA sequence converted into a numerical vector is combined with a convolutional neural network to obtain a feature extraction set;

[0049] The multiple sequence alignment layer (MSA) aligns the feature extraction set with the reference sequences in the data alignment library, and uses RAxML to calculate the phylogenetic tree and evaluate the Bootstrap support rate between the feature extraction set and closely related species.

[0050] The reference sequences of closely related species with a coverage greater than 70% were screened and considered as reliable support. The hidden Markov model was used to match conserved domains through HMMER, identify subspecies or closely related species-specific variations, and calculate the matching score.

[0051] The Jellyfish tool was used to count the k-mer frequencies of the feature extraction set, and the k-mer similarity matrix was constructed with the data alignment library. The cosine similarity was calculated and used as the sequence similarity. The similarity threshold corresponding to the sequencing data set was determined based on the HMMER layer.

[0052] Determine whether the sequence similarity exceeds the similarity threshold. If so, output the reference sequence with the highest similarity, and use the reference sequence with the highest similarity as the species identification result corresponding to the sample tissue to be identified.

[0053] On the other hand, the present invention also provides a species identification reagent, which includes an identification extraction reagent, a magnetic bead purification reagent, an agarose gel electrophoresis reagent, and an identification gel recovery reagent.

[0054] Another aspect of the present invention is to provide an application of the species identification method of the present invention in species identification.

[0055] Compared with the prior art, the embodiments of the present application have the following beneficial effects:

[0056] In an embodiment of the present invention, a species identification method is provided. Through the coordinated cooperation of identification extraction reagents and magnetic bead purification reagents, the sample tissue to be identified can be purified and impurities removed, and the sequencing quality of special samples such as degraded samples or samples containing inhibitors is significantly improved. It is particularly suitable for samples with high plant polysaccharide content. In addition, the identification analysis model adopts advanced DNA barcoding technology combined with a deep learning model, which can more accurately identify species, especially for morphologically similar and difficult to distinguish species. At the same time, the maximum likelihood method is introduced into the multiple sequence alignment layer (MSA), considering the frequency and inequality of base substitutions, and evaluating the support rate of the data comparison library sequence through Bootstrap resampling, thereby improving the accuracy of identification, realizing the evolutionary support rate of quantitative sequence variation, accurately distinguishing homologous variation from convergent evolution, simplifying the identification steps, and avoiding complex database classification storage and multiple search processes. Users only need to process and test samples according to the provided operating procedures to quickly obtain accurate identification results.

[0057] In the embodiments of the present invention, a dilution buffer is used to dilute the extracted supernatant, and then combined with a chloroform extraction method, impurities such as proteins, polysaccharides, and pigments can be effectively removed, thereby improving the purity of the DNA. The components of the dilution buffer (such as sodium lauryl sulfate, proteinase K, isopropanol, Tris-EDTA, polyvinyl pyrrolidone, etc.) cooperate with each other to help further purify the DNA. In addition, the physical and chemical treatment conditions during the pretreatment process are relatively mild, avoiding damage to the DNA by severe mechanical shear forces and strong acids and alkalis, and can effectively protect the integrity of the DNA, improve the integrity and quality of the DNA, and facilitate subsequent applications such as long-fragment DNA analysis and high-precision sequencing.

[0058] In the embodiments of the present invention, the porous nanoparticles in the magnetic bead purification reagent can significantly increase the surface area of ​​the beads, improving the adsorption capacity for target molecules such as DNA and primer dimers. Epoxysilane hydrolyzes to generate silanol groups (-Si-OH), which form hydrogen bonds with the phosphate backbone of DNA, enhancing adsorption specificity. EDC can activate carboxylic acid groups, promoting covalent binding between the magnetic bead surface and biotin-labeled molecules (such as streptavidin), improving stability, and thus reducing nonspecific adsorption on the magnetic bead surface.

[0059] In the embodiments of the present invention, the magnetic bead purification reagent and method are suitable for various types of DNA samples, including high-inhibitor, low-concentration or degraded samples, including DNA extracted from different tissues, cells or biological samples, and have wide applicability. Due to the high quality of the purified DNA, the success rate of downstream identification can be improved, and the identification duplication and reagent consumption caused by poor DNA quality can be reduced, thereby reducing the overall identification cost. The DNA purity and recovery rate are significantly better than traditional methods.

[0060] In an embodiment of the present invention, an identification and analysis model is provided. The identification and analysis model uses a support vector machine as an initial model, introduces a convolutional neural network, a multiple sequence alignment layer MSA, and a Bayesian hybrid model into the initial model, and combines the maximum likelihood method to consider the frequency and inequality of base substitutions to evaluate the support rate of the data comparison library sequence through Bootstrap resampling. Through adaptive threshold adjustment, it supports the rapid expansion of new species, and the MSA generates a phylogenetic tree and combines the maximum likelihood method to deeply analyze the base substitution rules, provide more valuable information for the model, improve the accuracy of identification, and can simultaneously process multiple bioinformatics data, such as DNA sequences, conserved domains, k-mer similarity, etc. HMM matches conserved domains through the HMMER tool, identifies subspecies or closely related species-specific variations, determines sequence similarity in combination with k-mer similarity, and efficiently processes complex bioinformatics data. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1It is a schematic diagram of the implementation process of the species identification method provided by the present invention.

[0062] Figure 2 The figure shows the electrophoresis diagram of nucleic acid after PCR amplification of the pre-treated sample tissue to be identified.

[0063] Figure 3 The figure shows the partial sequencing results of the sequencing amplification product bovine COI in the embodiment of the present invention by a PCR sequencer.

[0064] Figure 4 The results of partial sequencing of the bovine 12S rRNA sequenced amplification product according to the embodiment of the present invention by a PCR sequencer are shown.

[0065] Figure 5 These are the test results of the particle size of the porous nanomagnetic beads prepared in Examples 1-5 of the present invention.

[0066] Figure 6 Schematic diagram showing the results of DNA adsorption test on the purified sequencing amplification products of Examples 1-5 and Comparative Examples 1-4 of the present invention.

[0067] Figure 7 Schematic diagram showing the results of DNA purity testing of the purified sequencing amplification products of Examples 1-5 and Comparative Examples 1-4 of the present invention.

[0068] Figure 8 The figure shows the test results of comparing the identification and analysis model of the present invention with the conventional model. DETAILED DESCRIPTION

[0069] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of the application are only for the purpose of describing specific embodiments and are not intended to limit this application. The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of this application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.

[0070] Example 1

[0071] The embodiment of the present invention provides a species identification method, Figure 1 A schematic diagram of a species identification method implementation process is shown, wherein the species identification method specifically includes:

[0072] Step S10, obtaining a sample tissue to be identified, and pre-treating the sample tissue to be identified with an identification extraction reagent to remove impurities in the sample tissue to be identified;

[0073] It should be noted that the species identification method uses species identification reagents to achieve purification and sample impurity removal optimization. The species identification reagents include identification extraction reagents, magnetic bead purification reagents, agarose gel electrophoresis reagents, and identification gel recovery reagents. The identification extraction reagents and magnetic bead purification reagents are stored in a -20°C environment for a long time. If they need to be used within one week, they should be stored in a 4°C environment.

[0074] The method for pre-treating the sample tissue to be identified using the identification extraction reagent specifically includes:

[0075] Step S101: Take a sample tissue to be identified, determine the tissue type based on the tissue acquisition source, and pre-prepare an identification extraction reagent. If the tissue type is plant tissue or animal tissue, the identification extraction reagent includes 8 mL of modified extract solution, 180 μL of mixed enzyme solution, and 400 μL of proteinase K.

[0076] It should be noted that the concentration of the proteinase K mother solution is 20 mg / mL, the concentration of the mixed enzyme is 1 U / μL, the mixed enzyme can be a mixture of cellulase, pectinase and hemicellulase, and the volume ratio of cellulase, pectinase and hemicellulase can be 1:1:3.

[0077] Step S102: If the tissue type is animal tissue, 0.01-0.02 g of animal tissue is minced and placed in a 1.5 mL centrifuge tube. 100 μL of modified extract is added to the centrifuge tube and mixed for 1 minute. 5 μL of proteinase K is added to the centrifuge tube. The centrifuge tube containing the animal tissue sample is placed in a centrifuge and centrifuged for 10-15 seconds, then shaken for 10 seconds. The shaken centrifuge tube is placed in a metal bath for heating at 56°C for 30 minutes and 95°C for 10 minutes. After heating, the centrifuge tube containing the animal tissue sample is placed in a centrifuge and centrifuged for 3 minutes. The animal tissue supernatant is collected.

[0078] In an embodiment of the present invention, when pretreating animal tissues, Chelex-100 resin combined with β-glucanase can efficiently remove lipid inhibitors and decompose polysaccharide residues (such as fat in animal livers), and centrifugation and heating treatment (56°C cell lysis + 95°C enzyme inactivation) can significantly improve the DNA release rate, which is particularly suitable for high-fat, high-protein animal tissues (such as muscle and liver).

[0079] Step S103, if the tissue type is plant tissue, take the collected plant tissue and put it into liquid nitrogen for rapid cold storage. During the test, the roots and stems of the plant tissue are cut into pieces, and then the plant tissue is ground with liquid nitrogen. During the grinding, the plant tissue is placed in a mesh bag for filtration to release the plant tissue cells. 0.03 g of the ground plant tissue is weighed and placed in a centrifuge tube. Pure water is added to immerse the plant tissue. The mixed enzyme is added to the centrifuge tube at a volume ratio of 2:1 between the mixed enzyme and the plant tissue. After mixing evenly, the centrifuge tube is placed in a metal bath and enzymolyzed at 50-55°C for 6-8 hours. The centrifuge tube containing the enzymatic hydrolysis solution after enzymolysis is centrifuged for 1 minute, the residue is discarded, and the supernatant is recovered and placed in a high-speed centrifuge for 5 minutes. The precipitate is recovered and placed in a precipitate recovery tube. 100 μL of modified extract is added to the precipitate recovery tube. After mixing for 1 minute, 5 μL of proteinase K is added to the precipitate recovery tube. The precipitate recovery tube is centrifuged and heated to take the plant tissue supernatant;

[0080] In the embodiments of the present invention, during pretreatment of plant tissue, rapid cooling and grinding with liquid nitrogen are performed to avoid DNA degradation. A mixed enzyme (cellulase / pectinase) is combined to completely decompose the cell wall, thereby resolving the interference of plant polysaccharides and lignin in DNA extraction. Two centrifugal purifications (one step to remove residues and two steps to precipitate DNA) reduce residual impurities and improve the purity of DNA extracted from plant tissue.

[0081] Step S104, placing the extracted animal tissue supernatant / plant tissue supernatant in a purification tank, diluting the animal tissue supernatant / plant tissue supernatant with a dilution buffer to obtain a tissue dilution solution, extracting with chloroform and precipitating the tissue dilution solution to obtain sample tissue DNA.

[0082] In the embodiments of the present invention, a dilution buffer is used to dilute the extracted supernatant, and then combined with a chloroform extraction method, impurities such as proteins, polysaccharides, and pigments can be effectively removed, thereby improving the purity of the DNA. The components of the dilution buffer (such as sodium lauryl sulfate, proteinase K, isopropanol, Tris-EDTA, polyvinyl pyrrolidone, etc.) cooperate with each other to help further purify the DNA. In addition, the physical and chemical treatment conditions during the pretreatment process are relatively mild, avoiding damage to the DNA by severe mechanical shear forces and strong acids and alkalis, and can effectively protect the integrity of the DNA, improve the integrity and quality of the DNA, and facilitate subsequent applications such as long-fragment DNA analysis and high-precision sequencing.

[0083] In an embodiment of the present invention, the dilution buffer comprises the following raw materials in percentage by volume:

[0084] Sodium lauryl sulfate: 20%;

[0085] Proteinase K: 5%;

[0086] Isopropyl alcohol: 40%;

[0087] Tris-EDTA: 30%;

[0088] Polyvinylpyrrolidone: 5%;

[0089] Among them, sodium dodecyl sulfate can strongly lyse cell membranes, polyvinyl pyrrolidone can chelate polyphenols (such as tannic acid in plant roots and stems) to prevent oxidative damage to DNA, isopropanol can efficiently precipitate DNA, and combined with Tris-EDTA to maintain pH stability and inhibit nuclease activity.

[0090] The modified extract comprises the following raw materials in percentage by volume:

[0091] Chelex-100: 80%;

[0092] β-glucanase: 5%;

[0093] Citric acid: 5%;

[0094] Disodium hydrogen phosphate: 10%.

[0095] The present invention also provides a method for preparing a modified extract, comprising:

[0096] Step S201, weighing Chelex-100 resin, adding it to a mixed solution of citric acid and disodium hydrogen phosphate, adding a Tris-HCl solution five times the volume of the mixed solution to the mixed solution of Chelex-100 resin, citric acid, and disodium hydrogen phosphate, and vortexing overnight until completely dissolved, to obtain a Chelex solution;

[0097] In step S202 , 30 U / μL β-glucanase is dissolved in Tris-HCl buffer, and then the β-glucanase is added dropwise to the Chelex solution. After gently inverting to mix, the solution is allowed to stand at room temperature for 30 minutes to obtain a modified extract.

[0098] In an embodiment of the present invention, a modified extract and a preparation method are provided. The modified extract is composed of Chelex-100, β-glucanase, citric acid, and disodium hydrogen phosphate. Citric acid and disodium hydrogen phosphate can adjust the pH value of the solution and provide a suitable environment for DNA extraction. Citric acid can form a stable complex with metal ions, further reducing the damage of metal ions to DNA; disodium hydrogen phosphate helps to maintain the alkaline environment of the solution. The components of the modified extract can effectively inhibit the activity of DNA enzymes and RNA enzymes, preventing nucleic acids from being degraded during the extraction process. Chelex-100 can chelate metal ions that activate nucleases and reduce the activity of nucleases; components such as β-glucanase and citric acid also help to create an environment that is not conducive to nuclease activity, thereby protecting the integrity of DNA.

[0099] Soybean tissue, corn tissue, and bovine tissue were pretreated using the identification extraction reagent containing the modified extract prepared in Example 1 of the present invention. A conventional CTAB method was used as a control group for soybean tissue and corn tissue, respectively, while a commercially available kit A was used as a control group for bovine tissue. The DNA extraction products from soybean tissue, corn tissue, and bovine tissue were tested for concentration and purity using a NanoDrop 8000 spectrophotometer to demonstrate the effect of the modified extract. At least five replicates were performed for each sample tissue to ensure data accuracy and reproducibility. The average DNA concentration, A260 / A280 ratio, protein content, and polysaccharide content of each sample tissue were calculated. The test results are shown in Table 4.

[0100] Table 4

[0101]

[0102] As shown in Table 4, the A260 / A280 ratio for soybean tissue was 1.82, while that for the CTAB method was only 1.46, significantly lower, indicating that the CTAB method is subject to significant protein contamination in soybean tissue extraction. The A260 / A280 ratio for corn tissue was 1.87, while that for the CTAB method was only 1.67, still below the ideal range, indicating that protein contamination is also a concern in corn tissue extraction. The A260 / A280 ratio for beef tissue was 1.91, while that for commercial kit A was 1.62, significantly lower, indicating significant phenolic or protein contamination. This suggests that the high purity of the modified extract stems from its specific removal of impure proteins and nucleic acid-binding substances, avoiding the incomplete extraction of traditional methods or interference from residual protease inhibitors in the kit. Furthermore, the residual protein content indicates that the modified extract, by adding a proteinase K inhibitor and optimizing the pH of the lysis system, inhibits nuclease activity and reduces protein adsorption, thus avoiding the co-precipitation of protein and DNA in traditional methods. At the same time, the modified extract can disrupt the binding of polysaccharides to nucleic acids through chelating agents or special surfactants, which is superior to the traditional CTAB method and commercially available kit A. This shows that the modified extract can effectively extract high-purity and high-concentration DNA from soybean, corn, and cattle tissues, and is suitable for various molecular biology experiments such as PCR and sequencing.

[0103] Step S20, performing PCR amplification on the pre-treated sample tissue to be identified, extracting the amplified DNA product, and purifying the sequencing amplified product by agarose gel electrophoresis recovery method to obtain a purified amplified DNA product;

[0104] It should be noted that when performing PCR amplification on pretreated tissue samples to be identified, various reagents required for PCR are mixed in a sterile PCR tube or 96-well plate in a certain proportion. These reagents include a DNA template (DNA extracted from the pretreated tissue sample), primers (short single-stranded DNA fragments that specifically recognize the target DNA sequence), dNTPs (four deoxyribonucleotides that serve as raw materials for DNA synthesis), a buffer (to maintain the pH and ionic strength of the reaction system, providing a suitable reaction environment for Taq DNA polymerase), and Taq DNA polymerase (a high-temperature resistant DNA polymerase that catalyzes DNA chain extension). The prepared PCR reaction system is then placed in a PCR instrument. The temperature cycle program of the PCR instrument is set according to the characteristics of the target DNA sequence. The amplification process of the PCR reaction system by the PCR instrument is conventional technology and will not be described in detail here. Among them, during amplification, the primers include animal primer 1 (COI), animal primer 2 (12s rRNA), and animal primer 3 (16s rRNA). The target fragment size of animal primer 1 is 600-800, the target fragment size of animal primer 2 is 413-461, and the target fragment size of animal primer 3 is 589-632. The target fragment size of plant primer 1 (rbcL), plant primer 2 (psbA-trnH), plant primer 3 (trnL), and plant primer 1 is 700-800, the target fragment size of plant primer 2 is 291-559, and the target fragment size of plant primer 3 is 268-337. Considering that the amplifier model of each laboratory is different, the optimal number of cycles needs to be explored. 35 cycles are recommended for animal samples and 38 cycles are recommended for plant samples. The amplified DNA product can be subjected to agarose gel electrophoresis. The results of electrophoresis can be used to determine whether the sample can amplify the target fragment. If a single band can be seen ( Figure 2 ) then proceed to the next step of the experiment; if no single band is seen, adjust the amplification conditions or re-extract the sample for the experiment. Figure 2 The figure shows the electrophoresis diagram of nucleic acid after PCR amplification of the sample tissue to be identified after pretreatment. Figure 2 Specifically, this is the result of a multi-species nucleic acid fragment analysis of mitochondrial genes (COI, 12S, and 16S rRNA). Sample types: PCR amplification products or purified nucleic acids of mitochondrial genes (COI, 12S, and 16S rRNA) from cattle, sheep, pigs, and chickens. A single band is visible in the image, indicating that the next step of the experiment can be carried out.

[0105] In an embodiment of the present invention, the magnetic bead purification reagent comprises the following raw materials in parts by weight: 50 parts of porous nano-magnetic beads, 100 parts of magnetic bead washing solution, 100 parts of sodium chloride solution, 5 parts of epoxysilane, 10 parts of carbodiimide, and 5 parts of triethanolamine;

[0106] In the embodiment of the present invention, when the sequencing amplification product is purified by the agarose gel electrophoresis recovery method, the sequencing amplification product is purified by a combination of an agarose gel electrophoresis reagent and an identification gel recovery reagent;

[0107] Among them, the agarose gel electrophoresis reagent consists of 10g agarose powder, 30mL 50× buffer, 80μL color development solution, 240μL loading buffer, and 100μL DNA maker, while the identification gel recovery reagent includes 120 sets of adsorption columns and collection tubes, 72mL buffer B2, 24mL washing solution 1 (96mL anhydrous ethanol needs to be added), and 1.8mL eluent 1.

[0108] Step S30, taking the purified amplified DNA product, performing a sequencing reaction amplification on the amplified DNA product to obtain a sequencing amplification product, purifying the sequencing amplification product using a magnetic bead purification reagent, and performing bidirectional sequencing on the purified sequencing amplification product to obtain a sequencing data set representing the tissue information of the sample to be identified; Figure 3 The results of partial sequencing of the amplified product bovine COI by PCR sequencer in the embodiment of the present invention are shown. Figure 4 The following table shows the partial sequencing results of the bovine 12S rRNA sequenced as the amplified product of the present invention using a PCR sequencer. The data can be analyzed using the "Genescan" analysis software included with the sequencer, exported as a file with "fasta" or "seq", and then opened in Notepad to obtain a sequencing data set.

[0109] In the embodiments of the present invention, the porous nanoparticles in the magnetic bead purification reagent can significantly increase the surface area of ​​the beads, improving the adsorption capacity for target molecules such as DNA and primer dimers. Epoxysilane hydrolyzes to generate silanol groups (-Si-OH), which form hydrogen bonds with the phosphate backbone of DNA, enhancing adsorption specificity. EDC can activate carboxylic acid groups, promoting covalent binding between the magnetic bead surface and biotin-labeled molecules (such as streptavidin), improving stability, and thus reducing nonspecific adsorption on the magnetic bead surface.

[0110] The method for purifying the sequencing amplification product by using a magnetic bead purification reagent includes:

[0111] Step S301, dissolving triethanolamine in deionized water, taking CTAB, sodium salicylate and triethanolamine solution in a mass ratio of 3:1, mixing, heating to 75 ° C., and using magnetic stirring for 30 minutes, then adding 0.5 times the mass of sodium salicylate tetraethyl orthosilicate, stirring for 2 hours, centrifuging the product, washing with anhydrous ethanol 3 times, taking the prepared product and dispersing it in 3 times the volume of concentrated hydrochloric acid, condensing and refluxing at 55 ° C. for 5 hours, and dispersing the refluxed product in In 2 volumes of anhydrous ethanol, 0.5 times the mass of ammonia water and 3-aminopropyltriethoxysilane in tetraethyl orthosilicate were added to a mixed solution of the reflux product and anhydrous ethanol, and stirred at room temperature for 5 hours at a stirring speed of 2000 rpm to obtain porous nanomagnetic beads. CTAB and tetraethyl orthosilicate reacted synergistically to produce porous silicon-based magnetic beads with a pore size distribution of 5-20 nm. The high-temperature condensation reflux method can optimize the surface morphology of the magnetic beads, thereby enhancing mechanical stability.

[0112] Step S302: porous nanomagnetic beads are placed in a reactor and soaked in 30% sodium hydroxide solution for 15 minutes to remove unreacted silane groups on the surface of the beads and reduce nonspecific adsorption. The porous nanomagnetic beads are then poured into a round-bottom flask and washed with 70% ethanol to remove non-polar residues on the surface of the porous nanomagnetic beads. The porous nanomagnetic beads are mixed with isopropanol at a volume 3 times that of the porous nanomagnetic beads, and then ultrasonically cleaned twice to thoroughly remove residues between the beads and improve purification consistency. The solution is then set aside.

[0113] Step S303, ultrasonically dispersing the porous nanomagnetic beads and epoxysilane in anhydrous methanol 4 times the volume of the porous nanomagnetic beads at a mass ratio of 1:1, then heating to 65°C and stirring at a stirring speed of 500 rpm for 30 minutes, mixing carbodiimide and 3-aminopropyltriethoxysilane in MES buffer at a mass ratio of 2:1, and then mixing the carbodiimide and 3-aminopropyltriethoxysilane mixed solution with the porous nanomagnetic beads at room temperature for 15 minutes to obtain a modified magnetic bead mixture, taking triethanolamine and dissolving it in MES buffer, and then adding 0.02 times the volume of triethanolamine streptavidin to obtain a triethanolamine and streptavidin mixture, shaking the triethanolamine and streptavidin mixture with the modified magnetic bead mixture for 3 hours, washing with MES buffer 3 times, and then mixing the modified magnetic beads with sodium chloride solution to obtain a magnetic bead-sodium chloride solution;

[0114] Step S304: Take 0.5 μL of magnetic bead-sodium chloride solution and sequencing amplification product, add the magnetic bead-sodium chloride solution and sequencing amplification product into anhydrous ethanol twice the volume of the sequencing amplification product, and use an oscillator to mix the sample to obtain a purified mixed system. Place the purified mixed system on a magnetic stand for 1 minute, add magnetic bead washing solution to wash impurities, wash 5 times, add 50 μL of elution buffer ethanol to elute the modified magnetic beads, and incubate at room temperature for 10 minutes to obtain the purified sequencing amplification product.

[0115] In the embodiments of the present invention, the magnetic bead purification reagent and method are suitable for various types of DNA samples, including high-inhibitor, low-concentration or degraded samples, including DNA extracted from different tissues, cells or biological samples, and have wide applicability. Due to the high quality of the purified DNA, the success rate of downstream identification can be improved, and the identification duplication and reagent consumption caused by poor DNA quality can be reduced, thereby reducing the overall identification cost. The DNA purity and recovery rate are significantly better than traditional methods.

[0116] Step S40, capturing publicly available genomic information, pre-building a data comparison library based on the genomic information, and pre-training an identification and analysis model using the genomic information in the data comparison library to obtain an identification and analysis model based on maximum likelihood method and support vector machine;

[0117] It should be noted that, in an embodiment of the present invention, the identification and analysis model uses a support vector machine as the initial model, introduces a convolutional neural network into the initial model, embeds a multiple sequence alignment layer MSA and a Bayesian hybrid model into the convolutional neural network, introduces the maximum likelihood method into the multiple sequence alignment layer MSA, and uses the multiple sequence alignment layer MSA to compare the DNA sequences of multiple species in the sequencing data set. The maximum likelihood method is combined with the base substitution frequency and inequality to evaluate the support rate of the data comparison library sequence through Bootstrap resampling. A hidden Markov model is introduced after the convolutional neural network. The hidden Markov model matches conserved domains through HMMER to identify subspecies or closely related species-specific variations, and combines the k-mer similarity to determine the similarity between the sequencing data set and the data comparison library sequence.

[0118] The method of pre-training an identification and analysis model using genomic information in a data comparison library includes:

[0119] Step S401, pre-constructing a data comparison library, capturing published genome information of various species from public databases, organizing and annotating the collected genome information, establishing a local data comparison library, pre-processing the genome information, annotating the genome sequences in the genome information, and obtaining unique labels corresponding to the genome sequences, wherein the pre-processing of the genome information includes sequence cleaning (removing low-quality sequences (such as coverage <90%, N ratio >5%), short fragments), de-redundancy (using CD-HIT (similarity threshold 95%) to remove duplicate sequences), and format conversion (converting FASTA format to Phylip format for multiple sequence alignment);

[0120] Step S402: traverse the data comparison library, grab the training set, validation set, and test set from the data comparison library, set the hyperparameters, loss function, and training rounds of the identification analysis model, where the learning rate hyperparameter can be 0.02, the loss function can be the cross entropy loss function, and the training rounds are set to 100-120 times;

[0121] Step S403: Load the training set, encode the training set using one-hot encoding, freeze the hidden Markov model and support vector machine layers, pre-train the CNN-MSA architecture using the training set, generate a phylogenetic tree using the multiple sequence alignment layer (MSA), calculate the bootstrap support rate, adjust the sequence weights using the Bayesian mixture model, evaluate the SVM classification accuracy using the validation set after each round of training, dynamically adjust hyperparameters, freeze the HMM and support vector machine layers during the pre-training phase, and focus on training the CNN-MSA architecture so that the model initially learns the characteristics and patterns of the sequencing data, reduces training interference, and improves pre-training efficiency and quality;

[0122] Step S404: After training is completed, the HMMER and k-mer similarities are integrated and the similarity threshold is dynamically determined. When the similarity threshold is dynamically determined, the HMMER layer is unfrozen, the test set sequence is input, the conserved domain is matched, the k-mer similarity matrix is ​​calculated (using the Jellyfish tool), and the BLAST result is weighted and fused. The SVM and CNN-MSA architectures are jointly trained. The loss function of the CNN-MSA architecture adds the k-mer similarity constraint. The similarity threshold of a single label category is determined based on the loss function of the CNN-MSA architecture. The similarity threshold is dynamically determined during the model training and prediction process to avoid the risk of misjudgment or missed judgment caused by a fixed threshold. The conserved domain is matched by HMM and the k-mer similarity matrix is ​​calculated. The weighted fusion of the BLAST results is combined, the SVM and CNN-MSA architectures are jointly trained, and finally the similarity threshold of a single label category is determined based on the loss function of the CNN-MSA architecture, thereby improving the adaptability and accuracy of the model.

[0123] Step S405, load the verification set, use the verification set to evaluate the pre-trained identification and analysis model, calculate the evaluation indicators (such as accuracy, recall rate, F1 value, etc.), check whether the identification and analysis model has overfitting or underfitting, if the evaluation indicator of the verification set is greater than the preset indicator threshold and the identification and analysis model does not have overfitting or underfitting, output the converged identification and analysis model.

[0124] In an embodiment of the present invention, an identification and analysis model is provided. The identification and analysis model uses a support vector machine as an initial model, introduces a convolutional neural network, a multiple sequence alignment layer MSA, and a Bayesian hybrid model into the initial model, and combines the maximum likelihood method to consider the frequency and inequality of base substitutions to evaluate the support rate of the data comparison library sequence through Bootstrap resampling. Through adaptive threshold adjustment, it supports the rapid expansion of new species, and the MSA generates a phylogenetic tree and combines the maximum likelihood method to deeply analyze the base substitution rules, provide more valuable information for the model, improve the accuracy of identification, and can simultaneously process multiple bioinformatics data, such as DNA sequences, conserved domains, k-mer similarity, etc. HMM matches conserved domains through the HMMER tool, identifies subspecies or closely related species-specific variations, determines sequence similarity in combination with k-mer similarity, and efficiently processes complex bioinformatics data.

[0125] Step S50: Load the sequencing data set, identify and analyze the sequencing data set based on the pre-trained identification and analysis model, calculate the sequence similarity between the sequencing data set and the data comparison library, and dynamically determine the similarity threshold based on the sequencing data set to determine whether the sequence similarity exceeds the similarity threshold. If it exceeds the similarity threshold, output the species identification result corresponding to the sample tissue to be identified.

[0126] In an embodiment of the present invention, a species identification method is provided. Through the coordinated cooperation of identification extraction reagents and magnetic bead purification reagents, the sample tissue to be identified can be purified and impurities removed, and the sequencing quality of special samples such as degraded samples or samples containing inhibitors is significantly improved. It is particularly suitable for samples with high plant polysaccharide content. In addition, the identification analysis model adopts advanced DNA barcoding technology combined with a deep learning model, which can more accurately identify species, especially for morphologically similar and difficult to distinguish species. At the same time, the maximum likelihood method is introduced into the multiple sequence alignment layer (MSA), considering the frequency and inequality of base substitutions, and evaluating the support rate of the data comparison library sequence through Bootstrap resampling, thereby improving the accuracy of identification, realizing the evolutionary support rate of quantitative sequence variation, accurately distinguishing homologous variation from convergent evolution, simplifying the identification steps, and avoiding complex database classification storage and multiple search processes. Users only need to process and test samples according to the provided operating procedures to quickly obtain accurate identification results.

[0127] An embodiment of the present invention further provides a method for identifying and analyzing a sequencing dataset based on a pre-trained identification and analysis model, the method comprising:

[0128] Step S501: Obtain a sequencing data set, use FastQC to remove low-quality reads, convert clean reads to FASTA format, merge overlapping reads, obtain preprocessed DNA sequences, and filter reads with low coverage (<90%) and high N ratio (>5%) using quality scores to reduce noise interference and ensure the reliability of subsequent analysis.

[0129] Merge overlapping reads: solve the problem of short fragment overlap in high-throughput sequencing, increase the effective sequence length, and enhance the efficiency of subsequent alignment;

[0130] Step S502: Convert the DNA sequence into a numerical vector using k-mer encoding, and extract features from the converted DNA sequence using a convolutional neural network to obtain a feature extraction set. Here, k can be 15. The k-mer encoding divides the DNA sequence into short segments, calculates the frequency distribution, and captures local base patterns. The k-mer encoding focuses on local signals, while the CNN extracts global context information. The combination of the two improves the model's ability to identify complex variations (such as indels).

[0131] Step S503: The multiple sequence alignment layer (MSA) aligns the feature extraction set with the reference sequences in the data alignment library, and uses RAxML to calculate the phylogenetic tree to evaluate the Bootstrap support rate between the feature extraction set and closely related species.

[0132] It should be noted that the Multiple Sequence Alignment (MSA) layer (e.g., MAFFT) aligns reference sequences from multiple species, thereby preserving evolutionary signals and addressing the site bias problem of single sequence alignments. Furthermore, by evaluating the Bootstrap support ratio of feature extraction sets and closely related species, it can enhance evolutionary constraints, retaining only highly supported reference sequences from closely related species, reducing noise interference from distantly related species and improving the specificity of downstream analysis. Furthermore, the construction of phylogenetic trees based on the maximum likelihood method can quantify the divergence time between species, with support ratios (Bootstrap ≥ 70%) identifying reliable evolutionary branches.

[0133] Step S504, screen more than 70% of the reference sequences of closely related species, and regard more than 70% of the reference sequences of closely related species as reliable support. The hidden Markov model matches the conserved domains through HMMER, identifies subspecies or closely related species-specific variations, and calculates the matching score. HMMER matches the conserved domains (such as protein functional domains) in the Pfam database, and can accurately identify subspecies-specific variations (such as viral capsid protein mutations). In combination with the hidden Markov model (HMM), the matching threshold can be dynamically adjusted to reduce false positives (such as random mutation interference).

[0134] Step S505: Use the Jellyfish tool to count the k-mer frequencies of the feature extraction set, construct a k-mer similarity matrix with the data comparison library, calculate the cosine similarity, use the cosine similarity as the sequence similarity, and determine the similarity threshold corresponding to the sequencing data set based on the HMMER layer;

[0135] It is important to note that the Jellyfish tool was used to calculate the k-mer frequencies of the feature extraction set and construct a k-mer similarity matrix with the data alignment library. This step provides a method for quantifying the similarity between sequencing data and reference sequences. By calculating cosine similarity, the similarity value between sequences can be obtained, providing a numerical basis for species identification. By combining information from the HMMER layer to determine the similarity threshold corresponding to the sequencing dataset, the species identification criteria can be dynamically adjusted. This step allows for flexible determination of the similarity threshold based on different datasets and analysis requirements, improving the adaptability and accuracy of species identification.

[0136] Step S506: determining whether the sequence similarity exceeds a similarity threshold; if so, outputting a reference sequence with the highest similarity, and using the reference sequence with the highest similarity as the species identification result corresponding to the sample tissue to be identified;

[0137] In this embodiment of the present invention, by determining whether the sequence similarity exceeds a similarity threshold, the reference sequence with the highest similarity is output as the species identification result corresponding to the sample tissue to be identified. This step can quickly and accurately determine the final conclusion of species identification, ensuring the reliability and accuracy of the identification results. The output of the reference sequence with the highest similarity and its species label supports one-click traceability (such as directly linking individual information in a database in forensic medicine).

[0138] The sequence similarity is calculated using the following formula:

[0139]

[0140]

[0141]

[0142] Among them, K-mer_sim represents sequence similarity. When calculating sequence similarity, the numerator is the dot product of the k-mer frequency of the sample and the database, and the denominator is the product of the module length of each frequency vector, which conforms to the definition of cosine similarity. It is applicable to high-dimensional sparse data (such as k-mer frequency vectors) and is sensitive to local features. t(i) and DB(i) respectively represent the frequency of occurrence of the i-th K-mer_ in the feature extraction set and the summary frequency of the i-th K-mer_ in the conserved structural domain in the data comparison library. k represents the number of continuous subsequences i in the feature extraction set, and d ci ,sc pi They represent the likelihood value of the i-th K-mer_ in the feature extraction set and the matching score between the i-th K-mer_ and the conserved domain, respectively. logL(i|DB) represents the log likelihood of the conserved domain. In formula (3), the k-mer with the largest log likelihood in the conserved domain is selected as the weight benchmark. By maximizing the log likelihood L(i|DB), the k-mer that best matches the database is selected to improve the reliability of weight distribution.

[0143] In the embodiment of the present invention, the identification analysis model is compared with the conventional model, wherein the conventional model includes SVM, LSTM, CNN and multiple sequence alignment layer MSA. The test results are shown in Table 5 and Figure 8 As shown, Figure 8 The figure shows the test results of comparing the identification and analysis model of the present invention with the conventional model.

[0144] Table 5

[0145] Model Accuracy Accuracy Recall F1 Identification analysis model 91.36 92.18 90.62 91.65 Support Vector Machine 82.67 85.33 83.68 82.51 LSTM 83.62 84.66 84.12 84.68 CNN 87.65 86.35 85.47 85.61 MSA 88.45 89.02 87.23 89.34

[0146] From Table 5 and Figure 8 It can be seen that compared with traditional machine learning methods and other deep learning methods (LSTM, CNN), the identification and analysis models in this embodiment have achieved better results, which shows that the model of the present invention has higher reliability and efficiency in species identification tasks and can provide more powerful tools for research and applications in related fields.

[0147] Example 2

[0148] The steps of the species identification method in this embodiment are similar to those in Example 1. The difference from Example 1 is that the magnetic bead purification reagent includes the following raw materials in parts by weight: 60 parts of porous nano-magnetic beads, 120 parts of magnetic bead washing solution, 120 parts of sodium chloride solution, 12 parts of epoxysilane, 15 parts of carbodiimide, and 15 parts of triethanolamine;

[0149] Among them, the method of purifying the sequencing amplification product by magnetic bead purification reagent includes:

[0150] Step S301, dissolving triethanolamine in deionized water, taking CTAB, sodium salicylate and triethanolamine solution in a mass ratio of 3:1, mixing, heating to 80°C, and using magnetic stirring for 45 minutes, then adding 0.5 times the mass of sodium salicylate tetraethyl orthosilicate, stirring for 2 hours, collecting the product by centrifugation, washing with anhydrous ethanol 5 times, taking the prepared product and dispersing it in 3 volumes of concentrated hydrochloric acid, condensing and refluxing at 65°C for 5 hours, dispersing the refluxed product in 2 volumes of anhydrous ethanol, adding 0.5 times the mass of ammonia water and 3-aminopropyltriethoxysilane to the mixed solution of the refluxed product and anhydrous ethanol, stirring at room temperature for 6 hours at a stirring speed of 2200 rpm, to obtain porous nanomagnetic beads;

[0151] Step S302: porous nano-magnetic beads are placed in a reactor and soaked in a 30% sodium hydroxide solution for 20 minutes. The porous nano-magnetic beads are then poured into a round-bottom flask and washed with 70% ethanol to remove non-polar residues on the surface of the porous nano-magnetic beads. The porous nano-magnetic beads are mixed with isopropyl alcohol (3 times the volume of the porous nano-magnetic beads) and ultrasonically cleaned 5 times for later use.

[0152] Step S303, ultrasonically dispersing the porous nanomagnetic beads and epoxysilane in anhydrous methanol 4 times the volume of the porous nanomagnetic beads at a mass ratio of 1:1, then heating to 65°C and stirring at a stirring speed of 520 rpm for 35 minutes, mixing carbodiimide and 3-aminopropyltriethoxysilane in MES buffer at a mass ratio of 2:1, and then mixing the carbodiimide and 3-aminopropyltriethoxysilane mixed solution with the porous nanomagnetic beads at room temperature for 15 minutes to obtain a modified magnetic bead mixture, taking triethanolamine and dissolving it in MES buffer, and then adding 0.02 times the volume of triethanolamine streptavidin to obtain a triethanolamine and streptavidin mixture, shaking and mixing the triethanolamine and streptavidin mixture with the modified magnetic bead mixture for 3 hours, washing with MES buffer 5 times, and then mixing the modified magnetic beads with sodium chloride solution to obtain a magnetic bead-sodium chloride solution;

[0153] Step S304: Take 0.5 μL of magnetic bead-sodium chloride solution and sequencing amplification product, add the magnetic bead-sodium chloride solution and sequencing amplification product into anhydrous ethanol twice the volume of the sequencing amplification product, and use an oscillator to mix the sample to obtain a purified mixed system. Place the purified mixed system on a magnetic rack for 1 minute, add magnetic bead washing solution to wash impurities, wash 5 times, add 100 μL of elution buffer ethanol to elute the modified magnetic beads, and incubate at room temperature for 10 minutes to obtain the purified sequencing amplification product.

[0154] Example 3

[0155] The steps of the species identification method in this embodiment are similar to those in Example 1. The difference from Example 1 is that the magnetic bead purification reagent includes the following raw materials in parts by weight: 52 parts of porous nano-magnetic beads, 102 parts of magnetic bead washing solution, 102 parts of sodium chloride solution, 6 parts of epoxysilane, 11 parts of carbodiimide, and 6 parts of triethanolamine;

[0156] Among them, the method of purifying the sequencing amplification product by magnetic bead purification reagent includes:

[0157] Step S301, dissolving triethanolamine in deionized water, taking CTAB, sodium salicylate and triethanolamine solution in a mass ratio of 3:1, mixing, heating to 77°C, and using magnetic stirring for 32 minutes, then adding 0.5 times the mass of sodium salicylate tetraethyl orthosilicate, stirring for 2 hours, collecting the product by centrifugation, washing with anhydrous ethanol 4 times, taking the prepared product and dispersing it in 3 volumes of concentrated hydrochloric acid, condensing and refluxing at 56°C for 5 hours, dispersing the refluxed product in 2 volumes of anhydrous ethanol, adding 0.5 times the mass of ammonia water and 3-aminopropyltriethoxysilane to the mixed solution of the refluxed product and anhydrous ethanol, stirring at room temperature for 5.2 hours at a stirring speed of 2010 rpm, to obtain porous nanomagnetic beads;

[0158] Step S302: porous nano-magnetic beads are placed in a reactor and soaked in a 30% sodium hydroxide solution for 16 minutes. The porous nano-magnetic beads are then poured into a round-bottom flask and washed with 70% ethanol to remove non-polar residues on the surface of the porous nano-magnetic beads. The porous nano-magnetic beads are mixed with isopropyl alcohol (3 times the volume of the porous nano-magnetic beads) and ultrasonically cleaned 3 times for later use.

[0159] Step S303, ultrasonically dispersing the porous nanomagnetic beads and epoxysilane in anhydrous methanol 4 times the volume of the porous nanomagnetic beads at a mass ratio of 1:1, then heating to 65°C and stirring at a stirring speed of 510 rpm for 31 minutes, mixing carbodiimide and 3-aminopropyltriethoxysilane in MES buffer at a mass ratio of 2:1, and then mixing the carbodiimide and 3-aminopropyltriethoxysilane mixed solution with the porous nanomagnetic beads at room temperature for 15 minutes to obtain a modified magnetic bead mixture, taking triethanolamine and dissolving it in MES buffer, and then adding 0.02 times the volume of triethanolamine streptavidin to obtain a triethanolamine and streptavidin mixture, shaking and mixing the triethanolamine and streptavidin mixture with the modified magnetic bead mixture for 3 hours, washing with MES buffer 4 times, and then mixing the modified magnetic beads with sodium chloride solution to obtain a magnetic bead-sodium chloride solution;

[0160] Step S304: Take 0.5 μL of magnetic bead-sodium chloride solution and sequencing amplification product, add the magnetic bead-sodium chloride solution and sequencing amplification product into anhydrous ethanol twice the volume of the sequencing amplification product, and use an oscillator to mix the sample to obtain a purified mixed system. Place the purified mixed system on a magnetic rack for 1 minute, add magnetic bead washing solution to wash impurities, wash 5 times, add 60 μL of elution buffer ethanol to elute the modified magnetic beads, and incubate at room temperature for 6 minutes to obtain the purified sequencing amplification product.

[0161] Example 4

[0162] The steps of the species identification method in this embodiment are similar to those in Example 1. The difference from Example 1 is that the magnetic bead purification reagent includes the following raw materials in parts by weight: 58 parts of porous nano-magnetic beads, 118 parts of magnetic bead washing solution, 117 parts of sodium chloride solution, 11 parts of epoxysilane, 14 parts of carbodiimide, and 13 parts of triethanolamine;

[0163] Among them, the method of purifying the sequencing amplification product by magnetic bead purification reagent includes:

[0164] Step S301, dissolving triethanolamine in deionized water, taking CTAB, sodium salicylate and triethanolamine solution in a mass ratio of 3:1, mixing, heating to 78°C, and using magnetic stirring for 44 minutes, then adding 0.5 times the mass of sodium salicylate tetraethyl orthosilicate, stirring for 2 hours, collecting the product by centrifugation, washing with anhydrous ethanol 4 times, taking the prepared product and dispersing it in 3 times the volume of concentrated hydrochloric acid, condensing and refluxing at 62°C for 5 hours, dispersing the refluxed product in 2 times the volume of anhydrous ethanol, adding 0.5 times the mass of ammonia water and 3-aminopropyltriethoxysilane to the mixed solution of the refluxed product and anhydrous ethanol, stirring at room temperature for 6 hours at a stirring speed of 2200 rpm, to obtain porous nanomagnetic beads;

[0165] Step S302: porous nano-magnetic beads are placed in a reactor and soaked in a 30% sodium hydroxide solution for 19 minutes. The porous nano-magnetic beads are then poured into a round-bottom flask and washed with 70% ethanol to remove non-polar residues on the surface of the porous nano-magnetic beads. The porous nano-magnetic beads are mixed with isopropyl alcohol (3 times the volume of the porous nano-magnetic beads) and ultrasonically cleaned four times for later use.

[0166] Step S303, ultrasonically dispersing the porous nanomagnetic beads and epoxysilane in anhydrous methanol 4 times the volume of the porous nanomagnetic beads at a mass ratio of 1:1, then heating to 65°C and stirring at a stirring speed of 518 rpm for 34 minutes, mixing carbodiimide and 3-aminopropyltriethoxysilane in MES buffer at a mass ratio of 2:1, and then mixing the carbodiimide and 3-aminopropyltriethoxysilane mixed solution with the porous nanomagnetic beads at room temperature for 15 minutes to obtain a modified magnetic bead mixture, taking triethanolamine and dissolving it in MES buffer, and then adding 0.02 times the volume of triethanolamine streptavidin to obtain a triethanolamine and streptavidin mixture, shaking and mixing the triethanolamine and streptavidin mixture with the modified magnetic bead mixture for 3 hours, washing with MES buffer 5 times, and then mixing the modified magnetic beads with sodium chloride solution to obtain a magnetic bead-sodium chloride solution;

[0167] Step S304: Take 0.5 μL of magnetic bead-sodium chloride solution and sequencing amplification product, add the magnetic bead-sodium chloride solution and sequencing amplification product into anhydrous ethanol twice the volume of the sequencing amplification product, and use an oscillator to mix the sample to obtain a purified mixed system. Place the purified mixed system on a magnetic rack for 1 minute, add magnetic bead washing solution to wash impurities, wash 5 times, add 90 μL of elution buffer ethanol to elute the modified magnetic beads, and incubate at room temperature for 5-10 minutes to obtain the purified sequencing amplification product.

[0168] Example 5

[0169] The steps of the species identification method in this embodiment are similar to those in Example 1. The difference from Example 1 is that the magnetic bead purification reagent includes the following raw materials in parts by weight: 55 parts of porous nano-magnetic beads, 110 parts of magnetic bead washing solution, 110 parts of sodium chloride solution, 7 parts of epoxysilane, 13 parts of carbodiimide, and 10 parts of triethanolamine;

[0170] Among them, the method of purifying the sequencing amplification product by magnetic bead purification reagent includes:

[0171] Step S301, dissolving triethanolamine in deionized water, taking CTAB, sodium salicylate and triethanolamine solution in a mass ratio of 3:1, mixing, heating to 80°C, and using magnetic stirring for 38 minutes, then adding 0.5 times the mass of sodium salicylate tetraethyl orthosilicate, stirring for 2 hours, collecting the product by centrifugation, washing with anhydrous ethanol 4 times, taking the prepared product and dispersing it in 3 times the volume of concentrated hydrochloric acid, condensing and refluxing at 60°C for 5 hours, dispersing the refluxed product in 2 times the volume of anhydrous ethanol, adding 0.5 times the mass of ammonia water and 3-aminopropyltriethoxysilane to the mixed solution of the refluxed product and anhydrous ethanol, stirring at room temperature for 5.5 hours at a stirring speed of 2100 rpm, to obtain porous nanomagnetic beads;

[0172] Step S302: porous nanomagnetic beads are placed in a reactor and soaked in a 30% sodium hydroxide solution for 18 minutes. The porous nanomagnetic beads are then poured into a round-bottom flask and washed with 70% ethanol to remove non-polar residues on the surface of the porous nanomagnetic beads. The porous nanomagnetic beads are mixed with isopropyl alcohol (3 times the volume of the porous nanomagnetic beads) and ultrasonically cleaned four times for later use.

[0173] Step S303, ultrasonically dispersing the porous nanomagnetic beads and epoxysilane in anhydrous methanol 4 times the volume of the porous nanomagnetic beads at a mass ratio of 1:1, then heating to 65°C and stirring at a stirring speed of 510 rpm for 33 minutes, mixing carbodiimide and 3-aminopropyltriethoxysilane in MES buffer at a mass ratio of 2:1, and then mixing the carbodiimide and 3-aminopropyltriethoxysilane mixed solution with the porous nanomagnetic beads at room temperature for 15 minutes to obtain a modified magnetic bead mixture, taking triethanolamine and dissolving it in MES buffer, and then adding 0.02 times the volume of triethanolamine streptavidin to obtain a triethanolamine and streptavidin mixture, shaking the triethanolamine and streptavidin mixture with the modified magnetic bead mixture for 3 hours, washing with MES buffer 3 times, and then mixing the modified magnetic beads with sodium chloride solution to obtain a magnetic bead-sodium chloride solution;

[0174] Step S304: Take 0.5 μL of magnetic bead-sodium chloride solution and sequencing amplification product, add the magnetic bead-sodium chloride solution and sequencing amplification product into anhydrous ethanol twice the volume of the sequencing amplification product, and use an oscillator to mix the sample to obtain a purified mixed system. Place the purified mixed system on a magnetic rack for 1 minute, add magnetic bead washing solution to wash impurities, wash four times, add 70 μL of elution buffer ethanol to elute the modified magnetic beads, and incubate at room temperature for 8 minutes to obtain the purified sequencing amplification product.

[0175] Comparative Example 1

[0176] The species identification method and extraction reagents used in this comparative example are similar to those used in Example 5. Unlike Example 5, the magnetic bead purification reagent does not include carbodiimide. Specifically, the magnetic bead purification reagent comprises the following raw materials in parts by weight: 55 parts porous nano-magnetic beads, 110 parts magnetic bead washing solution, 105 parts sodium chloride solution, 5 parts epoxysilane, and 5 parts triethanolamine.

[0177] Comparative Example 2

[0178] The species identification method and identification and extraction reagents in this comparative example are similar to those in Example 5. Unlike Example 5, the magnetic bead purification reagent does not include porous nanomagnetic beads. Specifically, the magnetic bead purification reagent includes the following raw materials in parts by weight: 120 parts magnetic bead washing solution, 100 parts sodium chloride solution, 5 parts epoxysilane, 10 parts carbodiimide, and 5 parts triethanolamine.

[0179] Comparative Example 3

[0180] The species identification method and extraction reagents used in this comparative example are similar to those used in Example 5. Unlike Example 5, the magnetic bead purification reagent does not include carbodiimide or porous nanomagnetic beads. Specifically, the magnetic bead purification reagent includes the following raw materials in parts by weight: 110 parts magnetic bead washing solution, 110 parts sodium chloride solution, 7 parts epoxysilane, and 10 parts triethanolamine.

[0181] Comparative Example 4

[0182] The species identification method and identification extraction reagent in this comparative example are similar to those in Example 5. The difference from Example 5 is that the porous nanomagnetic beads in the raw materials of the magnetic bead purification reagent are replaced with commercially available ordinary silica gel magnetic beads. The magnetic bead purification reagent includes the following raw materials in parts by weight: 60 parts of silica gel magnetic beads, 100 parts of magnetic bead washing solution, 100 parts of sodium chloride solution, 12 parts of epoxy silane, 15 parts of carbodiimide, and 15 parts of triethanolamine.

[0183] Performance testing:

[0184] The porous nanomagnetic beads prepared in Examples 1-5 of the present invention were taken, and the particle size of the porous nanomagnetic beads was detected by laser diffraction method. The particle size distribution was analyzed by the diffraction spot of the particles to the laser. During the test, 5 samples were collected in each group of Examples 1-5, and three batches were tested independently. The test temperature was 25.6 ° C, the shading rate was 10.5%, and the laser beam was used to irradiate the sample. The detector received the diffraction spot and then calculated the particle size distribution (based on Fraunhofer or Mie theory). The particle size test results are shown in Tables 1 and Figure 5 shown.

[0185] Table 1

[0186] Group Example 1 Example 2 Example 3 Example 4 Example 5 Particle size (nm) 59.1±0.5 60.2±0.3 59.1±0.2 60.2±0.3 61.3±0.4

[0187] As can be seen from Table 1, the porous nanomagnetic beads prepared in Examples 1-5 of the present invention have relatively uniform particle sizes, which proves that the porous nanomagnetic beads have the advantage of balancing adsorption efficiency and fluidity, which helps to improve adsorption efficiency and ensure adsorption purity.

[0188] The DNA adsorption capacity of the purified sequencing amplification products of Examples 1-5 and Comparative Examples 1-4 was tested. The DNA binding capacity per unit mass of magnetic beads was calculated by measuring the adsorption capacity of the magnetic beads for DNA solutions of known concentrations. The DNA adsorption capacity (ng / μg) test results are shown in Tables 2 and Figure 6 shown.

[0189] Table 2

[0190]

[0191] As can be seen from Table 2, the DNA adsorption amount of the sequencing amplification products after purification in Examples 1-5 of the present invention is higher than that of the sequencing amplification products after purification in Comparative Examples 1-4, among which the adsorption amount of Comparative Example 3 is the lowest. It can be seen that carbodiimide activates the carboxylic acid group through EDC, promoting the covalent binding of the magnetic bead surface to DNA. When it is missing, the adsorption amount decreases, and the porous nanomagnetic beads have a high specific surface area and can provide more binding sites. When it is missing, the adsorption capacity is limited.

[0192] The purified sequencing amplification products of Examples 1-5 and Comparative Examples 1-4 were subjected to DNA purity tests. The DNA purity (A260 / A280 ratio) and integrity (band uniformity) were assessed by agarose gel electrophoresis and UV spectrophotometry. The DNA purity test results are shown in Tables 3 and Figure 7 shown.

[0193] Table 3

[0194]

[0195] As can be seen from Table 3, the purity of the sequencing amplification products purified in Examples 1-5 of the present invention is higher than that in Comparative Examples 1-4. This shows that the surface modification of the ordinary silica gel magnetic beads in Comparative Example 4 is insufficient, such as the increase in nonspecific adsorption of residual proteins, resulting in lower DNA purity.

[0196] In summary, the present invention provides species identification reagents, identification methods and applications. In an embodiment of the present invention, a species identification method is provided. By synergizing identification extraction reagents and magnetic bead purification reagents, the sample tissue to be identified can be purified and impurities removed, and the sequencing quality of special samples such as degraded samples or samples containing inhibitors is significantly improved. It is particularly suitable for samples with high plant polysaccharide content. The identification analysis model adopts advanced DNA barcoding technology combined with a deep learning model, which can more accurately identify species, especially for morphologically similar and difficult to distinguish species. At the same time, the maximum likelihood method is introduced into the multiple sequence alignment layer (MSA), considering the base substitution frequency and inequality, and evaluating the support rate of the data comparison library sequence through Bootstrap resampling, thereby improving the accuracy of identification, realizing the evolutionary support rate of quantitative sequence variation, accurately distinguishing homologous variation from convergent evolution, simplifying the identification steps, and avoiding complex database classification storage and multiple search processes. Users only need to process and detect samples according to the provided operating procedures to quickly obtain accurate identification results.

[0197] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the scope of protection of the invention. Obviously, the embodiments described are only some embodiments of the present invention, rather than all embodiments. Based on these embodiments, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in this field can still combine, add, delete or make other adjustments to the features in the various embodiments of the present invention according to the circumstances without conflict, without making creative work, so as to obtain different other technical solutions that do not deviate from the concept of the present invention in essence, and these technical solutions also fall within the scope of protection of the present invention.

Claims

1. A method for species identification, characterized in that: The species identification method comprises: Obtaining a sample tissue to be identified, pre-treating the sample tissue to be identified with an identification extraction reagent to remove impurities in the sample tissue to be identified; The pre-treated sample tissue to be identified is subjected to PCR amplification, the amplified DNA product is extracted, and the sequencing amplification product is purified by agarose gel electrophoresis recovery method to obtain the purified amplified DNA product; The purified amplified DNA product is taken and subjected to sequencing reaction amplification to obtain a sequencing amplification product, which is purified by a magnetic bead purification reagent and bidirectionally sequenced to obtain a sequencing data set representing the tissue information of the sample to be identified; Capturing publicly available genomic information, pre-building a data comparison library based on the genomic information, and pre-training an identification and analysis model using the genomic information in the data comparison library to obtain an identification and analysis model based on maximum likelihood method and support vector machine; Load the sequencing dataset, identify and analyze the sequencing dataset based on the pre-trained identification analysis model, calculate the sequence similarity between the sequencing dataset and the data comparison library, and dynamically determine the similarity threshold based on the sequencing dataset to determine whether the sequence similarity exceeds the similarity threshold. If so, output the species identification result corresponding to the sample tissue to be identified.

2. The species identification method according to claim 1, wherein: The method for pre-treating the sample tissue to be identified using the identification extraction reagent specifically includes: Take the sample tissue to be identified, determine the tissue type based on the tissue acquisition source, and pre-prepare the identification extraction reagent. Wherein, the tissue type is plant tissue or animal tissue, and the identification extraction reagent includes 8mL of modified extraction solution, 180μL of mixed enzyme solution, and 400μL of proteinase K; If the tissue type is animal tissue, take 0.01-0.02g of animal tissue, cut it into pieces and put it into a 1.5mL centrifuge tube, add 100μL of modified extract solution to the centrifuge tube, mix it for 1 minute, add 5μL of proteinase K to the centrifuge tube, place the centrifuge tube containing the animal tissue sample in a centrifuge and centrifuge for 10-15 seconds, then shake it for 10 seconds, take the shaken centrifuge tube and place it in a metal bath for heating treatment, the heating conditions are 56℃ for 30 minutes and 95℃ for 10 minutes, after heating, place the centrifuge tube containing the animal tissue sample in a centrifuge and centrifuge for 3 minutes, and take the animal tissue supernatant; If the tissue type is plant tissue, the collected plant tissue is placed in liquid nitrogen for rapid cold preservation. During the test, the roots and stems of the plant tissue are cut into pieces, and then the plant tissue is ground with liquid nitrogen. During grinding, the plant tissue is placed in a mesh bag for filtration to release the plant tissue cells. 0.03 g of the ground plant tissue is weighed and placed in a centrifuge tube. Pure water is added to immerse the plant tissue. The mixed enzyme is added to the centrifuge tube at a volume ratio of 2:1 between the mixed enzyme and the plant tissue. After mixing evenly, the centrifuge tube is placed in a metal bath and enzymolyzed at 50-55°C for 6-8 hours. The centrifuge tube containing the enzymatic hydrolysis solution after enzymolysis is centrifuged for 1 minute, the residue is discarded, and the supernatant is recovered and placed in a high-speed centrifuge for 5 minutes. The precipitate is recovered and placed in a precipitate recovery tube. 100 μL of modified extract is added to the precipitate recovery tube. After mixing for 1 minute, 5 μL of proteinase K is added to the precipitate recovery tube. The precipitate recovery tube is centrifuged and heated, and the plant tissue supernatant is taken; The extracted animal tissue supernatant / plant tissue supernatant is placed in a purification tank, and the animal tissue supernatant / plant tissue supernatant is diluted with a dilution buffer to obtain a tissue dilution solution, which is extracted with chloroform and precipitated to obtain sample tissue DNA.

3. The species identification method according to claim 1, wherein: The dilution buffer comprises the following raw materials in percentage by volume: Sodium lauryl sulfate: 20%; Proteinase K: 5%; Isopropyl alcohol: 40%; Tris-EDTA: 30%; Polyvinylpyrrolidone: 5%; The modified extract comprises the following raw materials in percentage by volume: Chelex-100: 80%; β-glucanase: 5%; Citric acid: 5%; Disodium hydrogen phosphate: 10%.

4. The species identification method according to claim 3, wherein: The modified extract preparation method comprises: Weigh Chelex-100 resin and place it in a mixed solution of citric acid and disodium hydrogen phosphate. Add five times the volume of Tris-HCl solution to the mixed solution of Chelex-100 resin, citric acid, and disodium hydrogen phosphate, and vortex overnight until completely dissolved to obtain a Chelex solution. β-glucanase at a concentration of 30 U / μL was dissolved in Tris-HCl buffer, and then the β-glucanase was added dropwise to the Chelex solution. After gently inverting to mix, the solution was allowed to stand at room temperature for 30 minutes to prepare a modified extract.

5. The species identification method according to claim 1, wherein: The magnetic bead purification reagent comprises the following raw materials in parts by weight: 50-60 parts of porous nano-magnetic beads, 100-120 parts of magnetic bead washing solution, 100-120 parts of sodium chloride solution, 5-12 parts of epoxysilane, 10-15 parts of carbodiimide, and 5-15 parts of triethanolamine; The method for purifying the sequencing amplification product by using a magnetic bead purification reagent includes: The triethanolamine was dissolved in deionized water, and CTAB, sodium salicylate and triethanolamine solution were mixed in a mass ratio of 3:

1. After mixing evenly, the mixture was heated to 75-80°C and magnetically stirred for 30-45 minutes. Then, 0.5 times the mass of sodium salicylate was added to tetraethyl orthosilicate, and the mixture was stirred for 2 hours. The product was centrifuged and collected, and washed with anhydrous ethanol for 3-5 times. The prepared product was dispersed in 3 times the volume of concentrated hydrochloric acid, condensed and refluxed at 55-65°C for 5 hours, and the refluxed product was dispersed in 2 times the volume of anhydrous ethanol. Ammonia water and 3-aminopropyltriethoxysilane with a mass of 0.5 times the mass of tetraethyl orthosilicate were added to the mixed solution of the refluxed product and anhydrous ethanol, and stirred at room temperature for 5-6 hours at a stirring speed of 2000-2200 rpm to obtain porous nanomagnetic beads. Place porous nanomagnetic beads in a reactor and soak them in 30% sodium hydroxide solution for 15-20 minutes. Then pour the porous nanomagnetic beads into a round-bottom flask and wash them with 70% ethanol to remove non-polar residues on the surface of the porous nanomagnetic beads. Mix the porous nanomagnetic beads with isopropanol (3 times the volume of the porous nanomagnetic beads) and then ultrasonically clean them for 2-5 times before use. The porous nanomagnetic beads and epoxysilane were ultrasonically dispersed in anhydrous methanol 4 times the volume of the porous nanomagnetic beads in a mass ratio of 1:1, and then the temperature was raised to 65°C and stirred at a stirring speed of 500-520 rpm for 30-35 minutes. Carbodiimide and 3-aminopropyltriethoxysilane were mixed in MES buffer at a mass ratio of 2:1, and the mixed solution of carbodiimide and 3-aminopropyltriethoxysilane was mixed with the porous nanomagnetic beads at room temperature for 15 minutes to obtain a modified magnetic bead mixture. Triethanolamine was dissolved in MES buffer, and then 0.02 times the volume of triethanolamine was added to streptavidin to obtain a triethanolamine and streptavidin mixture. The triethanolamine and streptavidin mixture and the modified magnetic bead mixture were oscillated and reacted for 3 hours. The mixture was washed with MES buffer 3-5 times, and then the modified magnetic beads were mixed with sodium chloride solution to obtain a magnetic bead-sodium chloride solution. Take 0.5 μL of magnetic beads-sodium chloride solution and sequencing amplification product, add the magnetic beads-sodium chloride solution and sequencing amplification product into anhydrous ethanol twice the volume of the sequencing amplification product, and use an oscillator to mix the sample to obtain a purified mixed system. Place the purified mixed system on a magnetic stand for 1 minute, add magnetic bead washing solution to wash impurities, wash 3-5 times, add 50-100 μL of elution buffer ethanol to elute the modified magnetic beads, incubate at room temperature for 5-10 minutes, and obtain the purified sequencing amplification product.

6. The species identification method according to claim 1, wherein: The identification and analysis model uses the support vector machine as the initial model, and introduces a convolutional neural network into the initial model. The multiple sequence alignment layer MSA and the Bayesian hybrid model are embedded in the convolutional neural network. The maximum likelihood method is introduced into the multiple sequence alignment layer MSA. The multiple sequence alignment layer MSA is used to align the DNA sequences of multiple species in the sequencing data set, and the support rate of the data alignment library sequence is evaluated by Bootstrap resampling in combination with the maximum likelihood method considering the frequency and inequality of base substitutions. The hidden Markov model is introduced after the convolutional neural network. The hidden Markov model matches conserved domains through HMMER to identify subspecies or closely related species-specific variations, and combines k-mer similarity to determine the similarity between the sequencing data set and the data alignment library sequence.

7. The species identification method according to claim 6, wherein: The method for pre-training an identification and analysis model using genomic information in a data comparison library comprises: Pre-build a data comparison library, capture the published genome information of various species from public databases, organize and annotate the collected genome information, establish a local data comparison library, pre-process the genome information, annotate the genome sequences in the genome information, and obtain the unique labels corresponding to the genome sequences; Traverse the data comparison library, grab the training set, validation set, and test set from the data comparison library, and set the hyperparameters, loss function, and training rounds of the identification analysis model; Load the training set, encode it using one-hot encoding, freeze the hidden Markov model and support vector machine layers, pre-train the CNN-MSA architecture using the training set, generate a phylogenetic tree using the multiple sequence alignment layer (MSA), calculate bootstrap support, adjust sequence weights using a Bayesian mixture model, evaluate the SVM classification accuracy using the validation set after each round of training, and dynamically adjust hyperparameters; After training, HMMER and k-mer similarity are integrated and the similarity threshold is dynamically determined. To dynamically determine the similarity threshold, the HMMER layer is unfrozen, the test set sequence is input, conserved domains are matched, the k-mer similarity matrix is ​​calculated, and weighted fusion is performed with the BLAST results. The SVM and CNN-MSA architectures are jointly trained. The k-mer similarity constraint is added to the loss function of the CNN-MSA architecture, and the similarity threshold for a single label category is determined based on the loss function of the CNN-MSA architecture. Load the validation set, use the validation set to evaluate the pre-trained identification and analysis model, calculate the evaluation index, and check whether the identification and analysis model has overfitting or underfitting. If the evaluation index of the validation set is greater than the preset index threshold and the identification and analysis model does not have overfitting or underfitting, output the converged identification and analysis model.

8. The species identification method according to claim 7, wherein: The method for identifying and analyzing a sequencing data set based on a pre-trained identification and analysis model comprises: Obtain sequencing data sets, use FastQC to remove low-quality reads, convert clean reads to FASTA format, merge overlapping reads, and obtain preprocessed DNA sequences; The DNA sequence is converted into a numerical vector through k-mer encoding, and the feature extraction of the DNA sequence converted into a numerical vector is combined with a convolutional neural network to obtain a feature extraction set; The multiple sequence alignment layer (MSA) aligns the feature extraction set with the reference sequences in the data alignment library, and uses RAxML to calculate the phylogenetic tree and evaluate the Bootstrap support rate between the feature extraction set and closely related species. The reference sequences of closely related species with a coverage greater than 70% were screened and considered as reliable support. The hidden Markov model was used to match conserved domains through HMMER, identify subspecies or closely related species-specific variations, and calculate the matching score. The Jellyfish tool was used to count the k-mer frequencies of the feature extraction set, and the k-mer similarity matrix was constructed with the data alignment library. The cosine similarity was calculated and used as the sequence similarity. The similarity threshold corresponding to the sequencing data set was determined based on the HMMER layer. Determine whether the sequence similarity exceeds the similarity threshold. If so, output the reference sequence with the highest similarity, and use the reference sequence with the highest similarity as the species identification result corresponding to the sample tissue to be identified.

9. A species identification reagent for implementing the species identification method according to any one of claims 1 to 8, characterized in that: The species identification reagents include identification extraction reagents, magnetic bead purification reagents, agarose gel electrophoresis reagents, and identification gel recovery reagents.

10. Use of the species identification method according to any one of claims 1 to 8 in species identification.

Citation Information

Patent Citations

  • Kit and method for quickly extracting DNA (deoxyribonucleic acid) from colla corii asini

    CN104046616A

  • Bacterial nucleic acid sequencing identification method and bacterial identification kit based on DNA characteristic sequence

    CN108410971A

  • Method for identifying industrial cannabis sativa based on DNA bar code sequence

    CN116334197A

  • Kit for identifying species of animal cells and use method of kit

    CN118667972A

  • Lignin degrading microorganism species classification model and method based on machine learning

    CN118888012A