EDNA method for identifying bumblebee honeycomb in permafrost region of Qinghai-Tibet Plateau
By combining multi-source data with a Bayesian hierarchical model and liquid nitrogen freezing technology, along with eDNA extraction and high-throughput sequencing analysis, the problem of bumblebee nest identification in the permafrost region of the Qinghai-Tibet Plateau was solved, achieving rapid, low-cost, and high-precision nest identification and quantity estimation.
Patent Information
- Application Number
- CN202511090563.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-11-21
AI Technical Summary
In the permafrost region of the Qinghai-Tibet Plateau, traditional methods are difficult to identify bumblebee nests efficiently and accurately, especially in high-altitude and harsh environments where nest surveys are challenging. Furthermore, traditional methods are time-consuming, labor-intensive, and produce distorted results, making it impossible to effectively determine the activity and number of nests.
By using multi-source data combined with a Bayesian hierarchical model to locate observation transects, and by using liquid nitrogen freezing technology to completely excavate nests, combined with eDNA extraction and high-throughput sequencing analysis, a Bayesian hierarchical model was constructed to determine nest characteristics, eliminate interference, and achieve high-precision identification.
It enables rapid, low-cost, and non-destructive identification of bumblebee nests in the permafrost region of the Qinghai-Tibet Plateau, improving monitoring efficiency and reducing manpower and equipment costs. It is also applicable to the identification of underground insect nests in other environmentally sensitive areas.
Smart Images

Figure BDA0005533791640000061 
Figure BDA0005533791640000065 
Figure BDA0005533791640000067
Abstract
Description
Technical Field
[0001] This invention relates to the field of ecological monitoring technology in high-altitude permafrost regions, and in particular to an eDNA method for identifying bumblebee nests in the permafrost region of the Qinghai-Tibet Plateau. Background Technology
[0002] The Qinghai-Tibet Plateau has an average altitude of over 4,000 meters, with permafrost covering approximately 1.32 million square kilometers, accounting for 46% of its total area. This region possesses a unique and fragile ecosystem with relatively few pollinators of varying numbers and species. Bumblebees play a crucial role in this ecosystem, playing a vital role in maintaining the structure of the plateau's plant communities and the functions of the ecosystem. Currently, factors such as permafrost degradation and habitat fragmentation caused by climate change, as well as human activities, are increasingly threatening the survival and reproduction of bumblebee populations in the permafrost habitats of the Qinghai-Tibet Plateau. Therefore, effectively monitoring and assessing the health status and survival risks of bumblebee populations, accurately understanding their distribution, abundance, and key habitats (especially nest locations), and studying their habitat preferences and responses to environmental changes are of great significance for assessing their population status, understanding their ecological roles, their response mechanisms to climate change, and developing effective biodiversity conservation strategies.
[0003] However, conducting beehive surveys and identification in the extreme environment of the Qinghai-Tibet Plateau permafrost region faces severe challenges. High altitude, complex terrain, and harsh climate make field surveys exceptionally difficult. Traditional beehive identification relies heavily on specialized taxonomic knowledge and experienced experts, using visual morphological observation of bees and nest structures. This is particularly inconvenient in the high-altitude environment, often resulting in quantitative deficiencies, being time-consuming and labor-intensive, and producing distorted results. On the one hand, many beehives are hidden and difficult to find, or their structural characteristics may be atypical in the unique high-altitude environment, increasing the difficulty and error rate of morphological identification and making it impossible to determine whether a nest is solitary or social. On the other hand, traditional methods struggle to determine whether a target nest is in use, and due to the existence of closely related bee species with similar morphological characteristics, determining whether bumblebees are merely passing through or are the nest's inhabitants remains limited. This poses a significant challenge to estimating the number of bumblebees at a given location.
[0004] In recent years, environmental DNA (eDNA) technology, as a non-invasive biodiversity monitoring method, has achieved significant success in aquatic ecosystem research. Although eDNA technology has shown great potential in ecological monitoring, its application in terrestrial ecosystems, especially for specific biological structures (such as bumblebee nests), is still in the exploratory and developmental stage. The interior of a beehive is a complex and unique microenvironment, and its eDNA sources may include the bees themselves (adults, larvae, shed cells), food stored within the hive (pollen, nectar), symbiotic organisms (such as mites, microorganisms), and substances introduced from the external environment. How to effectively collect, enrich, and extract high-quality eDNA representative of the target bee species from the hive matrix, how to determine whether a nest is in use or abandoned, and how to eliminate interference from DNA of symbiotic organisms within the nest or organisms passing through the hive entrance are all key scientific problems and technical bottlenecks that urgently need to be solved for the application of this technology in beehive identification. Summary of the Invention
[0005] The technical problem to be solved by this invention is to provide a simple, rapid and low-cost method for identifying the eDNA of bumblebee nests in the permafrost region of the Qinghai-Tibet Plateau.
[0006] To address the above problems, the present invention provides a method for identifying the eDNA of bumblebee nests in the permafrost region of the Qinghai-Tibet Plateau, comprising the following steps:
[0007] S1 is based on multi-source data combined with a Bayesian hierarchical model to locate the observation transect and sampling point layout;
[0008] (1) Obtain multi-source data on bumblebees in the permafrost region of the Qinghai-Tibet Plateau over the years, including known distribution points, satellite remote sensing images, and UAV thermal infrared imaging observation data;
[0009] (2) Based on the analysis of known multi-source data related to nesting sites over the years, the probability of potential nesting site distribution within the transect is initially calculated using the maximum entropy model to narrow the observation range. This probability is then used as the prior information p(θ) in the Bayesian hierarchical model.
[0010] (3) Introduce a Bayesian hierarchical model, combine newly collected multi-source data to calculate the likelihood p(data|θ), update the posterior probability, and designate areas with ≥70% of the data as high-priority survey areas;
[0011] (4) Set up sampling points in high-priority survey areas:
[0012] Multiple core sampling points were set up along the intersection of bumblebee foraging paths, permafrost fissure zones, and micro-topographic depressions. The surrounding environment of suspected areas was carefully observed for debris accumulation, spider web coverage, and frequent bumblebee activity. Auditory and olfactory judgments were used to comprehensively determine whether a suspected bumblebee nest sampling point could be used, i.e., a suspected nest entrance. At the same time, field control sampling points were set up in the same habitat type area with an altitude difference of ≤50 meters and vegetation similarity of ≥80%, with one field control sampling point corresponding to each suspected nest entrance.
[0013] S2 performs deep freezing and complete excavation of the matrix surrounding and inside all sampling points;
[0014] All sampling points were subjected to layered freezing using a liquid nitrogen spray system to ensure uniform freezing of each layer of the nest and to stabilize the temperature of each layer at -196℃ for more than 30 seconds. Once the frozen block hardness was ≥80 Shore D, tools were used for excavation. The intact frozen blocks containing the complete nest structure and internal matrix were subjected to three-dimensional laser scanning, and the number, diameter, and spatial distribution characteristics of the nest cells were recorded to establish a three-dimensional digital model.
[0015] S3 implements stratified sampling, anti-degradation treatment, and establishes a spatial attribute database:
[0016] In a -20℃ ultra-clean cold room or liquid nitrogen operating table, the frozen nest blocks obtained in step S2 are systematically decomposed based on a three-dimensional digital model to establish a spatial attribute database, and the data is uploaded in real time using consortium blockchain technology; then, soil matrix samples, wax samples, fecal samples, and plant swab samples are collected according to the different structural layers after decomposition.
[0017] S4 eDNA extraction and purification:
[0018] The corresponding eDNA was extracted from each sample, and its integrity was confirmed by electrophoresis before purification.
[0019] S5 identifies the nest host based on PCR amplification, high-throughput sequencing, and sequence alignment analysis.
[0020] Based on the COI gene sequences of nine common bumblebee species from the Qinghai-Tibet Plateau, all samples in S4 were amplified by PCR, sequenced using high-throughput sequencing, and the DNA sequence information was digitized to obtain the original sequence. The original sequence was then filtered and denoised to generate high-fidelity amplicon sequence variants. These amplicon sequence variants were then screened at the genus level, validated for spatial distribution, and excluded by negative controls. Finally, if ≥2 core area samples from a sample group originating from the same frozen block simultaneously passed the above process, the nest host was determined to be a bumblebee.
[0021] S6 combines biological information to construct a Bayesian hierarchical model to determine hive characteristics:
[0022] Construct a Bayesian hierarchical model to determine hive characteristics; analyze bumblebee hives and output results including the determination of active hives, the estimated value of individual bumblebees in the hive and their corresponding posterior probabilities, and a spatial distribution heatmap.
[0023] In step S2, layered freezing refers to spraying the surface layer (0-10cm) continuously at a flow rate of 5L / min for 3 minutes, the middle layer (10-30cm) at a flow rate of 8L / min for 5 minutes, and the bottom layer (30-40cm) at a flow rate of 10L / min for 8 minutes.
[0024] In step S3, systematic decomposition refers to first layering along the natural structure of the nest, and then dividing the frozen block into 5 axial layers and 3 radial rings; simultaneously recording the substrate type, color according to the Munsell color card number, and texture evaluated by the proportion of sand, silt, and clay particles.
[0025] In step S3, the spatial attribute database contains data information accurate to the minute, such as sample collection time, initial temperature of frozen blocks, and sampling tool number.
[0026] In step S5, if ≥2 core area samples from a sample group originating from the same frozen block pass genus-level screening, spatial distribution verification, and negative control exclusion, then the nest host is determined to be a bumblebee.
[0027] In step S6, active beehives are determined as follows: when the bumblebee eDNA abundance is ≥100 reads / ng total eDNA, and the abundance of symbiotic bacteria in the sample is significantly positively correlated with the bumblebee eDNA abundance.
[0028] Compared with the prior art, the present invention has the following advantages:
[0029] 1. This invention utilizes environmental DNA (eDNA) technology to detect extremely small amounts of bumblebee eDNA fragments in the environment, which can be effectively detected even when the nest is buried deep underground or the signs of activity are weak.
[0030] 2. This invention involves filling the target nest with liquid nitrogen to cool the substrate and completely excavating it. It captures eDNA fragments (such as skin cells, excrement, mucus, hair, etc.) shed by organisms from samples of the nest entrance and internal environment (such as soil, sediment, air, etc.). By verifying universal primers with common local bumblebee species and performing PCR amplification, interference from other insects or environmental microorganisms is eliminated, thereby inferring whether the nest host is a bumblebee and determining the nest characteristics and the number of bumblebees in the nest.
[0031] 3. Compared with traditional methods that rely on manual exploration, ground-penetrating radar, or large drilling equipment, this invention is simpler and faster, requires less professional expertise from personnel, significantly improves the efficiency of large-scale screening and monitoring, and reduces labor and equipment costs; at the same time, samples are easy to preserve and transport.
[0032] 4. The eDNA technology in this invention has significant advantages such as high sensitivity, minimal interference to target organisms, no limitation on the life cycle stage of organisms, and simultaneous detection of multiple species. It is not only applicable to the permafrost region of the Qinghai-Tibet Plateau, with the core advantages of non-destructive and high sensitivity, but also applicable in principle to the identification and research of underground insect nests in other environmentally sensitive areas or areas where physical exploration is difficult. It is particularly suitable for monitoring rare, hidden, or inaccessible organisms. Detailed Implementation
[0033] This invention uses multi-source data combined with a Bayesian interlayer structure model to locate observation transects, collects and comprehensively analyzes multi-source genetic markers (including bumblebee-specific DNA, nest-associated microbial community characteristics, and genetic abundance) in the entrance and internal environmental media (such as soil and plant residues) of suspected bumblebee nests, and combines layered eDNA extraction, targeted amplification, high-throughput sequencing, and bioinformatics analysis to achieve accurate and efficient identification of bumblebee nests at specific locations. Furthermore, it uses a Bayesian model to determine nest activity and estimate the number of bumblebees within the nest. This significantly improves the feasibility and reliability of broad-coverage, high-precision identification and distribution assessment of bumblebee nests, a key pollinator, in the sensitive and fragile permafrost ecosystem of the Qinghai-Tibet Plateau. This enables the timely development of effective colony protection and habitat management strategies, ensuring the safety and stability of pollination services in alpine grassland ecosystems. The specific process is as follows:
[0034] A method for identifying the eDNA of bumblebee nests in the permafrost region of the Qinghai-Tibet Plateau includes the following steps:
[0035] S1 uses multi-source data combined with a Bayesian hierarchical model to locate observation transects and deploy sampling points:
[0036] (1) Obtain multi-source data on bumblebees in the permafrost region of the Qinghai-Tibet Plateau over the years, including known distribution points, satellite remote sensing images, and UAV thermal infrared imaging observation data.
[0037] Data on known distribution points of bumblebees in the permafrost region of the Qinghai-Tibet Plateau from 2018 to 2024 (a total of 327 valid records) were compiled. Environmental variables related to bumblebee nests, such as vegetation type, vegetation cover, and topography (elevation, slope, aspect, etc.), were extracted from Sentinel-2 satellite remote sensing imagery (one scene per quarter, totaling 28 scenes, covering over 85% of the permafrost area). Image processing software such as ENVI was used to classify and calculate the images, obtaining relevant information for each influencing unit in the study area. The data was imported into geographic information system software such as ArcGIS for processing and analysis, yielding data layers for elevation, slope, and aspect. Similarly, information on surface microtemperature anomalies was extracted from UAV thermal infrared imaging data (collected monthly during the growing season, totaling 35 times), covering 24 key areas, with each scene covering an area of 0.5 km × 0.5 km. Using thermal infrared image processing software, areas with temperatures 0.5-2℃ higher than the surrounding area and exhibiting a circular distribution were identified and converted into vector data format for subsequent overlay analysis with other data.
[0038] Collecting historical data from permafrost regions can provide information such as topography and temperature, helping to delineate large transects where bumblebees may nest.
[0039] (2) The probability of potential nesting sites within the transect is initially calculated using the maximum entropy model to narrow the observation range. This probability is then used as the prior information p(θ) in the Bayesian hierarchical model.
[0040] The multi-source data from S1 was input into the maximum entropy model (MaxEnt). Linear, quadratic, product, and threshold feature types were selected, and a 10-fold cross-validation method was used. The regularization parameter was set to 1.0 to balance model complexity and predictive ability. When running the model, 70% of the distribution points were used as training data and 30% as test data. The AUC value was calculated using the Receiver Operating Characteristic (ROC) curve to evaluate model performance. An AUC value ≥ 0.8 was considered a good predictive performance. The model then outputs a probability map of potential bumblebee nests, with probability values ranging from 0 to 1. Higher values indicate a greater likelihood of bumblebee nests in the area. Finally, ArcGIS software was used to reclassify the probability map, designating areas with a probability ≥ 0.6 as potentially high-probability areas to narrow the observation range. This probability data layer serves as prior information p(θ) in the Bayesian hierarchical model.
[0041] (3) Introduce a Bayesian hierarchical model, calculate the likelihood p(data|θ) by combining newly collected multi-source data, update the posterior probability, and designate areas with ≥70% of these areas as high-priority survey areas.
[0042] A Bayesian Hierarchical Model (BHM) is introduced to infer the location of bumblebee nest entrances. The distribution of bumblebee nests is considered a regional phenomenon influenced by multiple spatial and environmental factors. It is assumed that the outcome variable Y related to bumblebee nest entrances is observed in regions = 1, ..., n. i (For example, the probability of the nest entrance existing) is set using the Besag-York-Mollié (BYM) model as follows:
[0043] Y i =Normal(μ) i ,σ 2 )……(1)
[0044] μ i =z i β+u i +v i ……(2)
[0045] Where: μ i Y is within region i i The corresponding theoretical potential mean; σ 2 It is the variance of the error term; z i β is a fixed effect; z i =(1,z) i1 ,…,z ip ) is a vector consisting of the intercept for region i and p covariates (elevation (4000-5200 meters), slope aspect (preferably south and southeast), soil type (sand loam ≥ 60%), etc.), β = (β0, ..., β p ) represents the coefficient vector; u i This refers to spatial random effects, used to explain spatial dependence, indicating that regions that are close to each other may have similar values. It can be modeled using its intrinsic conditional autoregressive (CAR) model, which smooths the data based on a specific neighborhood structure. Specifically:
[0046]
[0047] Where: N is short for Normal distribution; u j It is the adjacent region j (where j∈δ) i Spatial random effects value, δ i and Let i represent the set of all other regions that are spatially associated with region i and the number of regions contained therein; This represents the conditional variance.
[0048] v i These are unstructured commutative components used to model uncorrelated noise.
[0049]
[0050] Where: the mean is 0; This represents the marginal variance.
[0051] The likelihood was calculated by combining the newly acquired multi-source data, and the posterior probability was updated according to the Bayesian formula. Finally, the areas within the transect with a posterior probability ≥70% were designated as high-priority survey areas.
[0052]
[0053] (4) Set up sampling points in high-priority survey areas:
[0054] Multiple core sampling points were set up along the intersection of bumblebee foraging paths (≤5 meters from nectar-producing plants), frozen soil fissures (0.1-0.5 meters wide), and micro-topographic depressions. The surrounding environment of suspected areas was carefully observed for debris accumulation (such as cleared soil particles and wax debris), the degree of spider web coverage (active nests typically do not have dense spider webs), and the frequency of individual bumblebees entering and exiting. Auditory judgment (listening close to the ground for continuous buzzing sounds inside the nest) and olfactory judgment (detecting the distinctive scents of honey, pollen, or beeswax) were used to comprehensively determine whether a location could be considered a suspected bumblebee nest sampling point, i.e., a suspected nest entrance. Simultaneously, to effectively eliminate background interference and assess detection specificity, field control sampling points were set up in the same habitat type area with an altitude difference ≤50 meters and vegetation similarity ≥80%, with one field control sampling point corresponding to each suspected nest entrance.
[0055] S2 performed deep freezing and complete excavation of the matrix surrounding and inside all sampling points:
[0056] Centered on the suspected nest entrance / on-site control sampling point, a 12cm radius and 40cm depth area was subjected to stratified freezing using a liquid nitrogen spraying system. The liquid nitrogen spraying system operated at a pressure of 0.6MPa and was equipped with a real-time temperature monitoring probe (accuracy ±0.5℃), recording the temperature every 10 seconds to ensure uniform freezing of all layers of the nest, preventing structural collapse caused by unfrozen or uneven freezing, and stabilizing the temperature of each layer at -196℃ for more than 30 seconds. Stratified freezing involved spraying the top 0-10cm layer at a flow rate of 5L / min for 3 minutes, the middle layer (10-30cm) at a flow rate of 8L / min for 5 minutes, and the bottom layer (30-40cm) at a flow rate of 10L / min for 8 minutes.
[0057] Then, a low-temperature hardness tester was used to test the hardness of the frozen block to be ≥80 Shore D. Three points were randomly selected for each frozen block to test the hardness, and the average value was used as the judgment basis. If the hardness of a single point was <80 Shore D, the corresponding area was re-frozen.
[0058] Next, excavation tools treated with DNA removal (soaked in 10% sodium hypochlorite for 30 minutes and then irradiated with ultraviolet light for 2 hours) were used to excavate the frozen areas of all field sampling points. The frozen blocks were wrapped in sterile aluminum foil throughout the process and placed in a -80℃ constant temperature transport box (temperature fluctuation ≤ ±2℃).
[0059] Three-dimensional laser scanning (accuracy 0.1 mm) was performed on the frozen blocks of nests that were completely removed, including the complete nest structure and internal matrix. The number of nest cells and their spatial distribution characteristics were recorded, and a three-dimensional digital model was established for subsequent stratified sampling planning.
[0060] S3 implements stratified sampling, anti-degradation treatment, and establishes a spatial attribute database:
[0061] In a -20℃ ultra-clean cold chamber or liquid nitrogen operating table, the frozen blocks of nests, containing complete nest structures and internal matrix obtained in step S2, were systematically dissociated based on a three-dimensional digital model to ensure the stability of the eDNA in the original samples. First, the frozen blocks were layered along the natural nest structure, and then divided into 5 axial layers (thickness proportional to the natural axial structure of the nest) and 3 radial rings (e.g., core area 0-6cm; transition area 6-12cm; outer area 12-20cm). The control frozen blocks were divided into 5 layers of equal height, and 3 rings were equally spaced with the geometric center as the origin. Then, the matrix type (soil, wax, feces), color according to the Munsell color chart number, and texture evaluated by the percentage of sand, silt, and clay particles were recorded simultaneously. Based on this, a spatial attribute database was established to ensure that subsequent analyses can be traced back to the original spatial location. The spatial attribute database contains data information accurate to the minute, such as sample collection time, initial frozen block temperature, and sampling tool number. It uses consortium blockchain technology to upload data in real time, providing a foundation for subsequent spatial distribution analysis of eDNA and ensuring the authenticity and reliability of the samples.
[0062] Soil matrix samples, wax samples, fecal samples, and plant swab samples were collected based on the different structural layers after dissociation, collecting the target amount as much as possible.
[0063] (1) Nest entrance channel layer: Carefully separate the frozen soil around the entrance channel to a depth of 0-5cm. Use a pre-cooled sterile sampling drill (2cm inner diameter) to set up sampling points in a radial ring. Set up 1-2 sampling points in each ring of each layer. After obtaining columnar samples, cut them into three sections: "top-middle-bottom". Take 6g of the matrix from each section and put them into 50mL Falcon tubes containing 25mL Longmire's preservation solution (containing 100mM Tris-HCl, 100mM EDTA, 100mM NaCl, 1% SDS, pH 8.0). Shake in an ice bath for 10 minutes (200rpm) and then immediately store in a -80℃ freezer.
[0064] (2) Core layer of the nest chamber: Dissect the frozen matrix inside the nest chamber layer by layer and collect special matrix. For the waxy layer of the nest wall, use a pre-cooled special sterile spoon to take a sample and weigh it accurately to 80mg using a microbalance (accuracy 0.1mg) in a sterile operating table; take 200mg of pollen storage chamber residue; take 30mg of larval feces. Add 2mL, 3mL and 1mL of pre-cooled protective solution containing 0.1% β-mercaptoethanol (to inhibit DNase activity) to the above samples respectively. At the same time, scrape the soil / waxy matrix from the inner wall of the nest chamber, search for and collect frozen bumblebee fecal particles, undigested pollen clusters, larval food residues or fragments of dead individuals, and transfer them to sterile cryovials using sterilized forceps that have been treated with DNA removal (soaked in 1% sodium hypochlorite for 30 minutes and then irradiated with ultraviolet light for 2 hours).
[0065] (3) Associated microbial layer: Soil samples reflecting the nest microenvironment were collected in the adjacent area of the nest chamber (not the direct nest chamber structure). The sampling method was the same as that of the entrance channel layer. The total amount of each sampling point was not less than 40g. The samples were put into a pre-cooled 50mL Falcon tube, immediately frozen in liquid nitrogen, and then stored at -80℃.
[0066] Repeat the above operation for the same locations as the control frozen block and the nest frozen block.
[0067] During the exfoliation process, all samples were given the corresponding protective solution in real time and thoroughly mixed, maintaining a low temperature environment below 0°C throughout. Each sample was precisely labeled with its X / Y / Z three-dimensional spatial coordinates (a rectangular coordinate system was established with the geometric center of the frozen block as the origin), and the data was simultaneously entered into the spatial attribute database.
[0068] S4 eDNA extraction and purification:
[0069] eDNA was extracted from each sample using a method adapted to its specific sample type. The eDNA was quantified using Qubit 4.0 to ensure a concentration ≥10 ng / μL. Quantification was performed three times for each sample, and the average value was taken. If the coefficient of variation was >10%, the eDNA was extracted again. Electrophoresis was performed on a 1.5% agarose gel. A main band ≥1000 bp was considered intact, with no obvious diffuse bands. Unqualified samples were re-extracted. Purification was then performed.
[0070] In the laboratory, different extraction and purification protocols are used for different types of samples:
[0071] (1) Soil matrix sample: Take 5g of the frozen and thawed matrix, add 15mL of CTAB extraction buffer (containing 2% CTAB, 1.4M NaCl, 20mM EDTA, 100mM Tris-HCl, pH 8.0) and 100μL of proteinase K (20mg / mL), and shake in a water bath at 65℃ (150rpm) for 2 hours. After centrifugation at 12000g for 5 minutes, take the supernatant and pass it through a PVPP column (500mg packing material) to remove polyphenols, and then perform secondary purification using the Qiagen DNeasy PowerSoil Pro Kit Soil DNA Extraction Kit (the operation is strictly in accordance with the kit instructions: mix the treated supernatant thoroughly with the lysis buffer in the kit, incubate at a suitable temperature to promote cell lysis, then centrifuge to adsorb DNA onto the centrifuge column, and obtain the preliminarily purified eDNA by elution), with a final elution volume of 50μL.
[0072] (2) Waxy samples: The waxy layer was dissolved using a chloroform-methanol mixture (2:1), centrifuged at 8000g for 10 minutes at 4℃ to collect the precipitate, 1-2 mL of sterile water was added to the precipitate, vortexed to mix, and centrifuged at 8000g for 5 minutes at 4℃. The supernatant was discarded. The sample was washed once to completely remove residual chloroform-methanol. The precipitate was retained after washing. Extraction was performed using the Bioline ISOLATE II Fecal DNA Kit (which uses high-density grinding beads in the kit to rapidly lyse the sample by bead milling, and purifies the lysate through a silica gel membrane column to remove impurities such as humic acid / polyphenols without the need for hazardous reagents such as phenols, and the procedure can be completed within 15 minutes). An additional RNase A (10 μg / mL) treatment step was added to remove RNA interference.
[0073] (3) Fecal sample: First, mix the fecal sample with the lysis buffer. Use the Bioline ISOLATE II Fecal DNA Kit to quickly and effectively lyse the sample using the high-density grinding beads in the kit. Purify the lysate through a silica gel membrane column to remove impurities such as humic acid / polyphenols and promote the release of high-quality DNA. If the sample contains a large amount of coarse fiber or undigested food particles, centrifuge at 10,000g for 1 minute before mixing with the lysis buffer to remove large particles and prevent them from clogging the grinding beads. The extraction can be completed in as little as 15 minutes.
[0074] (4) Plant swab samples: Place the swab samples into a tube containing lysis buffer, and release nucleic acids by destroying the cell structure through the lysis buffer; then add magnetic beads to make the nucleic acids bind to the magnetic beads, and use an external magnetic field to separate the magnetic beads that have bound nucleic acids. Use washing buffer to remove residual impurities such as proteins, polysaccharides, and cell debris from the surface of the magnetic beads; finally use elution buffer to release the pure nucleic acids from the magnetic beads, thus completing the extraction and purification.
[0075] S5 identifies the nest host based on PCR amplification, high-throughput sequencing, and sequence alignment analysis.
[0076] Based on the COI gene sequences of nine common bumblebee species from the Qinghai-Tibet Plateau (Source: Liu-Hao Wang, han Liu, Yu-Jie Tang, et al. Zookeys, 2020, 1007:1–21. DOI:10.3897 / zookeys.1007.34105), universal primers for the bumblebee genus level (Bombus-COI-F1: 5'-GGTCAACAAATCATAAAGATATTGG-3'; Bombus-COI-R1: 5'-TAAACTTCAGGGTGACCAAAAAATCA-3') (Source: Folmer, O., Black, M., Hoeh, W., Lutz, R., & Vrijenhoek, R. Molecular Marine Biology and Biotechnology, 1994, 3(5):294-299. DOI:10.1007 / BF02446414) Applicability verification was performed: (1) It was confirmed that the primer binding site was highly conserved in the COI sequences of the nine bumblebee species, which could effectively amplify all target species and avoid false negatives due to genetic differences; (2) The similarity of the primer pair with sequences of non-bumblebee closely related species (especially other bees that may coexist on the Qinghai-Tibet Plateau) was verified to be <85% using the Primer-BLAST tool, ensuring genus-level specificity; (3) Its amplification efficiency was confirmed to be stable (coefficient of variation CV <5%) in three replicate experiments. The amplification primers were used for subsequent experiments after verification that they met the above requirements. Otherwise, they were re-verified after local adjustments (optimization of conditions or replacement of primers).
[0077] PCR amplification was performed on all samples in S4 to enrich a 700 bp COI gene fragment from complex environmental samples. The PCR reaction system (25 μL) contained: 12.5 μL 2×Taq Pro Master Mix, 1 μL each of forward and reverse primers (10 μM), 2 μL extracted eDNA template, and 8.5 μL ddH2O. The reaction conditions were: 95℃ pre-denaturation for 5 minutes; followed by 35 cycles (94℃ denaturation for 30 seconds, 52℃ annealing for 45 seconds, and 72℃ extension for 1 minute); and a final extension at 72℃ for 10 minutes.
[0078] To ensure the reliability of the experimental results, a triple control sample was strictly set up: a positive control (containing 1 ng / μL of known bumblebee DNA) to verify the effectiveness of the system; a negative control (ddH2O without template) to monitor for exogenous contamination; and a blank extraction control (containing only extraction buffer) to exclude contamination during the extraction process.
[0079] After amplification, the product was verified by 1.2% agarose gel electrophoresis (suitable for separating 500-1000bp fragments) to confirm the specific band of the expected size (calculated based on the target gene sequence and primer binding sites, the expected size was 700bp). After confirming successful amplification, the product was purified using Agencourt AMPure XP magnetic beads, and the final product concentration was adjusted to 20ng / μL to prepare for subsequent high-throughput sequencing analysis.
[0080] After completing PCR amplification and obtaining a mixture containing bumblebee DNA fragments, the mixture was verified by agarose gel electrophoresis to ensure a single target band (no extraneous bands). Paired-end sequencing (2×300bp) was performed on the Illumina MiSeq platform to ensure that each sample obtained no less than 50,000 high-quality reads. The DNA sequence information was then digitized to obtain the original sequence.
[0081] The original sequences were then subjected to rigorous quality control: low-quality sequences (Phred quality value Q≥20) were filtered using Trimmomatic and adapter contamination was removed, retaining sequences ≥200bp in length; noise reduction was performed using DADA2 software to generate high-fidelity amplicon sequence variants (ASVs). The processed ASVs need to undergo: (1) genus-level screening, cross-matching with the COI sequences of the bumblebee genus in the bumblebee reference sequence library (sources include: "Illustrated Guide to Tibetan Honeybees of the Second Comprehensive Scientific Expedition to the Qinghai-Tibet Plateau", public databases such as BOLD, GenBank), and a sequence similarity of ≥95% is considered a successful match of the bumblebee genus; (2) spatial distribution verification: the relative abundance of bumblebee ASVs in the core area samples is ≥60% of the total bumblebee ASVs in the frozen block sample group; high abundance of bumblebee ASVs (relative abundance ≥10%) is detected in the wax layer of the nest chamber / larval feces samples, while the abundance of ASVs in the nest entrance passage layer is significantly lower than that in the core area (paired t test, p<0.01); (3) exclusion of negative controls: the relative abundance of bumblebee ASVs in the corresponding core area samples of the field control sampling points is extremely low (<0.1%) or not detected.
[0082] Ultimately, if ≥2 core area samples from a sample group originating from the same frozen block simultaneously pass genus-level screening, spatial distribution verification, and exclusion by negative controls, then the nest host is determined to be a bumblebee.
[0083] S6 combines biological information to construct a Bayesian hierarchical model to determine hive characteristics (including hive activity and the number of bumblebees):
[0084] Based on the data processed above, a Bayesian hierarchical model is constructed to determine honeycomb features.
[0085] The model uses the No-U-Turn-Sampler (NUTS) algorithm, sets up 4 Markov chains, and runs 5000 iterations per chain (including 1000 warm-up iterations). Convergence is evaluated by Rhat value (<1.01) and effective sample size (ESS>1000).
[0086] The first layer of the model (detection layer) calculates the detection probability of each ASV and sets the prior distribution to a Beta (1.5, 1.5) symmetric distribution (to suppress false positives by using zero boundary density and to optimize model convergence by using weak information constraints to accommodate low abundance ASV detection). The false positive rate is calibrated based on negative control sample data (≤0.01) to ensure detection reliability.
[0087] The second layer of the model (abundance layer): Using sequencing depth as a covariate, a negative binomial distribution model is established to estimate the absolute abundance of species (the absolute content of bumblebee eDNA in a unit sample). The formula is: Abundance = exp(α0 + β0 × log(sequencing depth) + ε), where α0 is the species-specific intercept; β0 is the depth effect coefficient; and ε is the random error term.
[0088] The third layer of the model (activity determination layer): The posterior probability of an active hive is calculated through Markov chain Monte Carlo (MCMC) simulation. When the bumblebee eDNA abundance is ≥100 reads / ng total eDNA, and the abundance of symbiotic bacteria (such as Lactobacillus spp.) in the sample is significantly positively correlated with the bumblebee eDNA abundance (r≥0.6, p<0.05), it is determined to be an active hive.
[0089] The fourth layer of the model (quantity estimation layer): Utilizing a Bayesian framework, the posterior distribution of λ_species output from the abundance layer is integrated, and the following abundance-quantity transformation relationship is derived through model derivation.
[0090] (Quantity ~ Normal(a×λ_species+b,σ_conv)……(6)
[0091] Where: λ_spectes refers to the posterior distribution characteristics of the absolute abundance of bumblebee species estimated by the abundance layer; a is the transformation coefficient; b is the basic offset term, both of which are determined by experimental calibration based on bumblebee characteristics; σ_conv is the transformation error term, which is given weak information prior (such as Half-Cauchy) and directly estimates the number of individual bumblebees in the nest and its uncertainty (posterior distribution) through MCMC simulation.
[0092] Finally, the bumblebee nests are analyzed, and the results include the determination of whether the nest is active (active / not active), the estimated value of each bumblebee in the nest and its corresponding posterior probability, and a spatial distribution heatmap.
[0093] Active hives were determined as follows: when the bumblebee eDNA abundance was ≥100 reads / ng total eDNA, and the abundance of symbiotic bacteria in the sample was significantly positively correlated with the bumblebee eDNA abundance (r≥0.6, p<0.05).
Claims
1. A method for identifying the eDNA of bumblebee nests in the permafrost region of the Qinghai-Tibet Plateau, comprising the following steps: S1 is based on multi-source data combined with a Bayesian hierarchical model to locate the observation transect and sampling point layout; (1) Obtain multi-source data on bumblebees in the permafrost region of the Qinghai-Tibet Plateau over the years, including known distribution points, satellite remote sensing images, and UAV thermal infrared imaging observation data; (2) Based on the analysis of known multi-source data related to nesting sites over the years, the probability of potential nesting site distribution within the transect is initially calculated using the maximum entropy model to narrow down the observation range. This probability is then used as prior information in the Bayesian hierarchical model. p ( θ ); (3) Introduce a Bayesian hierarchical model and calculate the likelihood using newly collected multi-source data. p ( data | θ Update the posterior probability and designate areas with ≥70% of these areas as high-priority survey areas; (4) Set up sampling points in high-priority survey areas: Multiple core sampling points were set up along the intersection of bumblebee foraging paths, permafrost fissure zones, and micro-topographic depressions. The surrounding environment of suspected areas was carefully observed for debris accumulation, spider web coverage, and frequent bumblebee activity. Auditory and olfactory judgments were used to comprehensively determine whether a suspected bumblebee nest sampling point could be used, i.e., a suspected nest entrance. At the same time, field control sampling points were set up in the same habitat type area with an altitude difference of ≤50 meters and vegetation similarity of ≥80%, with one field control sampling point corresponding to each suspected nest entrance. S2 performs deep freezing and complete excavation of the matrix surrounding and inside all sampling points; All sampling points were subjected to layered freezing using a liquid nitrogen spray system to ensure uniform freezing of each layer of the nest and to stabilize the temperature of each layer at -196℃ for more than 30 seconds. Once the frozen block hardness was ≥80 Shore D, tools were used for excavation. The intact frozen blocks containing the complete nest structure and internal matrix were subjected to three-dimensional laser scanning, and the number, diameter, and spatial distribution characteristics of the nest cells were recorded to establish a three-dimensional digital model. S3 implements stratified sampling, anti-degradation treatment, and establishes a spatial attribute database: In a -20℃ ultra-clean cold room or liquid nitrogen operating table, the frozen nest blocks obtained in step S2 are systematically decomposed based on a three-dimensional digital model to establish a spatial attribute database, and the data is uploaded in real time using consortium blockchain technology; then, soil matrix samples, wax samples, fecal samples, and plant swab samples are collected according to the different structural layers after decomposition. S4 eDNA extraction and purification: The corresponding eDNA was extracted from each sample, and its integrity was confirmed by electrophoresis before purification. S5 identifies the nest host based on PCR amplification, high-throughput sequencing, and sequence alignment analysis. Based on the COI gene sequences of nine common bumblebee species from the Qinghai-Tibet Plateau, all samples in S4 were amplified by PCR, sequenced using high-throughput sequencing, and the DNA sequence information was digitized to obtain the original sequence. The original sequence was then filtered and denoised to generate high-fidelity amplicon sequence variants. These amplicon sequence variants were then screened at the genus level, validated for spatial distribution, and excluded by negative controls. Finally, if ≥2 core area samples from a sample group originating from the same frozen block simultaneously passed the above process, the nest host was determined to be a bumblebee. S6 combines biological information to construct a Bayesian hierarchical model to determine hive characteristics: Construct a Bayesian hierarchical model to determine hive characteristics; analyze bumblebee hives and output results including the determination of active hives, the estimated value of individual bumblebees in the hive and their corresponding posterior probabilities, and a spatial distribution heatmap.
2. The method for identifying the eDNA of bumblebee nests in the permafrost region of the Qinghai-Tibet Plateau as described in claim 1, characterized in that: In step S2, layered freezing refers to spraying the surface layer (0-10cm) continuously at a flow rate of 5L / min for 3 minutes, the middle layer (10-30cm) at a flow rate of 8L / min for 5 minutes, and the bottom layer (30-40cm) at a flow rate of 10L / min for 8 minutes.
3. The method for identifying the eDNA of bumblebee nests in the permafrost region of the Qinghai-Tibet Plateau as described in claim 1, characterized in that: In step S3, systematic decomposition refers to first layering along the natural structure of the nest, and then dividing the frozen block into 5 axial layers and 3 radial rings; simultaneously recording the substrate type, color according to the Munsell color card number, and texture evaluated by the proportion of sand, silt, and clay particles.
4. The method for identifying the eDNA of bumblebee nests in the permafrost region of the Qinghai-Tibet Plateau as described in claim 1, characterized in that: In step S3, the spatial attribute database contains data information accurate to the minute, such as sample collection time, initial temperature of frozen blocks, and sampling tool number.
5. The method for identifying the eDNA of bumblebee nests in the permafrost region of the Qinghai-Tibet Plateau as described in claim 1, characterized in that: In step S5, if ≥2 core area samples from a sample group originating from the same frozen block pass genus-level screening, spatial distribution verification, and negative control exclusion, then the nest host is determined to be a bumblebee.
6. The method for identifying the eDNA of bumblebee nests in the permafrost region of the Qinghai-Tibet Plateau as described in claim 1, characterized in that: In step S6, active beehives are determined as follows: when the bumblebee eDNA abundance is ≥100 reads / ng total eDNA, and the abundance of symbiotic bacteria in the sample is significantly positively correlated with the bumblebee eDNA abundance.
Citation Information
Cited By
Gull gull breeding island habitat optimization method based on nest site preference and energy saving
CN121420849A