Method for monitoring macroalgae biodiversity of seaweed field based on environmental DNA technology
Polysaccharides and humic acid were removed by density gradient centrifugation and modified magnetic bead adsorption. PCR bias was corrected by combining specific amplicon primers and a hidden Markov model. Convolutional neural networks were used to identify base variation sites, and a spatiotemporal coupled dynamic evolution model was constructed. This solved the problem of trace DNA extraction and amplification bias in seaweed field sediments, and achieved efficient, accurate, and dynamic analysis of seaweed biodiversity monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG ACAD OF MARINE SCI (QINGDAO NAT MARINE SCI RES CENT)
- Filing Date
- 2026-06-24
- Publication Date
- 2026-07-21
AI Technical Summary
Existing technologies for extracting trace DNA from seaweed farm sediments are hampered by humic acid and polysaccharides, leading to the loss of DNA fragments from low-abundance key species. Furthermore, amplification bias masks signals from closely related rare species, making it difficult to accurately monitor seaweed biodiversity.
Polysaccharides and humic acid were removed by density gradient centrifugation and modified magnetic bead adsorption. Specific amplicon primers were designed and combined with a hidden Markov model to correct PCR amplification bias. Convolutional neural networks were used to identify base variation sites, and a spatiotemporal coupled dynamic evolution model was constructed. Environmental factors and species annotation results were integrated to eliminate systematic bias interference.
It significantly improves the recovery rate and extraction purity of low-abundance DNA, accurately identifies closely related rare species, dynamically analyzes the differences in seaweed community structure, and supports the tracking of ecological restoration processes.
Smart Images

Figure CN122428028A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of marine ecological monitoring technology, and in particular to a method for monitoring the biodiversity of large seaweeds in seaweed farms based on environmental DNA technology. Background Technology
[0002] As a typical nearshore marine ecosystem, seaweed farms have significant ecological service value. Their biodiversity status is a key indicator for assessing the health and function of the ecosystem. Traditional monitoring methods rely heavily on morphological identification, making it difficult to comprehensively capture the structure and spatiotemporal dynamics of seaweed communities. However, environmental DNA technology, by detecting genetic material in water or sediments, can efficiently identify the composition of seaweed species and has become an important tool for monitoring marine biodiversity.
[0003] Existing technologies for extracting eDNA from sediments are subject to strong interference from humic acids and polysaccharides in seaweed beds, leading to the easy loss of trace DNA fragments from low-abundance key species. Furthermore, in the context of low purity, strong amplification bias further masks the true signals of closely related rare species when amplifying homologous barcode regions of mixed seaweed communities. At the same time, under the dual constraints of incomplete extraction and identification bias, it is difficult to stably compare the community dynamics of artificial and natural seaweed beds, thus limiting the dynamic tracking of ecological restoration processes in seaweed biodiversity monitoring. Summary of the Invention
[0004] To overcome the aforementioned problems in the existing technology, this invention proposes a method for monitoring the biodiversity of large seaweeds in seaweed farms based on environmental DNA technology.
[0005] The technical solution adopted by this invention to solve its technical problem is as follows: This invention provides a method for monitoring the biodiversity of large seaweeds in seaweed farms based on environmental DNA technology, including the following steps: Step 1: Identify the target area, collect sediment and water samples from the seaweed farm, use density gradient centrifugation to remove macromolecular polysaccharides, retain low-abundance seaweed DNA components, effectively enrich trace amounts of seaweed DNA, and significantly improve the sensitivity of subsequent detection. Step 2: Based on the improved magnetic bead adsorption method, modified magnetic beads are used to target and remove humic acid, obtain DNA extraction purity data and integrity information, improve extraction purity and integrity, efficiently remove humic acid interference, and ensure high-quality DNA recovery and complete preservation. Step 3: Design seaweed-specific amplicon primers for the purified DNA, combine them with a hidden Markov model to correct PCR amplification bias, obtain a bias correction parameter set, suppress the masking of closely related species signals, accurately correct amplification bias, and restore the true abundance ratio of closely related species. Step 4: Perform high-throughput sequencing on the amplified products, introduce convolutional neural networks to identify base variation sites, perform debiased typing on mixed amplicon to distinguish closely related rare species in seaweed farms, correct species annotation results, establish a large seaweed biodiversity database of seaweed farms, support accurate identification, use deep learning to distinguish closely related species, and build a high-confidence dynamic database. Step 5: Integrate the DNA extraction purity data and bias correction parameter set extracted in the previous step to construct a spatiotemporal coupled dynamic evolution model of seaweed community, realize full-chain data alignment, and use temporal Bayesian network to hierarchically fuse environmental factors and species annotation results to dynamically analyze the differences in community structure between artificial and natural seaweed farms, eliminate comparative interference, eliminate systematic bias interference, and realize objective comparative analysis of community structure. Step 6: Deploy a graph convolutional long short-term memory network to drive the spatiotemporal coupled dynamic evolution model of seaweed communities, outputting a spatiotemporal dynamic map of biodiversity of large seaweeds in seaweed farms, supporting the tracking and evaluation of ecological restoration processes, realizing precise monitoring of the entire chain of biodiversity of large seaweeds in seaweed farms, dynamically visualizing the restoration process, and accurately identifying the ecological succession stage of seaweed farms.
[0006] Preferably, step 1 specifically includes: Surface sediments and bottom water samples were collected from the target seaweed field. After removing large particles by low-temperature centrifugation, the supernatant was transferred to a continuous density gradient medium constructed by iodixanol and sucrose in a volume ratio of 1:2. This effectively eliminated interference from coarse particles and ensured the efficiency of subsequent density gradient separation. Ultracentrifugation at 4℃ utilizes the significant difference in sedimentation coefficients between macromolecular polysaccharides and DNA components in a composite medium, causing polysaccharides to migrate to the upper layer of the gradient, while low-abundance DNA is enriched in a high-density characteristic layer with a density of 1.25-1.35 g / mL, achieving efficient stratified separation of polysaccharides and DNA components and significantly reducing interference from viscous inhibitors. A puncture-layer collection device was used to extract high-density characteristic layer components. After ultrafiltration desalting and nucleic acid concentration column combined treatment, a low-abundance seaweed DNA crude extract was obtained. The polysaccharide removal rate was increased to more than 95% while ensuring that the trace DNA recovery rate was not less than 70%. Low-abundance DNA components were obtained in a targeted manner to maximize the preservation of genetic information of key species.
[0007] Preferably, step 2 specifically includes: Bifunctional superparamagnetic nanoparticles with surface simultaneously modified with carboxyl and polyethyleneimine were mixed with low-abundance DNA crude extract at a mass-volume ratio of 1:5. The pH of the reaction system was adjusted to 4.0-5.0 to enhance the electrostatic adsorption of phenolic hydroxyl groups in humic acid molecules with the surface of magnetic beads, thereby enhancing the adsorption selectivity of humic acid and reducing the loss of target DNA. Incubation at 4°C under low-temperature oscillation conditions for a preset time, by introducing a competitive binding inhibitor (sodium pyrophosphate) to block the non-specific adsorption of magnetic beads on target DNA, allows modified magnetic beads to selectively capture humic acid, effectively protecting the integrity of target DNA and improving recovery efficiency; An alternating magnetic field is applied to cause magnetic beads to migrate in a directional manner and form a chain-like aggregate structure, thereby achieving efficient separation of the magnetic bead-humic acid complex. After collecting the supernatant, the sample is dynamically eluted with a low-salt buffer to obtain a preliminarily purified seaweed DNA sample. The humic acid magnetic bead complex is then rapidly separated to obtain a high-purity DNA elution buffer.
[0008] Preferably, step 2 further includes: The fragmentation level of pre-purified DNA samples was analyzed at the single-molecule level using a combination of nucleic acid fluorescent dyes and nanoflow cytometry. The absorbance at 230 nm, 260 nm and 280 nm was measured at the same time, and the A260 / A230 ratio was calculated to quantify the residual humic acid level. This effectively distinguishes between intact DNA and fragments and accurately quantifies the residual humic acid. The retention efficiency of high molecular weight DNA fragments was evaluated by pulsed field gel electrophoresis combined with fluorescence signal integration algorithm. The fragment length distribution entropy value was introduced as a DNA integrity index. The closer the value is to 1, the more uniform the fragment distribution is, so as to accurately evaluate the retention efficiency and fragment distribution uniformity of high molecular weight DNA. A two-dimensional decision boundary for purity and integrity is constructed (A260 / A230 ≥ 2.0 and DNA integrity index ≥ 0.75). DNA samples that do not meet the threshold are automatically rejected and a backtracking optimization process is triggered (adjusting magnetic bead incubation time or competitive inhibitor concentration). Qualified DNA is screened for subsequent amplicon library construction. The dual indicators strictly control DNA quality, and automatic optimization ensures successful amplification despite obstacles.
[0009] Preferably, step 3 specifically includes: Multiple specific amplicon primer pairs with molecular interior studs were designed for brown algae and red algae, targeting the V4 region of the 18S rRNA gene and the 5' hypervariable region of the COI gene. The mismatch recognition ability was improved by locking nucleic acid at the 3' end of the primers, which effectively reduced non-target amplification and improved the differentiation of closely related species. By constructing an orthogonal gradient calibrator simulating seaweed communities for PCR amplification, a three-parameter bias response surface covering primer dimer formation, GC bias, and template secondary structure was established to achieve multidimensional quantification of amplification bias and improve the fitting accuracy of the calibration model. Based on the Hidden Markov Model, a state transition penalty factor is introduced to learn the nonlinear preference pattern of base incorporation in the amplification cycle. Inverse weighted correction is performed on the early stage of exponential amplification of low abundance templates to generate a bias correction parameter set to suppress the relative masking of signals from closely related species, significantly restore the true abundance of low abundance species, and avoid the signal being masked by dominant species.
[0010] Preferably, step 4 specifically includes: After quality filtering and chimera removal of high-throughput sequencing data, a multi-level base context coding strategy is used to convert each read into a three-dimensional feature tensor containing the coupling relationship between adjacent triplet sites, effectively preserving the evolutionary conservation and context-dependent structure of the base sequence. A residual convolutional neural network architecture with an attention mechanism was constructed. Using a haplotype reference set of closely related species verified by single-cell sequencing as the training object, the long-range dependencies between variant sites and species-specific evolutionary fingerprints were learned, which significantly improved the classification boundary discrimination of closely related species in high-dimensional feature space. Hierarchical progressive classification reasoning is performed on mixed amplicon reads. First, clustering is performed according to genus-level features, and then fine typing is performed according to species-level variation sites. The species labels corresponding to each read and the confidence scores after Bayesian calibration are output, which greatly improves the accuracy and reliability of identifying low-abundance closely related rare species.
[0011] Preferably, step 4 further includes: Based on the species label confidence score output by the convolutional neural network, the graph cut algorithm is used to globally optimize the allocation boundary of low confidence reads. By constructing a sequence similarity graph, the co-amplified closely related rare species are modularly segmented, which significantly improves the resolution of closely related rare species and reduces the classification error rate. The debiased typing results are compared with dynamically updated Hidden Markov alignment maps. An arbitration mechanism based on phylogenetic inconsistency scoring is introduced to correct species annotations for alignment conflict areas, which effectively resolves alignment conflicts and improves the consistency and accuracy of species annotations. By integrating the revised species annotation results, an incremental Bayesian update strategy was adopted to construct a large-scale algal biodiversity database covering all sample plots of the target seaweed farm. This database supports real-time assimilation and anomaly detection of new sequencing data, enabling real-time updates and dynamic maintenance of the database and enhancing the early warning capability for abnormal changes.
[0012] Preferably, step 5 specifically includes: The obtained DNA purity and integrity data, bias correction parameter set, and species abundance table after bias removal and typing by convolutional neural network are jointly standardized based on Bayesian hierarchical framework to eliminate the differences in dimensionality and error distribution between multiple steps, effectively unify the scale of multi-source data, and improve the accuracy of subsequent fusion modeling. Using the three-dimensional spatial coordinates (longitude, latitude, and water depth) of sampling points and multi-time point sampling windows as coupling nodes, a multi-layer heterogeneous data association matrix is constructed to achieve backpropagation alignment of errors across the entire chain, from DNA extraction purity constraints to species identification confidence, enabling error tracing and correction layer by layer and enhancing the consistency of data across the entire chain. A spatiotemporally non-stationary Gaussian process regression method is used to replace traditional Kriging interpolation. Uncertainty quantification is performed on the community composition inference for deletion sites where DNA extraction fails due to humic acid interference. This generates a spatiotemporally continuous probability field of algal species distribution, quantifies the uncertainty of deletion site inference, and ensures the spatiotemporal continuity of the community distribution map.
[0013] Preferably, step 5 further includes: High-frequency time-series environmental factor data from each sampling point are collected, including water temperature, salinity, dissolved oxygen, turbidity, bottom sediment particle size distribution, and nutrient molar ratio. The data are then processed by adaptive Kalman filtering for non-equidistant time series interpolation and normalization to ensure the continuous comparability of environmental factors over time and improve the stability of the model input. A four-layer temporal Bayesian network model with hidden state layers was constructed, with environmental factors as observation parent nodes, DNA extraction purity and bias correction parameters as hidden state nodes, and species annotation results as output child nodes. The model was used to learn the differences in conditional dependence structure and causal path between artificial and natural seaweed farms, and to reveal the key environmental driving factors and causal paths of different seaweed farm types. By comparing posterior probabilities, the system offset in community structure between artificial seaweed farms under construction, mature artificial seaweed farms, and natural seaweed farms is inferred dynamically. This eliminates ecological comparison interference caused by humic acid interference, amplification bias, and loss of rare species, thereby achieving quantitative analysis of community structure differences and eliminating the impact of multi-source interference on comparison.
[0014] Preferably, step 6 specifically includes: Using each sampling point as a graph node, a dynamic adjacency matrix is generated by a composite spatial distance kernel constructed based on hydrological connectivity and similarity of substrate types. This forms an adaptive graph structure that can update topological relationships with the seasons, effectively capturing the spatial dependence of seaweed farms and improving the accuracy of community structure modeling. The posterior species probability field and uncertainty interval output by the temporal Bayesian network are used as the input sequence. Spatial dependency features are extracted through graph attention convolutional layers with residual connections. Then, the evolutionary patterns of forward and backward time windows are captured through a bidirectional long short-term memory network. This fully integrates spatiotemporal features and significantly improves the accuracy of species dynamic change prediction. The parameters of the joint adversarial training graph convolutional network and long short-term memory network are incorporated with DNA-extracted quality labels as regularization constraints. The output is a spatiotemporal dynamic map of biodiversity of large seaweed farms, including species composition, migration paths of dominant species, and appearance / disappearance events of rare species. This map is used to identify the stages of ecological restoration in artificial seaweed farms, enhance the robustness of the model, and enable quantifiable identification of ecological restoration stages.
[0015] Compared with the prior art, the beneficial effects of the present invention are: 1. This method for monitoring the biodiversity of large seaweeds in seaweed farms based on environmental DNA technology effectively removes highly interfering macromolecular polysaccharides and humic acids from seaweed farm sediments through the synergistic effect of density gradient centrifugation and modified magnetic bead adsorption. It significantly improves the recovery rate and extraction purity of trace DNA from low-abundance key species, providing high-quality templates for downstream species identification.
[0016] 2. This method for monitoring the biodiversity of large seaweeds in seaweed farms based on environmental DNA technology establishes a differentiated reverse weighted correction mechanism for templates of different abundances through specific primer design and amplification bias correction driven by a hidden Markov model. This significantly reduces the signal suppression effect of high abundance templates on low abundance templates, inhibits the relative masking of closely related rare species signals by amplification bias, and improves the detection accuracy of low abundance species in mixed communities.
[0017] 3. This method for monitoring the biodiversity of large seaweeds in seaweed farms based on environmental DNA technology constructs a multi-layer Bayesian network that integrates environmental factors, DNA purity information, and amplification bias parameters. Through posterior probability comparison inference and system offset quantification, it achieves dynamic analysis of the differences in community structure between artificial and natural seaweed farms, effectively eliminating ecological comparison interference caused by incomplete extraction and identification bias. Attached Figure Description
[0018] Figure 1 This is a schematic diagram illustrating the workflow of the method for monitoring the biodiversity of large seaweeds in seaweed farms based on environmental DNA technology, as described in this invention. Figure 2 This is a line graph showing the effect of the Hidden Markov Model correction of this invention on the amplification efficiency of seaweed templates with different abundances; Figure 3 This is a scatter plot showing the distribution of initial copy number of seaweed species and CNN annotation confidence in this invention. Figure 4 This is a graph showing the fitting relationship between the purity of the eDNA and the number of macroalgae species detected in this invention. Detailed Implementation
[0019] To enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0020] Example 1 Please see Figures 1 to 4 This invention provides a technical solution: a method for monitoring the biodiversity of large seaweeds in seaweed farms based on environmental DNA technology, comprising the following steps: Step 1: Identify the target area and collect sediment and water samples from the seaweed beds. Density gradient centrifugation is used to remove large polysaccharides while retaining low-abundance seaweed DNA components, effectively enriching trace amounts of seaweed DNA and significantly improving subsequent detection sensitivity. Surface sediment and bottom water samples are collected from the target seaweed beds. After removing large particles through low-temperature centrifugation, the supernatant is transferred to a continuous density gradient medium constructed from iodixanol and sucrose at a volume ratio of 1:2. This effectively eliminates interference from coarse particles and ensures efficient density gradient separation. Ultracentrifugation is then performed at 4°C, utilizing the interaction between large polysaccharides and DNA components in the composite medium. The significant difference in sedimentation coefficients allows polysaccharides to migrate to the upper layer of the gradient, while low-abundance DNA accumulates in a high-density characteristic layer with a density of 1.25-1.35 g / mL. This achieves efficient separation of polysaccharide and DNA components, significantly reducing interference from viscous inhibitors. A puncture-based stratification collection device is used to extract components from the high-density characteristic layer. After ultrafiltration desalting and combined treatment with a nucleic acid concentration column, a crude extract of low-abundance algal DNA is obtained. This increases the polysaccharide removal rate to over 95% while ensuring a trace DNA recovery rate of no less than 70%. This targeted acquisition of low-abundance DNA components maximizes the preservation of key species genetic information.
[0021] It should be noted that samples were immediately placed in an ice bath after collection and returned to the laboratory within 4 hours for pretreatment. The samples were centrifuged at 4°C and 5000×g for 15 minutes to remove large particles such as coarse sand, mineral particles, and cell debris. The supernatant was then slowly transferred to a pre-prepared continuous density gradient medium, which was a mixture of iodixanol and sucrose at a volume ratio of 1:2, with a density range of 1.10–1.40 g / mL. Subsequently, the centrifuge tubes were placed in an ultracentrifuge at 4°C and 100,000×g. After centrifugation for 3 hours, the large polysaccharides, due to their low sedimentation coefficient, migrated to the upper layer of the gradient, while the low-abundance DNA fraction was enriched in a high-density characteristic layer with a density of 1.25–1.35 g / mL. This region appeared as a light yellow transparent band, which was easily identifiable. A puncture-based stratification collection device was used to precisely extract the high-density characteristic layer. Specifically, the puncture needle was slowly lowered to the bottom of the characteristic layer, and the chromatography medium was continuously extracted at a flow rate of 0.5 mL / min, collecting approximately 2.5 mL of enriched solution. This enriched solution was then transferred to an ultrafiltration tube (molecular weight cutoff 30 kDa). The sample was desalted by centrifugation at 4℃ and 4000×g for 20 minutes to remove iodixanol and sucrose residues. The retentate from the ultrafiltration tube was collected and transferred to a nucleic acid concentrating column. Following the column instructions, two volumes of binding buffer were added, the mixture was vortexed and allowed to stand for 2 minutes, then centrifuged at 12000×g for 1 minute. The flow-through was discarded, and 750 μL of washing buffer was added for a second wash. Finally, 50 μL of elution buffer preheated to 55℃ was added, the mixture was allowed to stand for 2 minutes, and then centrifuged at 12000×g for 1 minute. The collected eluent was the crude low-abundance algal DNA. Extraction: The crude DNA extract, after processing, ensured a polysaccharide removal rate of over 95% as determined by the phenol-sulfuric acid method. The micro-DNA recovery rate was quantitatively assessed using the Qubit fluorescence method, showing a result of no less than 70%. This significantly reduced the residual amount of PCR inhibitors such as humic acid in the crude extract. The A260 / A230 ratio was between 1.6 and 1.9, indicating that its purity met the basic requirements for subsequent nucleic acid amplification experiments. At the same time, 1% agarose gel electrophoresis confirmed that the high molecular weight genomic DNA was well preserved without obvious tailing or diffuse degradation.
[0022] Step 2: Based on the improved magnetic bead adsorption method, modified magnetic beads are used to target and remove humic acid, obtaining DNA extraction purity and integrity information, improving extraction purity and integrity, efficiently removing humic acid interference, and ensuring high-quality DNA recovery and complete preservation. Bifunctional superparamagnetic nanoparticles with surfaces simultaneously modified with carboxyl and polyethyleneimine are mixed with low-abundance crude DNA extract at a mass-to-volume ratio of 1:5. The pH of the reaction system is adjusted to 4.0-5.0 to enhance the electrostatic adsorption between the phenolic hydroxyl groups in humic acid molecules and the magnetic bead surface, thereby enhancing the selectivity of humic acid adsorption and reducing... The target DNA was lost, and the mixture was incubated at 4°C under low-temperature oscillation for a preset time. By introducing a competitive binding inhibitor (sodium pyrophosphate), the non-specific adsorption of the magnetic beads on the target DNA was blocked, allowing the modified magnetic beads to selectively capture humic acid, effectively protecting the integrity of the target DNA and improving recovery efficiency. An alternating magnetic field was applied to cause the magnetic beads to migrate in a directional manner and form a chain-like aggregate structure, achieving efficient separation of the magnetic bead-humic acid complex. After collecting the supernatant, the sample was dynamically eluted with a low-salt buffer to obtain a preliminarily purified seaweed DNA sample. The humic acid magnetic bead complex was quickly separated to obtain a high-purity DNA eluent.
[0023] It should be noted that bifunctional superparamagnetic nanoparticles with surfaces simultaneously modified with carboxyl and polyethyleneimine were used as the humic acid adsorption medium. These magnetic beads had a particle size range of 150-200 nm and a saturation magnetization of not less than 50 emu / g. The magnetic beads were mixed with the obtained low-abundance algal DNA crude extract at a mass-to-volume ratio of 1:5, i.e., 1 mg of magnetic beads was added to every 5 mL of crude extract. The pH of the mixture was slowly adjusted to 4.5 using 0.1 M hydrochloric acid. Under this pH condition, the phenolic hydroxyl groups in the humic acid molecules were less deprotonated, resulting in enhanced electrostatic adsorption with the carboxyl and amine groups on the surface of the magnetic beads. The mixture was placed in a low-temperature shaking incubator and incubated at 4°C with shaking at 150 rpm for 30 minutes to ensure sufficient contact and stability between the magnetic beads and the humic acid. The entire process was performed in a clean bench to avoid exogenous DNA contamination. Sodium pyrophosphate was added to the incubation system as a competitive binding inhibitor at a final concentration of 10 mM. Pyrophosphate ions preferentially occupy the non-specific DNA binding sites on the magnetic bead surface, blocking the electrostatic adsorption of the target DNA fragment by the magnetic beads, without affecting the affinity of the magnetic beads for humic acid. After incubation, an alternating magnetic field with a frequency of 50 Hz and a field strength of 0.5 T was applied to the reaction tube for 3 minutes, causing the magnetic beads to migrate directionally along the magnetic field lines and form a chain-like aggregate structure. After the magnetic field was removed, the magnetic bead-humic acid complex quickly settled to the bottom of the tube. The supernatant was carefully collected using a pipette and transferred to a new centrifuge tube. An equal volume of low-salt buffer (10 mM) was added to the collected supernatant. The mixture was gently inverted and mixed with mM Tris-HCl (pH 8.0, 50 mM NaCl). Dynamic elution was performed at 4°C for 10 minutes to remove residual sodium pyrophosphate and small amounts of nonspecific conjugates. After dynamic elution, an alternating magnetic field (50 Hz, 0.5 T) was applied again for 2 minutes to ensure that any remaining magnetic beads were completely aggregated on the tube wall or bottom. While maintaining the magnetic field, the supernatant was transferred to a new sterile centrifuge tube. This supernatant is the preliminarily purified seaweed DNA sample. The A260 / A230 ratio of the processed sample was increased from 1.6-1.9 in the crude extract stage to 2.1-2.3, and the residual humic acid was significantly reduced. The sample was verified by 1% agarose gel electrophoresis, showing intact high-molecular-weight DNA bands without significant tailing or degradation. This preliminarily purified product can be directly used for subsequent PCR amplification and library construction without additional purification steps. All operations were performed at room temperature. The magnetic beads were for single use only and not recycled.
[0024] Furthermore, step 2 also includes: using nucleic acid fluorescent dyes combined with nanoflow cytometry to analyze the fragmentation level of the pre-purified DNA sample at the single-molecule level, simultaneously measuring the absorbance at 230nm, 260nm, and 280nm, calculating the A260 / A230 ratio to quantify the residual humic acid level, effectively distinguishing between intact DNA and fragments, accurately quantifying residual humic acid, evaluating the retention efficiency of high molecular weight DNA fragments through pulsed-field gel electrophoresis combined with a fluorescence signal integration algorithm, introducing the fragment length distribution entropy value as a DNA integrity index, the closer the value is to 1, the more uniform the fragment distribution, accurately evaluating the retention efficiency of high molecular weight DNA and the uniformity of fragment distribution, constructing a purity-integrity two-dimensional decision boundary (A260 / A230 ≥ 2.0 and DNA integrity index ≥ 0.75), automatically removing DNA samples that do not meet the threshold and triggering a backtracking optimization process (adjusting the magnetic bead incubation time or the concentration of competitive inhibitors), screening qualified DNA for subsequent amplicon library construction, strictly controlling DNA quality with dual indicators, and automatically optimizing to ensure successful amplification despite obstacles.
[0025] It should be noted that the use of SYBR Gold nucleic acid fluorescent dye combined with nanoflow cytometry was employed to analyze the fragmentation level of pre-purified algal DNA samples at the single-molecule level. The DNA samples were diluted to 0.5 ng / μL, and SYBR Gold dye was added to a final concentration of 1×. After incubation in the dark for 5 minutes, the samples were analyzed using the nanoflow cytometer, which was set to an excitation wavelength of 488 nm and collected fluorescence signals at 530 nm. At least 10,000 fluorescence events were collected from each sample. Simultaneously, 2 μL of... The DNA stock solution was placed in an ultra-micro spectrophotometer, and the absorbance at 230 nm, 260 nm, and 280 nm was measured. The A260 / A230 ratio was automatically calculated, which was used to quantify the residual humic acid level: a ratio below 1.8 indicates significant humic acid interference, between 1.8 and 2.0 is acceptable, and 2.0 and above indicates excellent purity. Nanoflow cytometry results showed that intact DNA molecules produced uniform high-intensity fluorescence pulses, while fragmented DNA exhibited low-intensity and discretely distributed pulse signals. The proportion of intact DNA molecules in the sample could be quantitatively assessed by the distribution of pulse peak height and peak width. Pulsed-field gel electrophoresis combined with a fluorescence signal integration algorithm was used to evaluate the retention efficiency of high molecular weight algal DNA fragments. A 1% pulsed-field grade agarose gel was prepared, and the electrophoresis buffer was 0.5×TBE. The electrophoresis conditions were set as follows: initial conversion time 0.5 seconds, final conversion time 15 seconds, voltage gradient 6 V / cm, angle 120°, electrophoresis time 12 hours, temperature 14℃, and 200 ng sample per well. DNA was added along with a high molecular weight DNA marker. After electrophoresis, SYBR Gold staining was performed for 30 minutes, and images were acquired using a gel imaging system. ImageJ software was used to integrate the fluorescence signal in the lanes, dividing the electrophoretic pattern into three intervals based on molecular weight: <10kb, 10-50kb, and >50kb. The fragment length distribution entropy was introduced as a DNA integrity index, calculated using the following formula: ,in The percentage of fluorescence signal in each molecular weight range is represented by H, which is the DNA integrity index. A value closer to 1 indicates a more uniform distribution of fragments across molecular weight ranges and more complete preservation of high-molecular-weight DNA. An H value below 0.75 indicates significant degradation in the sample. A two-dimensional decision boundary between purity and integrity is constructed as the DNA sample qualification standard. The horizontal axis represents the A260 / A230 ratio, and the vertical axis represents the DNA integrity index. The judgment rule is set as follows: samples with A260 / A230 ≥ 2.0 and a DNA integrity index ≥ 0.75 are considered qualified and allowed to proceed to the subsequent amplicon library construction process. Samples failing to meet either indicator are considered unqualified, and an automatic backtracking optimization process is triggered for unqualified samples: if A260 / A230 < 2.0 but the integrity index ≥ 0.75, it indicates high humic acid residue, and the magnetic bead incubation time is extended from 30 minutes to 45 minutes; if the integrity index < 0.75 and A260 / A230 ≥ 2.0, it indicates severe DNA shearing, and the final sodium pyrophosphate concentration is reduced from 10... The mM is reduced to 5 mM; if both indicators fail to meet the standard, the two parameters are adjusted simultaneously, and the optimized sample is re-executed in step 2 until the two-dimensional decision boundary threshold is met. Qualified DNA samples are stored at -20℃ for later use.
[0026] Step 3: For the purified DNA, design seaweed-specific amplicon primers, and use a hidden Markov model to correct PCR amplification bias, obtain a bias correction parameter set, suppress signal masking by closely related species, accurately correct amplification bias, and restore the true abundance ratio of closely related species. For brown algae and red algae, design multiplex specific amplicon primer pairs containing molecular internal nail structures to target the 18S... By modifying the V4 region of the rRNA gene and the 5' hypervariable region of the COI gene with 3' locked nucleic acid modification of the primers, the mismatch recognition ability is enhanced, effectively reducing non-target amplification and improving the differentiation of closely related species. PCR amplification is performed by constructing an orthogonal gradient calibrator simulating seaweed communities, and a three-parameter bias response surface covering primer dimer formation, GC bias, and template secondary structure is established to achieve multidimensional quantification of amplification preference and improve the fitting accuracy of the calibration model. Based on a hidden Markov model, a state transition penalty factor is introduced to learn the nonlinear preference pattern of base incorporation in the amplification cycle. Inverse weighted correction is performed on the initial stage of exponential amplification of low-abundance templates to generate a bias correction parameter set to suppress the relative masking of signals from closely related species, significantly restore the true abundance of low-abundance species, and avoid the signal being masked by dominant species.
[0027] It should be noted that for seaweed species in the phyla Phaeophyta and Phaeophyta, multiplex specific amplicon primer pairs containing molecular internal studs were designed to target the V4 region of the 18S rRNA gene and the 5' hypervariable region of the COI gene, respectively. The primer length was 22-28 nucleotides, the annealing temperature was controlled between 58-62℃, and the GC content was 45%-55% (GC content refers to the percentage of guanine (G) and cytosine (C) bases in a DNA or RNA molecule). A locked nucleic acid modification monomer was introduced at the penultimate phosphodiester bond position at the 3' end of each primer. This modification enhances the mismatch recognition ability between the primer and the template. When a single nucleotide mismatch occurs between the template base and the 3' end of the primer, the elongation efficiency decreases by more than 80%. The molecular internal stud structure introduces an additional pair of complementary short molecules within the primer sequence. Sequences were generated to form a stem-loop structure of 4-6 base pairs, which reduces the free energy for primer dimer formation, decreasing the dimer formation probability by approximately 60%. The designed multiplex primer mixture was prepared into working solutions, with each primer having a final concentration of 0.2-0.4 μM. These solutions were diluted with nuclease-free water, aliquoted, and stored at -20°C in the dark. Orthogonal gradient calibrators simulating algal communities were constructed. Representative species from the brown and red algae phyla with abundance gradients spanning at least three orders of magnitude were selected. Genomic DNA was extracted from each species and its concentration was determined. Mixed templates were prepared according to the orthogonal experimental design, with the initial copy number of low-abundance templates controlled within the range of 10-100 copies / μL and the initial copy number of high-abundance templates controlled within the range of 10-100 copies / μL. 4 -10 5PCR amplification was performed using multiple primers within the copy / μL range. Cyclic parameters were set as follows: 95℃ pre-denaturation for 3 minutes; 95℃ denaturation for 30 seconds, 58℃ annealing for 45 seconds, and 72℃ extension for 45 seconds, for a total of 35 cycles; a final extension at 72℃ for 5 minutes. After amplification, the ratio of expected to measured abundance in each calibrator was obtained by high-throughput sequencing. Bias response surface models were established for three parameters: primer dimer formation rate, GC bias coefficient, and template secondary structure free energy. Primer dimer formation rate was quantified by the proportion of nonspecific peak area in the melting curve. The GC bias coefficient was calculated based on a template with 50% GC content to determine relative amplification efficiency. The template secondary structure free energy was predicted using mFold software to determine its ΔG value. Fluorescence signal intensity data of the amplified products were collected at each cycle number to construct an observation sequence reflecting the dynamics of base incorporation during the amplification cycle. A hidden Markov model was set. The three latent states correspond to the initial stage of exponential amplification, the linear amplification stage, and the plateau stage, respectively. A state transition penalty factor is introduced to constrain the probability of state switching between adjacent cycles. The penalty factor is set to 0.7 to suppress misjudgment of states caused by random fluctuations. The base incorporation preference parameter in each state is estimated by the Baum-Welch algorithm to obtain the amplification efficiency correction coefficient of different templates in each cycle. The amplification efficiency of low abundance templates in the initial stage of exponential amplification (cycles 15-25) is reverse-weighted and corrected. The correction weight is negatively correlated with the logarithm of the initial copy number of the template. The lower the initial copy number, the greater the correction weight. The correction coefficients are integrated to form a bias correction parameter set. This parameter set contains the specific amplification correction factor for each species at each abundance level. It is used to correct the relative abundance of species in subsequent measured samples on a template-by-template basis to suppress the relative masking of closely related species signals caused by PCR amplification preference.
[0028] like Figure 3 As shown, the horizontal axis represents the initial copy number (copies / μL), and the vertical axis represents the species annotation confidence level. The scatter plot distribution is as follows: low abundance (10... 1 -10 2 ): Confidence level 0.70-0.85; Moderate abundance (10) 3 -10 4 ): Confidence level 0.85-0.95; High abundance (10) 5 ): Confidence level 0.92-0.98, of which, closely related rare species: confidence level 0.75-0.82, brown algae = blue, red algae = green, closely related rare species = red, and a confidence level ≥ 0.7 is the effective typing threshold; The expression for the amplification efficiency correction coefficient is as follows: ; In the formula: is the amplification efficiency correction factor; k represents the kth seaweed species template, corresponding to a specific species in the calibrator or a homologous sequence in its subsequent test samples; j represents the jth cycle in the PCR amplification process, where the critical cycle interval targeted by the correction is the 15th to 25th cycles (the initial stage of exponential amplification). This represents the copy number of the k-th template at the start of PCR amplification (initial copy number). This represents the maximum initial copy number among all templates in the same amplification reaction; α represents the base incorporation preference parameter estimated by the Hidden Markov Model in the j-th cycle. This parameter is learned from the amplification dynamics observation sequence using the Baum-Welch algorithm and reflects the systematic influence of the amplification stage (early exponential amplification, linear amplification, or plateau phase) on the template amplification efficiency; α represents the overall correction strength coefficient, which is a pre-set normal number; the initial copy number N k,0 The lower, The smaller the ratio, the better. The smaller the subtraction term in the formula, the smaller the correction factor C. k,j The closer the coefficient is to 1 or greater than 1, the higher the correction weight is given to the low abundance template. Conversely, the correction coefficient of the high abundance template is moderately reduced to suppress the relative signal masking caused by amplification preference.
[0029] like Figure 2 As shown, the horizontal axis represents the number of PCR cycles (15, 20, 25, 30, 35), and the vertical axis represents the relative amplification efficiency. The curve data are as follows: before high abundance correction, 0.95, 1.10, 1.20, 1.00, 0.80; after high abundance correction, 0.92, 0.98, 1.00, 0.95, 0.82; before low abundance correction, 0.30, 0.40, 0.50, 0.60, 0.55; after low abundance correction, 0.70, 0.85, 0.95, 0.92, 0.80. The black solid line represents the result after high abundance correction; the black dashed line represents the result before high abundance correction; the red solid line represents the result after low abundance correction; and the red dashed line represents the result before low abundance correction. The graph shows that the correction effect is most significant during the exponential amplification phase (cycles 15-25).
[0030] Step 4: Perform high-throughput sequencing on the amplified products, introduce convolutional neural networks to identify base variation sites, perform debiased typing on mixed amplicon to distinguish closely related rare species in seaweed farms, and correct species annotation results to establish a large-scale seaweed biodiversity database for seaweed farms, supporting accurate identification, deep learning to distinguish closely related species, and constructing a high-confidence dynamic database.
[0031] Step 5: Integrate the DNA extraction purity data and bias correction parameter set extracted in the previous step to construct a spatiotemporal coupled dynamic evolution model of seaweed community, realize full-chain data alignment, and use a temporal Bayesian network to hierarchically fuse environmental factors and species annotation results to dynamically analyze the differences in community structure between artificial and natural seaweed farms, eliminate comparative interference, eliminate systematic bias interference, and realize objective comparative analysis of community structure.
[0032] Step 6: Deploy a graph convolutional long short-term memory network to drive the spatiotemporal coupled dynamic evolution model of seaweed communities, outputting a spatiotemporal dynamic map of biodiversity of large seaweeds in seaweed farms, supporting the tracking and evaluation of ecological restoration processes, realizing precise monitoring of the entire chain of biodiversity of large seaweeds in seaweed farms, dynamically visualizing the restoration process, and accurately identifying the ecological succession stage of seaweed farms.
[0033] Example 2 like Figures 1 to 4 As shown, based on Example 1, the present invention provides a technical solution: Step 4 specifically includes: after quality filtering and chimera removal of high-throughput sequencing data, a multi-level base context coding strategy is used to convert each read into a three-dimensional feature tensor containing the coupling relationship of adjacent triplet sites, effectively preserving the evolutionary conservation and context-dependent structure of the base sequence, constructing a residual convolutional neural network architecture with an attention mechanism, using a haplotype reference set of closely related species verified by single-cell sequencing as the training object, learning the long-range dependence between variant sites and species-specific evolutionary fingerprints, significantly improving the classification boundary discrimination of closely related species in high-dimensional feature space, performing hierarchical progressive classification reasoning on mixed amplicon reads, first clustering by genus-level features and then finely typing by species-level variant sites, outputting the species label corresponding to each read and the confidence score after Bayesian calibration, greatly improving the identification accuracy and reliability of low-abundance closely related rare species; It should be noted that quality filtering was performed on paired-end reads from the high-throughput sequencing platform. A sliding window length of 5 bases was set, with an average window quality threshold of Q20. Reads with an average window quality below this threshold were removed, along with sequences shorter than 75% of the original read length. VSEARCH software was used with default parameters to identify and remove chimeric sequences, retaining high-quality, unique reads for subsequent analysis. After preprocessing, a multi-level base context encoding strategy was implemented for each read: first-level context captured the transfer preference of adjacent bases, second-level context extracted codon-level co-occurrence patterns of triplets, and third-level context... To model longer-range base association rules, multi-order features are stacked into a three-dimensional feature tensor, with the tensor dimension being read length × feature order × base embedding dimension. The coupling relationship between adjacent triplet sites is quantified using the log-likelihood ratio of the co-occurrence matrix. This maps discrete base sequences to a continuous vector space, preserving the contextual dependency structure implied by evolutionary conservation. A residual convolutional neural network architecture is constructed with a network depth of 34 layers, each with a 3×3 kernel size and an initial number of filters of 64. The number of filters is doubled every other residual block. A compressed-excitation attention module is embedded after each residual block. The feature map is compressed into channel descriptors using global average pooling. Then, channel weights are generated through two fully connected layers (the first layer has 1 / 16 the number of channels and uses ReLU activation; the second layer restores the original number of channels and uses Sigmoid activation). The original feature map is adaptively recalibrated. The network is trained using a reference set of haplotypes of closely related species verified by single-cell sequencing. This reference set covers haplotype sequences of 38 brown algae species and 42 red algae species recorded in the target seaweed farm, with at least 3 independent haplotypes for each species. Cross-entropy loss is used during training, and the optimizer is Ada. With m as the initial learning rate and batch size of 32, the network learns the long-range dependencies between variant sites through an attention mechanism, extracts species-specific evolutionary fingerprint features, and forms classification boundaries that distinguish closely related species in a high-dimensional feature space. A hierarchical progressive strategy is adopted for classification reasoning of mixed amplicon reads. The first level is genus-level feature clustering: the three-dimensional feature tensor of the read to be classified is input into the trained residual convolutional network, and the feature vector (dimension 512) before the last fully connected layer is extracted. Cosine similarity is used to measure similarity with representative feature vectors of each genus in the reference set, with a similarity threshold set to 0.Reads with a confidence level below 85 are marked as unclassified. The second level is fine-tuning of species-level variant sites: Based on the genus-level clustering results, for each species within a genus, the 32 variant sites with the highest attention weights are extracted as species-level distinguishing features and input into the classification layer of the network for fine-tuning. For taxa with more than 5 species within a genus, a one-to-many strategy is used for individual discrimination. Each read ultimately outputs a species label and a Bayesian-calibrated confidence score. The confidence score is calculated by multiplying the posterior probability by the sequence diversity factor of the species in the reference set. This factor is defined as the ratio of the average pairwise sequence distance within a species to the minimum sequence distance between species. Reads with a confidence score below 0.7 are marked as questionable and are not included in subsequent abundance calculations.
[0034] Furthermore, step 4 also includes: based on the species label confidence score output by the convolutional neural network, using a graph cut algorithm to globally optimize the allocation boundary of low-confidence reads, modularly segmenting co-amplified closely related rare species by constructing a sequence similarity map, significantly improving the resolution of closely related rare species and reducing the classification error rate, comparing the biased typing results with the dynamically updated Hidden Markov Alignment Map, introducing an arbitration mechanism based on phylogenetic inconsistency scoring to correct species annotations for alignment conflict areas, effectively resolving alignment conflicts, improving the consistency and accuracy of species annotations, integrating the corrected species annotation results, and using an incremental Bayesian update strategy to construct a biodiversity database of large algae covering all sample plots of the target seaweed farm, supporting real-time assimilation and anomaly detection of new sequencing data, realizing real-time database updates and dynamic maintenance, and enhancing the early warning capability of abnormal changes.
[0035] It should be noted that for each read output by the convolutional neural network, its species label and Bayesian-calibrated confidence score are extracted. A threshold of 0.7 is set as the doubt threshold, and all reads with a confidence score below 0.7 are marked as low-confidence reads and participate in the graph cut optimization process. Using low-confidence reads as nodes, edge weights are constructed using the inverse of the edit distance between sequences to form an undirected weighted sequence similarity graph. The edit distance is calculated over the entire read length, and the truncation threshold is set to 0.15. Reads exceeding this threshold are not connected. The α-expansion graph cut algorithm is used to solve the multi-label segmentation problem of this graph. The data term in the energy function uses the log probability of the original confidence score output by the convolutional neural network, the smoothing term uses the Potts model, and the penalty coefficient is set to 0.5. After the algorithm converges, reads that were originally in the ambiguous region of the classification boundary are forcibly assigned to the most closely connected neighborhood category in the similarity graph, realizing modular segmentation of co-amplified closely related rare species. For reads that still cannot be assigned or whose confidence score is still below 0 after segmentation, the algorithm continues to optimize the graph.Reads 7 and 8 were marked as unclassified and removed from downstream analysis. The graph-cut optimized, debiased genotyping results were compared read-by-read with a dynamically updated Hidden Markov Model (HMM) alignment map. This map was constructed based on the full-length mitochondrial COI and 18S ribosomal reference sequences of 38 brown algae species and 42 red algae species from the target seaweed field, using a ProfileHMM architecture. The insertion and deletion state transition probabilities at each site were obtained through Byesian adaptive estimation. A map recalibration was triggered every 100 new high-quality sequencing reads. When the classification label of a read did not match the best-matching species in the HMM alignment map, it was identified as an alignment conflict region. For conflict regions, an arbitration mechanism based on phylogenetic inconsistency scoring was introduced: based on the constructed phylogenetic trees of brown and red algae in the target region (based on maximum likelihood, with 1000 bootstrap repetitions), the phylogenetic distance between the two candidate species to which the read belonged was calculated. The reciprocal of this distance was used as the inconsistency score weight, and the arbitration label was the species with the highest weighted score. After arbitration correction, the original convolutional neural network output labels are replaced, and the corrected annotation results enter the database integration stage. An incremental Bayesian update strategy is used to gradually assimilate the corrected species annotation results into the macroalgae biodiversity database. The database uses sampling point, sampling time, and species as three-level indexes. The existence probability of each species in each sample plot in each quarter is modeled using a Beta distribution, with an initial prior parameter of 1. For each new batch of sequencing data (defined as the read set of a single sequencing run), the arbitration-corrected species label for each read is read, and the occurrence frequency of each species is calculated per read. This frequency is used as the input to the likelihood function to update the posterior parameter of the corresponding Beta distribution, achieving real-time database assimilation. The latency for a single batch of data entering the database does not exceed 30 seconds. Simultaneously, the database has a built-in anomaly detection module: if the posterior existence probability of any species decreases by more than 0.3 compared to the previous time window and the current abundance is less than 20% of the historical median of that sample plot, an anomaly marker is triggered, indicating that the species may be declining or there may be sampling bias in that sample plot.
[0036] Step 5 specifically includes: performing joint standardization processing based on a Bayesian hierarchical framework on the obtained DNA purity and integrity data, bias correction parameter set, and species abundance table after debiasing and typing by convolutional neural network, eliminating differences in dimensions and error distribution between multiple steps, effectively unifying the scale of multi-source data, and improving the accuracy of subsequent fusion modeling; constructing a multi-layer heterogeneous data association matrix with the three-dimensional spatial coordinates (longitude, latitude, and water depth) of sampling points and multi-time point sampling windows as coupling nodes, realizing the backpropagation alignment of errors from DNA extraction purity constraints to species identification confidence, realizing error tracing and correction layer by layer, enhancing the consistency of the entire chain of data; using a spatiotemporal non-stationary Gaussian process regression method to replace traditional Kriging interpolation, performing uncertainty quantification of community composition inference for missing sites caused by humic acid interference that lead to DNA extraction failure, generating a spatiotemporally continuous seaweed species distribution probability field, quantifying the uncertainty of missing site inference, and ensuring the spatiotemporal continuity of the community distribution map.
[0037] It should be noted that the obtained DNA purity and integrity data (characterized by the A260 / A230 ratio and the DNA integrity index H), the generated bias correction parameter set (containing specific amplification correction factors for each species at each abundance level), and the species abundance table (output in read count form) after graph cut optimization and phylogenetic arbitration are imported into the Bayesian hierarchical framework. The hyperparameters of the data layer are set as follows: the variance of DNA purity observations follows a Gamma distribution, the shape parameter is set to 2.0, and the scale parameter is set to 0.5; the species abundance count follows a negative binomial distribution. The discrete parameters of the model are pre-calculated from the reference set using the method of moments, with a value of 0.15. Weak prior information is introduced at the parameter level, with the prior mean of the dimension transformation factor being 1.0 and the standard deviation being 0.2. A No-U-Turn sampler is used for posterior sampling, executing four Markov chains, each iterating 2000 times. The first 500 iterations are discarded as warm-up iterations. The convergence criterion is set to a latent scale reduction factor less than 1.05. After standardization, all data sources are mapped to a unified logarithmic scale space, and the dimensional differences are compressed to within 5% of the original amplitude. The error distribution tends to have a symmetrical unimodal shape. Based on standardized data, the three-dimensional spatial coordinates of each sampling point (longitude uses the WGS84 coordinate system; water depth is in meters) and multi-time point sampling windows (expressed in days per year, with a sampling interval of no less than 30 days) form coupling nodes. When constructing a multi-layer heterogeneous data association matrix, the DNA extraction purity constraint is used as the first-layer variable, the deviation correction parameter as the second-layer variable, and the species identification confidence level as the third-layer output variable. Error backpropagation uses a chain rule to calculate partial derivatives layer by layer, and the damping factor is set to 0 in the propagation path. To prevent gradient explosion, in practice, when the A260 / A230 ratio of a sample is lower than 2.0 in the quality control judgment, the abundance of all species associated with that sample is assigned a decay weight of 0.6 in backpropagation; when the DNA integrity index H is lower than 0.75, the decay weight is further reduced to 0.4; for missing sampling points (approximately no more than 8% of the total sampling plan) due to DNA extraction failure or quality control failure caused by humic acid interference, spatiotemporal non-stationary Gaussian process regression is used to infer community composition, and the covariance function is selected as Maternity. The 3 / 2 core was used, with the spatial length scale parameter set to 1500 meters (matching the spacing of seaweed beds) and the temporal length scale parameter set to 60 days (matching the quarterly sampling interval). Non-stationarity was achieved by applying a depth Gaussian process prior to the length scale parameter. The mean of this prior varied linearly with the water depth gradient, with a coefficient of 0.05 per meter. Hyperparameter optimization adopted the marginal likelihood maximization method, which was performed for 300 iterations. The convergence tolerance was set to 10 to the power of -3. The inference output was presented in the form of a probability field: the mean probability of existence of each species corresponding to each missing site and the 95% confidence interval. The spatial grid resolution was 100 meters × 100 meters, and the temporal interpolation step size was 30 days. For posterior confidence intervals with a width exceeding 0.The inferred site at number 4 was automatically labeled as "high uncertainty" and was not included in subsequent trend analysis. The final output probability field of algal species distribution covers the entire study area, is spatially continuous and temporally smooth.
[0038] Furthermore, step 5 also includes: collecting high-frequency time-series environmental factor data from each sampling point, including water temperature, salinity, dissolved oxygen, turbidity, sediment particle size distribution, and nutrient molar ratio; performing non-equidistant time series interpolation and normalization processing using adaptive Kalman filtering to ensure the time series continuity and comparability of environmental factors, improving the stability of model input; constructing a four-layer time-series Bayesian network model with hidden state layers, using environmental factors as observation parent nodes, DNA extraction purity and bias correction parameters as hidden state nodes, and species annotation results as output child nodes; learning the conditional dependence structure and causal path differences between artificial and natural seaweed farms; revealing the key environmental driving factors and causal paths of different seaweed farm types; and inferring the system offset in community structure of artificial seaweed farms under construction, mature artificial seaweed farms, and natural seaweed farms through posterior probability comparison; eliminating ecological comparison interference caused by humic acid interference, amplification bias, and loss of rare species; realizing quantitative analysis of community structure differences; and eliminating the influence of multi-source interference on comparison.
[0039] It should be noted that high-frequency time-series environmental factor data were simultaneously collected at various sampling points in the target seaweed farm. These included water temperature, salinity, dissolved oxygen, turbidity, sediment particle size distribution (measured by a laser particle size analyzer, classifying sand, silt, and clay into three grades), and nutrient molar ratio (the ratio of dissolved inorganic nitrogen to reactive phosphate). Due to the non-uniformity of the field sampling intervals, an adaptive Kalman filter was used to perform non-uniform interval interpolation and normalization on each environmental factor sequence. Specifically, the state transition matrix was set as a diagonal matrix, the initial value of the process noise covariance was set to 0.01, and the observation noise covariance was set separately based on the calibration error of each sensor. The filtering convergence criterion was that the zero-mean test of the innovation sequence passed. Normalized environmental factor data were mapped to a standard space with a mean of 0 and a standard deviation of 1 to eliminate the influence of dimensional differences on subsequent modeling. Preprocessed environmental factors were used as observation parent nodes, DNA extraction purity (A260 / A230 ratio) and factors of each dimension of the bias correction parameter set were used as first hidden state nodes, the statistics of the PCR amplification efficiency correction coefficient (median and interquartile range) were used as second hidden state nodes, and species annotation results (read counts after graph cut and arbitration) were used as output child nodes. A four-layer temporal Bayesian network model was constructed. The K2 algorithm was used for network structure learning, the search space was limited to no more than 3 parent nodes, and maximum likelihood estimation was used for conditional probability table parameter estimation. A Dirichlet prior was applied, with hyperparameters set as the reciprocals of the number of states of each parent node. For artificial seaweed farms (mature and under construction) and natural seaweed farms, network parameters were trained independently, resulting in three sets of conditional probability distributions. By calculating the posterior probability ratio of the two networks under the same observational evidence, environmental factor nodes and causal paths with discriminative power for different seaweed farm types were identified, quantifying the systematic impact of humic acid interference and amplification bias on species annotation. Based on the aforementioned Bayesian network, posterior probability comparison inference was calculated for each sampling time window. Specifically, the observed parent nodes (i.e., measured environmental factors) were fixed, and the network models for artificial and natural seaweed farms were substituted into them, outputting the respective... The expected probability of species existence is defined as the system offset, which is the ratio of the logarithmic posterior probabilities of the two network outputs. A positive value indicates that the relative abundance of the species in the artificial seaweed field is higher than that in the natural control area, while a negative value indicates the opposite. To further eliminate the comparison interference caused by DNA extraction failure or quality control failure, a decay weight is introduced in the offset calculation: when the A260 / A230 ratio of a sample is lower than 2.0, the posterior probability contribution of the sample is multiplied by 0.6; when the DNA integrity index is lower than 0.75, it is further multiplied by 0.4. The species-specific offset error bar of the final output is estimated using the Bootstrap method (1000 resamplings). A 95% confidence interval crossing zero is considered as having no significant offset.
[0040] like Figure 4As shown, the horizontal axis represents the A260 / A230 ratio, and the vertical axis represents the number of species detected. The scatter plot data are as follows: 1.50→12 species, 1.80→18 species, 2.00→24 species, 2.20→29 species, 2.30→31 species. The smoothed line is a LOESS smoothed fit, and R... 2 =0.96. As shown in the figure, the number of species detected enters a plateau period when A260 / A230≥2.0.
[0041] Step 6 specifically includes: using each sampling point as a graph node, a dynamic adjacency matrix is generated by a composite spatial distance kernel constructed based on hydrological connectivity and substrate type similarity, forming an adaptive graph structure that can update topological relationships with the seasons, effectively capturing the spatial dependence of seaweed farms, improving the accuracy of community structure modeling, using the posterior species probability field and uncertainty interval output by the temporal Bayesian network as the input sequence, extracting spatial dependence features through a graph attention convolutional layer with residual connections, and then capturing the evolutionary patterns of forward and backward time windows through a bidirectional long short-term memory network, fully integrating spatiotemporal features, significantly improving the prediction accuracy of species dynamic changes, jointly training the parameters of the adversarial graph convolutional network and the long short-term memory network, incorporating DNA extraction quality labels as regularization constraints, and outputting a spatiotemporal dynamic map of biodiversity of large seaweed farms containing species composition, migration paths of dominant species, and appearance / disappearance events of rare species, which is used for stage discrimination of the ecological restoration process of artificial seaweed farms, enhancing the robustness of the model, and realizing quantifiable discrimination of ecological restoration stages.
[0042] It should be noted that, using each sampling point in the target seaweed farm as a graph node, a composite spatial distance kernel is constructed based on hydrological connectivity and sediment type similarity to generate a dynamic adjacency matrix. Hydrological connectivity is calculated using the main tidal current direction and velocity amplitude measured by an acoustic Doppler current profiler, with the water mass exchange rate between two nodes used as a weighting factor. Sediment type similarity is calculated using the Bray-Curtis dissimilarity index based on the volume percentages of sand, silt, and clay measured by a laser particle size analyzer. The composite spatial distance kernel weights and merges these two weighting factors in a 7:3 ratio to form an initial adjacency matrix. Based on this, with a quarterly time step, hydrological connectivity is reassessed according to newly collected environmental factors (water temperature, salinity). The adjacency matrix is updated with corresponding edge weights to achieve an adaptive graph structure that changes with the seasons. The spatial adjacency range is set to a distance threshold of 2000 meters; nodes exceeding this threshold are not directly connected. The posterior species probability field and 95% confidence interval output by the temporal Bayesian network are used as input sequences and fed into the graph attention convolutional long short-term memory network model. The number of attention heads in the graph attention convolutional layer is set to 4, and the output feature dimension of each head is 64. Residual connections are added after every two convolutional layers to alleviate the gradient vanishing problem in deep networks. The time window length for each sampling point is set to 12 months, and the sliding step is 30 days. The hidden layer dimension of the bidirectional long short-term memory network is set to 128, and the forward and backward outputs are... The network was fused using a concatenation method. The Adam optimizer was used during training with an initial learning rate of 0.0005, a batch size of 16, and 200 training epochs. An early stopping mechanism was triggered when the validation loss did not decrease for 10 consecutive epochs. The model output included the expected species abundance, dominant species ranking trend, and rare species detection probability time series for each sampling node at each time point. An adversarial training strategy was used to jointly optimize the parameters of the graph convolutional and long short-term memory networks. DNA extraction quality labels were introduced as regularization constraints. In adversarial training, the learning rates of both the generator and discriminator were set to 0.0002, and each was updated once per epoch during alternating training. The DNA extraction quality labels were determined using the A260 / A230 ratio and D... The product of the NA integrity indexes is used as the comprehensive score, and the regularization coefficient is set to 0.1. This score is added to the total loss function as an auxiliary loss term to constrain the model to reduce the prediction weight on samples with low DNA extraction quality. After the model is trained, a spatiotemporal dynamic map of biodiversity of large seaweeds in the seaweed farm is output. The migration path of the dominant species is identified using the Bayesian change point detection algorithm. The appearance or disappearance of rare species is determined based on the detection probability crossing the 0.5 threshold within three consecutive time windows. This spatiotemporal dynamic map is output in the form of a spatial grid resolution of 100m×100m and a temporal resolution of 30 days, which is used to distinguish the stages of the ecological restoration process of artificial seaweed farms (divided into three stages: initial colonization, structural development, and community stability).
[0043] The above embodiments are merely exemplary embodiments of the present invention and are not intended to limit the present invention. Those skilled in the art can make various modifications or equivalent substitutions to the present invention within its scope and spirit, and such modifications or equivalent substitutions should also be considered to fall within the scope of protection of the present invention.
Claims
1. A method for monitoring the biodiversity of macroalgae in seaweed farms based on environmental DNA technology, characterized in that, Includes the following steps: Step 1: Collect seaweed field sediments and water samples in the target area, and use density gradient centrifugation to remove macromolecular polysaccharides while retaining low-abundance seaweed DNA components. Step 2: Based on the improved magnetic bead adsorption method, modified magnetic beads are used to target and remove humic acid, thereby obtaining DNA extraction purity data and integrity information. Step 3: For the data obtained in Step 2, design seaweed-specific amplicon primers, combine them with a hidden Markov model to correct PCR amplification bias, obtain a bias correction parameter set, and suppress the masking of signals from closely related species. Step 4: Perform high-throughput sequencing on the amplified products, identify base variation sites using convolutional neural networks, perform debiased typing on the mixed amplicon to distinguish closely related rare species in the seaweed farm, correct the species annotation results, and establish a large seaweed biodiversity database for the seaweed farm. Step 5: Integrate the DNA extraction purity data obtained in Step 2 with the deviation correction parameter set obtained in Step 3 to construct a spatiotemporal coupled dynamic evolution model of seaweed communities. Then, use a temporal Bayesian network to hierarchically fuse environmental factors and species annotation results to dynamically analyze the differences in community structure between artificial and natural seaweed farms. Step 6: Deploy a graph convolutional long short-term memory network to drive the spatiotemporal coupled dynamic evolution model of seaweed communities, outputting a spatiotemporal dynamic map of biodiversity of large seaweeds in seaweed farms, and realizing accurate monitoring of the entire seaweed farm chain.
2. The method for monitoring the biodiversity of large seaweeds in seaweed farms based on environmental DNA technology according to claim 1, characterized in that, Step 1 specifically includes: Surface sediments and bottom water samples were collected from the target seaweed farm. After removing large particulate impurities by low-temperature centrifugation, the supernatant was transferred to a continuous density gradient medium constructed by iodixanol and sucrose in a volume ratio of 1:
2. The medium was then subjected to ultracentrifugation at 4°C. Utilizing the significant difference in sedimentation coefficients between macromolecular polysaccharides and DNA components in the composite medium, the polysaccharides migrated to the upper layer of the gradient, while low-abundance DNA was enriched in a high-density characteristic layer with a density of 1.25-1.35 g / mL. The high-density characteristic layer components were extracted using a puncture-layer collection device, and the resulting crude extract of low-abundance seaweed DNA was obtained after ultrafiltration desalting and a nucleic acid concentration column.
3. The method for monitoring the biodiversity of large seaweeds in seaweed farms based on environmental DNA technology according to claim 1, characterized in that, Step 2 specifically includes: Bifunctional superparamagnetic nanoparticles with surface simultaneously modified with carboxyl and polyethyleneimine were mixed with low-abundance crude DNA extract at a mass-volume ratio of 1:
5. The pH of the reaction system was adjusted to 4.0-5.0 to enhance the electrostatic adsorption between the phenolic hydroxyl groups in humic acid molecules and the surface of the magnetic beads. Incubate at 4°C under low-temperature oscillation conditions for a preset time, and introduce competitive binding inhibitors to block the non-specific adsorption of magnetic beads on target DNA, enabling the modified magnetic beads to selectively capture humic acid. An alternating magnetic field was applied to cause the magnetic beads to migrate in a directional manner and form a chain-like aggregate structure. After collecting the supernatant, the sample was dynamically eluted with a low-salt buffer to obtain a preliminarily purified seaweed DNA sample.
4. The method for monitoring the biodiversity of large seaweeds in seaweed farms based on environmental DNA technology according to claim 3, characterized in that, The fragmentation level of pre-purified DNA samples was analyzed at the single-molecule level using a combination of nucleic acid fluorescent dyes and nanoflow cytometry. The absorbance at 230 nm, 260 nm and 280 nm was measured, and the A260 / A230 ratio was calculated to quantify the residual humic acid level. The retention efficiency of high molecular weight DNA fragments was evaluated by pulsed field gel electrophoresis combined with fluorescence signal integration algorithm, and the fragment length distribution entropy value was introduced as a DNA integrity index. Construct a two-dimensional decision boundary for purity and integrity to automatically remove DNA samples that do not meet the threshold.
5. The method for monitoring the biodiversity of large seaweeds in seaweed farms based on environmental DNA technology according to claim 1, characterized in that, Step 3 specifically includes: Multiple specific amplicon primer pairs containing molecular interior studs were designed for brown algae and red algae to target the V4 region of the 18S rRNA gene and the 5' hypervariable region of the COI gene. Mismatch recognition ability was improved by locking nucleic acid at the 3' end of the primers. PCR amplification was performed by constructing orthogonal gradient calibrators to simulate seaweed communities, and a three-parameter bias response surface covering primer dimer formation, GC bias, and template secondary structure was established. Based on the Hidden Markov Model, a state transition penalty factor is introduced to learn the nonlinear preference pattern of base incorporation in the amplification cycle. Inverse weighted correction is performed on the early stage of exponential amplification of low-abundance templates to generate a bias correction parameter set to suppress the relative masking of closely related species signals.
6. The method for monitoring the biodiversity of large seaweeds in seaweed farms based on environmental DNA technology according to claim 1, characterized in that, Step 4 specifically includes: After quality filtering and chimera removal of the high-throughput sequencing data, a multi-level base context coding strategy is used to convert each read into a three-dimensional feature tensor containing the coupling relationship between adjacent triplet sites. We constructed a residual convolutional neural network architecture that incorporates an attention mechanism, and used a haplotype reference set of closely related species verified by single-cell sequencing as the training object to learn the long-range dependencies between variant sites and species-specific evolutionary fingerprints. Hierarchical progressive classification reasoning is performed on mixed amplicon reads. First, clustering is performed by genus-level features, and then fine typing is performed by species-level variation sites. The species labels and Bayesian-calibrated confidence scores for each read are output.
7. The method for monitoring the biodiversity of large seaweeds in seaweed farms based on environmental DNA technology according to claim 6, characterized in that: Step 4 also includes: Based on the species label confidence score output by the convolutional neural network, the graph cut algorithm is used to globally optimize the allocation boundary of low confidence reads, and modular segmentation of co-amplified closely related rare species is performed by constructing a sequence similarity graph. The debiased typing results were compared with the dynamically updated Hidden Markov alignment map, and an arbitration mechanism based on phylogenetic inconsistency score was introduced to correct species annotations for areas of alignment conflict. By integrating the revised species annotation results, an incremental Bayesian update strategy was used to construct a macroalgal biodiversity database covering all sample plots in the target seaweed farm.
8. The method for monitoring the biodiversity of large seaweeds in seaweed farms based on environmental DNA technology according to claim 1, characterized in that: Step 5 specifically includes: The obtained DNA purity and integrity data, bias correction parameter set, and species abundance table after debiasing and typing by convolutional neural network are jointly standardized based on Bayesian hierarchical framework to eliminate the differences in dimensionality and error distribution between multiple steps. A multi-layer heterogeneous data association matrix is constructed using the three-dimensional spatial coordinates of the sampling points and the multi-time point sampling window as coupling nodes; A spatiotemporally non-stationary Gaussian process regression method is used to replace traditional Kriging interpolation. Uncertainty quantification is performed on the community composition inference of missing sites where DNA extraction fails due to humic acid interference, generating a spatiotemporally continuous probability field of algal species distribution.
9. The method for monitoring the biodiversity of large seaweeds in seaweed farms based on environmental DNA technology according to claim 8, characterized in that: Step 5 further includes: High-frequency time-series environmental factor data from each sampling point were collected, including water temperature, salinity, dissolved oxygen, turbidity, bottom sediment particle size distribution, and nutrient molar ratio. The data were then processed by adaptive Kalman filtering for non-equidistant time series interpolation and normalization. A four-layer temporal Bayesian network model with hidden state layers is constructed, with environmental factors as observation parent nodes, DNA extraction purity and bias correction parameters as hidden state nodes, and species annotation results as output child nodes, to learn the differences in conditional dependence structure and causal path between artificial and natural seaweed farms. By comparing posterior probabilities, the system offset in community structure between artificial seaweed farms under construction, mature artificial seaweed farms, and natural seaweed farms is inferred through dynamic analysis, eliminating ecological comparison interference caused by humic acid interference, amplification bias, and loss of rare species.
10. The method for monitoring the biodiversity of large seaweeds in seaweed farms based on environmental DNA technology according to claim 1, characterized in that: Step 6 specifically includes: Using each sampling point as a graph node, a dynamic adjacency matrix is generated by a composite spatial distance kernel constructed based on hydrological connectivity and similarity of substrate type, forming an adaptive graph structure that can update topological relationships with the seasons; The posterior species probability field and uncertainty interval output by the temporal Bayesian network are used as the input sequence. Spatial dependency features are extracted through graph attention convolutional layers with residual connections, and then the evolutionary patterns of forward and backward time windows are captured through a bidirectional long short-term memory network. The parameters of the joint adversarial training graph convolutional network and long short-term memory network are incorporated with DNA-extracted quality labels as regularization constraints to output a spatiotemporal dynamic map of biodiversity in large seaweed farms, including species composition, migration paths of dominant species, and occurrence / disappearance events of rare species.