Microplastic stress environment microbial community key succession species identification method

By constructing a relative abundance matrix and classification model, and combining phylogenetic tree data, the assembly offset index and correlation coefficient were calculated to identify key successional species under microplastic stress. This solved the problems of abundance bias and high false positive rate in existing technologies, and achieved microbial community analysis with high accuracy and ecological mechanism support.

CN122493982APending Publication Date: 2026-07-31NORTHWEST A & F UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-12
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies suffer from abundance bias in identifying key successional species in environmental microbial communities under microplastic stress, leading to the neglect of low-abundance taxa and a lack of ecological mechanism support, resulting in a high false positive rate.

Method used

By constructing a relative abundance matrix and a classification model, and combining phylogenetic tree data, the assembly offset index and correlation coefficient were calculated to screen out candidate key successional species with low abundance and high contribution, and their response type was determined by time response curves.

Benefits of technology

It effectively identified a small number of microbial groups that change significantly under stress conditions, reduced the false positive rate, ensured that the screening results were supported by ecological mechanisms, and improved the accuracy and reliability of the analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493982A_ABST
    Figure CN122493982A_ABST
Patent Text Reader

Abstract

This invention discloses a method for identifying key successional species in microbial communities under microplastic stress, belonging to the fields of environmental microbiology and pollution ecological risk assessment. The method includes: S1, acquiring high-throughput sequencing data of samples, constructing a relative abundance matrix, and grouping them; S2, constructing a classification model using microbial relative abundance as input, extracting importance scores, and obtaining a set of candidate key successional species through screening; S3, calculating assembly shift indices based on phylogenetic distance matrices to quantify the relative contributions of deterministic and stochastic processes under each grouping condition; S4, calculating comprehensive change characteristic values ​​and assembly shift characteristic values, and then calculating the assembly shift correlation coefficient between the two, obtaining key successional species through quantitative correlation screening. This invention overcomes the shortcomings of traditional methods that neglect low-abundance functional groups and lack ecological mechanism support, significantly reducing the false positive rate and improving the reliability and technical specificity of key successional species identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of environmental microbiology, environmental bioinformatics and pollution ecological risk assessment, and specifically relates to a method for identifying key successional microorganisms in environmental samples under microplastic stress conditions. Background Technology

[0002] With the continuous accumulation of microplastics in environmental media such as soil, water, and sediments, their impact on environmental micro-ecosystems is receiving increasing attention. Microplastics can not only affect microbial habitats by altering the physicochemical properties of the media, but may also drive changes in the structure, function, and succession of environmental microbial communities through pathways such as particle surface adsorption, biofilm formation, additive release, and interfacial effects.

[0003] Identifying successional microorganisms that play a key role in responding to microplastic stress is crucial for environmental risk assessment and elucidating response mechanisms. However, existing technologies face several bottlenecks in identifying core biomarkers of the environmental microbiome: First, there is significant abundance bias. Current differential abundance analysis methods often over-rely on absolute species abundance, leading to the masking or neglect of many ecologically important but low-abundance taxa. Second, there is a lack of ecological mechanistic support, resulting in a high false positive rate. When simply using data-driven machine learning algorithms to extract feature importance, the high dimensionality and strong background noise of high-throughput environmental data can cause models to misinterpret abundance fluctuations caused by random community drift as genuine biological responses to microplastic stress, leading to a lack of clear ecological causal relationships among the screened microorganisms. Summary of the Invention

[0004] To overcome the shortcomings of existing technologies, such as the easy masking or neglect of low-abundance taxa and the lack of ecological mechanism support for screening results, resulting in a high false positive rate, this application provides a method for identifying key successional species in microbial communities under microplastic stress.

[0005] The technical solution of this application is as follows: The method for identifying key successional species in microbial communities under microplastic stress provided in this application includes the following steps: S1. Extract microbial DNA from environmental samples affected by microplastic stress and perform high-throughput sequencing. Generate abundance tables and phylogenetic tree data based on the high-throughput sequencing results and bioinformatics workflow, and obtain a sample metadata table recording microplastic type, exposure concentration, exposure time, and environmental site information. Standardize the abundance matrix in the abundance table into a relative abundance matrix. Group the environmental samples according to the microplastic type, exposure concentration, exposure time, and environmental site information in the sample metadata table. S2. Using the relative abundance of microbial groups as the input feature and the sample group as the target variable, construct a classification model and extract the importance score of each microbial group. Based on the relative abundance matrix, the average relative abundance of each microbial group was calculated; based on the average relative abundance and importance score, the microbial groups were stratified and screened, and low-abundance high-contribution groups and high-abundance high-contribution groups were extracted as candidate key succession species. S3. Construct a phylogenetic distance matrix between microbial groups based on the phylogenetic tree data. The phylogenetic distance matrix records the closeness of the evolutionary relationship between each pair of microbial groups. Based on the phylogenetic distance matrix and relative abundance matrix, the assembly offset index between different samples is calculated; the community assembly process type of each sample pair is determined based on the assembly offset index, and the determination results are statistically analyzed according to the grouping conditions; the relative contributions of deterministic assembly processes and stochastic processes under each grouping condition are obtained based on the statistical results. S4. Based on the relative abundance changes and importance score changes of each candidate key successional species under different grouping conditions, calculate the comprehensive change characteristic value of each candidate key successional species under different grouping conditions. Based on the relative contributions of the deterministic assembly process and the relative contributions of the stochastic process under each grouping condition, calculate the assembly offset characteristic value under the corresponding grouping condition. The assembly offset correlation coefficient of each candidate key successional species is calculated based on the comprehensive change feature value and the assembly offset feature value. The candidate key successional species are then screened a second time based on the assembly offset correlation coefficient to obtain a set of key successional species.

[0006] In one possible design, the identification method further includes step S5: S5. Construct a time response curve based on at least one of the abundance changes, importance score changes, or assembly offset correlation value changes of the key successional species at different exposure time points, and determine the response type of the key successional species based on the time response curve.

[0007] In one possible design, step S5 further includes: using Dynamic Time Warping (DTW) to calculate the time response distance between key successional species, and performing temporal clustering based on the time response distance.

[0008] In one possible design, in step S4, the comprehensive variation characteristic values ​​of each candidate key successional species under different grouping conditions are... Calculate using the following formula:

[0009] in, Indicates the first The candidate key successor species in the The comprehensive change characteristic value under each grouping condition; The candidate key succession species is indicated in the first... Relative abundance changes under each grouping condition; The candidate key succession species is indicated in the first... Changes in importance scores under each grouping condition; Represents a standardized function; and These are the weighting coefficients.

[0010] In one possible design, in step S4, the assembly offset characteristic value corresponds to the grouping condition. Calculate using the following formula:

[0011] in, Indicates the first Assembly offset characteristic values ​​under each grouping condition; Indicates the first The relative contribution of the deterministic assembly process under each grouping condition; Indicates the first The relative contribution of a stochastic process under each grouping condition; This represents the standardized function.

[0012] In one possible design, in step S4, the assembly offset correlation coefficients of each candidate key successional species are... Calculate using the following formula:

[0013] in, Indicates the first Assembly offset correlation coefficients of candidate key successional species; Indicates the total number of grouping conditions; Indicates the first The average value of the comprehensive change characteristics of the candidate key successional species; This represents the average value of the assembly offset characteristic value.

[0014] In one possible design, step S4 further includes: verifying the significance of the assembly offset correlation coefficient through a permutation test; when the assembly offset correlation coefficient is not lower than a preset threshold and the significance test result meets the preset significance level, the corresponding candidate key succession species is retained as a key succession species.

[0015] In one possible design, step S1) further includes obtaining a species classification annotation table based on the high-throughput sequencing results, and integrating the abundance table, species classification annotation table, sample metadata table, and phylogenetic tree data into a structurally unified and interconnected microbial community analysis data object.

[0016] In one possible design, step S1) also includes a step of quality control and filtering of the microbial community analysis data objects.

[0017] In one possible design, step S3) includes, before calculating the assembly offset index between different samples, a step of matching the phylogenetic distance matrix with the relative abundance matrix for common species and synchronizing the rearrangement.

[0018] The advantages of this application compared to the prior art are: First, in step S2, this application cross-combines the average relative abundance with the importance score output by the classification model to divide the species into four categories: low abundance with high contribution, high abundance with high contribution, high abundance with low contribution, and low abundance with low contribution. The first two categories are then selected as candidate key succession species. This stratified screening mechanism includes low-abundance taxa that are scarce but show significant changes under microplastic stress and are crucial for distinguishing different sample groups in the candidate list. This overcomes the shortcomings of traditional differential abundance analysis methods, which rely too heavily on absolute species abundance, leading to the masking or neglect of low-abundance functional taxa.

[0019] Secondly, this application calculates the community assembly shift index in step S3, quantifying the relative contributions of deterministic and stochastic processes. Step S4 calculates the comprehensive change characteristic value integrating abundance and importance score changes, as well as the assembly shift characteristic value reflecting the net intensity of community assembly deviating from the stochastic state. Based on this, the assembly shift correlation coefficient is calculated to measure the synchronicity between the trend of species response intensity changes and the trend of community assembly shift intensity changes. This quantitative correlation elevates the selection of candidate species from statistical correlation to mechanistic correlation, ensuring that the key successional species ultimately retained not only possess characteristic importance but are also highly coupled with the transformation of the overall community assembly mechanism. This effectively eliminates candidate taxa that generate falsely high contributions due to abundance fluctuations caused solely by random drift, significantly reducing the false positive rate.

[0020] Furthermore, this application processes the raw sequencing data into a standardized relative abundance matrix in step S1, eliminating systematic bias caused by differences in sequencing depth. Subsequently, through stratified screening in step S2, the analytical targets are compressed from tens of thousands of microbial taxa in the original matrix into a finite-sized candidate set. Each species in this candidate set undergoes explicit statistical screening, which not only significantly reduces the computational load of subsequent assembly offset association analysis but also effectively eliminates interference from a large number of background noise species on the statistical signal, providing high-quality focused input for the refined ecological validation in steps S3 and S4. Attached Figure Description

[0021] Figure 1 This is a flowchart of a preferred embodiment of the present invention.

[0022] Figure 2 This is a screening diagram of candidate key successional species based on average relative abundance and feature importance scores.

[0023] Figure 3 The graph shows the test results used to verify the significance of the prediction accuracy of the classification model.

[0024] Figure 4 A comparative graph showing the relative contributions of different environmental microbial community assembly process types under different microplastic exposure concentrations.

[0025] Figure 5 A time-dynamic succession heatmap of key successor species under different microplastic exposure times. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims and drawings of this application are intended to cover non-exclusive inclusion.

[0028] The following provides a detailed description of this application.

[0029] Figure 1A flowchart of a preferred embodiment of the present invention is provided below. Figure 1 The method for identifying key successional species in microbial communities under microplastic stress provided by this invention includes the following steps: S1. Extract microbial DNA from environmental samples affected by microplastic stress and perform high-throughput sequencing. Generate abundance tables and phylogenetic tree data based on the high-throughput sequencing results and bioinformatics workflow, and obtain a sample metadata table recording microplastic type, exposure concentration, exposure time, and environmental site information. Standardize the abundance matrix in the abundance table into a relative abundance matrix. Group the environmental samples according to the microplastic type, exposure concentration, exposure time, and environmental site information in the sample metadata table. S2. Using the relative abundance of microbial groups as the input feature and the sample group as the target variable, construct a classification model and extract the importance score of each microbial group. Based on the relative abundance matrix, the average relative abundance of each microbial group was calculated; based on the average relative abundance and importance score, the microbial groups were stratified and screened, and low-abundance high-contribution groups and high-abundance high-contribution groups were extracted as candidate key succession species. S3. Construct a phylogenetic distance matrix between microbial groups based on the phylogenetic tree data. The phylogenetic distance matrix records the closeness of the evolutionary relationship between each pair of microbial groups. Based on the phylogenetic distance matrix and relative abundance matrix, the assembly offset index between different samples is calculated; the community assembly process type of each sample pair is determined based on the assembly offset index, and the determination results are statistically analyzed according to the grouping conditions; the relative contributions of deterministic assembly processes and stochastic processes under each grouping condition are obtained based on the statistical results. S4. Based on the relative abundance changes and importance score changes of each candidate key successional species under different grouping conditions, calculate the comprehensive change characteristic value of each candidate key successional species under different grouping conditions. Based on the relative contributions of the deterministic assembly process and the relative contributions of the stochastic process under each grouping condition, calculate the assembly offset characteristic value under the corresponding grouping condition. The assembly offset correlation coefficient of each candidate key successional species is calculated based on the comprehensive change feature value and the assembly offset feature value. The candidate key successional species are then screened a second time based on the assembly offset correlation coefficient to obtain a set of key successional species.

[0030] Specifically, in step S1, after high-throughput sequencing, an amplicon sequence variant (ASV) abundance table or an operational taxonomic unit (OTU) abundance table can be obtained.

[0031] In the ASV abundance table, each row represents an ASV, which can be understood as a high-resolution microbial species, or a microbial group. Each column represents an environmental sample, and the value in the cell indicates the count or abundance of the ASV in that sample.

[0032] In the OTU abundance table, each row represents an OTU. An OTU can be understood as a high-resolution microbial species, that is, a microbial group. Each column represents an environmental sample, and the value in the cell indicates the count or abundance of the OTU in that sample.

[0033] Phylogenetic tree data are typically in Newick format or similar files, describing the evolutionary distances and phylogenetic relationships among all microbial groups in the abundance table.

[0034] In this step, to eliminate the differences in sequencing depth among different samples, the abundance matrix in the abundance table is standardized into a relative abundance matrix. Each value does not represent the original number of sequencing sequences, but rather the percentage or proportion of a certain microbial group in the total microbial community in a specific sample.

[0035] The samples were grouped based on the microplastic type (PE, PP), exposure concentration, exposure time and environmental site information in the sample metadata, and the microbial community analysis dataset required for subsequent analysis was constructed.

[0036] As an example, microplastic types include polylactic acid (PLA), polybutylene adipate / terephthalate (PBAT), polyethylene (PE), and a blank control (CK). When grouping samples, they can be grouped according to microplastic type, exposure concentration, exposure time, and environmental site information. For example, samples belonging to the same microplastic type can be grouped into one group, such as the PE group; samples with the same exposure time can be grouped into one group, such as the 7-day group; samples belonging to the same environmental site can be grouped into one group, such as the soil group; and samples with the same exposure concentration can be grouped into one group.

[0037] Example as follows:

[0038] After grouping: Grouped by microplastic type: {S001, S002, S005} → PE group; {S003} → PP group; {S004} → control group; Grouped by exposure time: {S001, S003, S004, S005} → 7-day group; {S002} → 14-day group; Grouped by exposure concentration: {S001, S002, S003} → High exposure concentration group; {S004} → No exposure concentration group; {S005} → Low exposure concentration group; Grouped by environmental site: {S001, S002, S003, S004, S005} → Soil group.

[0039] For step S2), where the sample group is the label of each group after sample grouping, the importance score is a quantitative indicator that measures the contribution of microbial taxa to the classification model in distinguishing sample groups. The higher the score, the more important the abundance change pattern of the taxa is in distinguishing different groups.

[0040] The classification model can be a random forest classification model, with the number of decision trees preferably set to 1000, and the feature importance evaluation function enabled. Other classification models can also be used, such as XGBoost, LightGBM, LASSO logistic regression, support vector machines, multilayer perceptrons, etc.

[0041] See Figure 2 , Figure 2 This is a candidate key successional species screening diagram based on average relative abundance and feature importance score in one embodiment. The horizontal axis represents average relative abundance in percentage, using a logarithmic scale. The further to the right, the higher the average proportion of the microorganism in the sample. The vertical and horizontal axes represent random forest importance scores. The higher up, the greater the contribution of the microorganism to the classification model in distinguishing different microplastic treatment groups. Each point represents a microbial taxonomy (ASV level) and is labeled with its taxonomic name and ASV number. The dashed lines in the diagram divide the two-dimensional space into four quadrants: the upper left quadrant represents low abundance and high contribution, the upper right quadrant represents high abundance and high contribution, the lower right quadrant represents high abundance and low contribution, and the lower left quadrant represents low abundance and low contribution.

[0042] Furthermore, to verify the reliability of the classification model prediction results in step S2, see [link to relevant documentation]. Figure 3 , Figure 3 This chart shows the test results used to verify the significance of the classification model's prediction accuracy. The bars represent the random accuracy distribution obtained by randomly shuffling the sample labels hundreds to thousands of times, retraining the classification model each time, and calculating the prediction accuracy. The red dashed line represents the accuracy of the classification model trained based on the real sample labels and real microbial abundance data. The greater the distance between the red dashed line and the random accuracy distribution, the less likely the prediction results of the classification model are caused by random factors, and the more reliable the model is. Figure 3 This indicates that the accuracy of the true model is significantly higher than the accuracy distribution under random permutation conditions, demonstrating a real and model-identifiable correlation between microbial abundance distribution and different microplastic treatment groups. This provides reliable model support for the screening results of candidate key successional species in step S2. For step S3, the phylogenetic distance matrix records the evolutionary relationship between each pair of microbial taxa (ASVs). It is a numerical representation calculated from the phylogenetic tree.

[0043] Assuming there are 3 ASVs: ASV_A, ASV_B, and ASV_C, then the phylogenetic distance matrix is:

[0044] The community assembly process types include heterogeneous selection, homogeneous selection, and stochastic processes. Heterogeneous selection and homogeneous selection are both deterministic assembly processes. Heterogeneous selection indicates that environmental differences widen community disparities, while homogeneous selection indicates that environmental pressures reduce community disparities. The community disparities in stochastic processes originate from random fluctuations.

[0045] In summary, in this application, step S1 completed the data preparation and grouping framework construction, processed the raw sequencing data into a standardized relative abundance matrix, eliminated the systematic bias caused by the difference in sequencing depth, and made different samples comparable; and grouped the samples according to four dimensions: microplastic type, exposure concentration, exposure time and environmental site, thus constructing a multi-dimensional stress response research framework.

[0046] Step S2 uses relative microbial abundance as the input feature and sample group as the target variable. A classification model quantifies the contribution of each microbial group to distinguishing different microplastic treatment conditions. The average relative abundance and importance score are cross-combined to create four categories. The low-abundance, high-contribution groups specifically correspond to microorganisms that, although few in number, exhibit significant changes under stress conditions and are crucial for group differentiation. This design compensates for the shortcomings of traditional differential abundance methods that neglect low-abundance functional groups.

[0047] Furthermore, the microbial abundance matrix constructed in step S1 is extremely high-dimensional, typically containing tens of thousands of ASVs or OTUs. Directly performing subsequent assembly offset association analysis on all these microorganisms would not only be computationally intensive, but also subject to significant interference from background noise species that would severely hinder the identification of statistical signals. Step S2 stratifies and filters microbial taxa using average relative abundance and importance scores, thereby reducing the analytical targets from tens of thousands of microbial taxa in the original matrix to a finite-sized candidate set. Each species in this candidate set significantly contributes to distinguishing different microplastic stress conditions.

[0048] Step S3, based on the grouping conditions set in step S1 (microplastic type, exposure concentration, etc.), calculates the relative contributions of deterministic and stochastic processes to each group. These two values ​​are directly output to step S4 to construct assembly offset eigenvalues. This step transforms abstract ecological mechanisms into quantifiable variables, providing quantifiable variables of community assembly processes to overcome the lack of ecological mechanisms.

[0049] In step S4, the comprehensive change eigenvalue integrates relative abundance changes and importance scores, reflecting the intensity of microbial species' response to microplastic stress. The assembly shift eigenvalue reflects the net intensity of the community assembly process deviating from a random state under the given grouping conditions. A larger value indicates stronger deterministic selection driven by stress conditions. The assembly shift correlation coefficient measures the synchronicity between the changing trends of species response intensity and community assembly shift intensity under different microplastic types, concentrations, times, and sites. By using the assembly shift correlation coefficient, the screening of candidate species is elevated from statistical correlation to mechanistic correlation, effectively reducing the false positive rate and directly addressing the problem of insufficient ecological mechanism support.

[0050] Furthermore, the method for identifying key successional species in microbial communities under microplastic stress provided in this application also includes step S5: S5. Construct a time response curve based on at least one of the abundance changes, importance score changes, or assembly offset correlation value changes of the key successional species at different exposure time points, and determine the response type of the key successional species based on the time response curve.

[0051] Specifically, the analysis in steps S1-S4 is mainly based on the comparison between grouping conditions. Although exposure time is one of the grouping criteria in step S1, the calculation of the assembly offset correlation coefficient in step S4 treats each time point as an independent grouping condition, without showing the continuous change trajectory of the same species at different time points. Thus, even if a species is known to be a key successor species, it is unknown at what time stage its key response occurred.

[0052] After introducing the time-response curve in step S5, each key successional species is assigned a dynamic trajectory over time. This allows researchers to see how a species gradually responds over time with microplastic exposure: whether it rises rapidly and then falls back, accumulates slowly and then explodes, or remains at a high level. Understanding the time response types of key successional species is helpful for subsequent research and applications.

[0053] Among them, early-response species are those whose response peaks appear in the early part of the time series. These species may represent pioneer groups or opportunists that respond rapidly to microplastic stress. Continuously responsive species are those whose responses remain at a high level over a longer period of time. These species may represent stable functional groups that are deeply involved in the process of stress adaptation. Lag response species are those whose response peaks appear in the later part of the time series, and may represent groups that produce secondary responses to earlier community changes or the accumulation of metabolites.

[0054] In some examples, key succession species whose peak response period is in the first 1 / 3 of all time points are identified as early-response species; key succession species whose peak response period is in the first 1 / 3 of all time points after the peak response are identified as sustained-response species; and key succession species whose peak response period is in the second 1 / 3 of all time points are identified as delayed-response species.

[0055] Furthermore, the peak response periods of each key successional species can be output. In this way, in environmental monitoring applications, appropriate time windows can be selected for targeted sampling based on the peak periods of different response types, thereby improving monitoring efficiency.

[0056] Furthermore, step S5 also includes: using Dynamic Time Warping (DTW) to calculate the temporal response distance between key successional species, and performing temporal clustering based on the time response distance.

[0057] Specifically, while the responses of different key successional species may differ in onset time and rate, their response patterns may be fundamentally similar in "shape." Traditional distance metrics require strict point-by-point alignment of two sequences on the timeline, failing to handle stretching, displacement, or phase shifts along the timeline. Consequently, two species with highly similar response patterns but slightly misaligned peak times would be judged as highly different by traditional methods.

[0058] The DTW algorithm uses dynamic programming to find the optimal nonlinear alignment path between two time series, tolerating local stretching or compression on the time axis, thus more realistically measuring the similarity of curve shapes. In this way, even if different species have different response rhythms, similar response patterns can be identified, improving the robustness and biological plausibility of clustering results.

[0059] Furthermore, time-based clustering can automatically group all key successional species into a limited number of groups based on the similarity of their curve morphology. This allows researchers to quickly identify the key species communities under microplastic stress and determine the typical response patterns, such as unimodal response, impulse response, and progressive response modes.

[0060] In some embodiments of this application, the comprehensive variation characteristic value F of each candidate key successional species under different grouping conditions is... ik Calculate using the following formula:

[0061] in, Indicates the first The candidate key successor species in the The comprehensive change characteristic value under each grouping condition; The candidate key succession species is indicated in the first... Relative abundance changes under each grouping condition; The candidate key succession species is indicated in the first... Changes in importance scores under each grouping condition; Represents a standardized function; and These are the weighting coefficients.

[0062] Specifically, in this embodiment, and By incorporating them into the same indicator through weighted summation, It also captured response information of candidate key succession species at both the level of quantitative changes and structural importance changes.

[0063] Formula requirements and Z-standardization was performed separately before weighted summation. Standardization transforms two indicators with different dimensions to the same scale, resulting in comparable means and variances. This avoids computational inappropriateness caused by the original value of one indicator being much larger than that of another, ensuring that changes in abundance and importance scores play their due roles in the comprehensive feature value according to their expected weights, thus improving the fairness and effectiveness of the comprehensive score.

[0064] Furthermore, ω1 and ω2 in the formula are weighting coefficients. Researchers can adjust the relative contributions of abundance changes and importance score changes based on specific research objectives or data characteristics. By default, both can be set to be equal (e.g., 0.5 each) to achieve a balanced assessment. In studies emphasizing the ability to distinguish community structure, ω2 can be increased to highlight changes in importance score.

[0065] It should be noted that the calculation of the comprehensive change characteristic value is not limited to the specific method provided in this application. The formula provided in this implementation essentially standardizes the change values ​​of multiple dimensions and then obtains a scalar through weighted summation to comprehensively measure the intensity of a species' response to stress. Based on this logic, alternative calculation methods for the comprehensive change characteristic value can also have other variations, such as weighted average form, Euclidean summation form, etc.

[0066] Furthermore, in some embodiments of this application, in step S4, the assembly offset characteristic value corresponding to the grouping condition is... Calculate using the following formula:

[0067] in, Indicates the first Assembly offset characteristic values ​​under each grouping condition; Indicates the first The relative contribution of the deterministic assembly process under each grouping condition; Indicates the first The relative contribution of a stochastic process under each grouping condition; This represents the standardized function.

[0068] Specifically, in this implementation, the formula is obtained through... and The difference operation compresses two parallel relative contribution values ​​into a single scalar. This scalar intuitively reflects the direction of the community assembly process under the given grouping conditions; a positive value indicates that deterministic selection is dominant, a negative value indicates that random selection is dominant, and the absolute value indicates the degree of deviation from the equilibrium state. This simplifies the complexity of subsequent association analysis.

[0069] Z() standardization allows for different grouping conditions. The values ​​are converted to a unified, standardized scale. This allows for direct comparison of assembly offset intensity caused by different microplastic types on the same scale; quantitative comparison of assembly offset effects under different exposure concentration gradients; and analysis of experimental results from different environmental sites within a unified framework. This enhances the applicability of the method in multi-condition, multi-batch experiments and the integrability of results, providing standardized quantitative indicators for large-scale meta-analysis and comparative studies.

[0070] It should be noted that the assembly offset characteristic value The calculation method is not limited to the specific method provided in this application. In other embodiments, the assembly offset characteristic value... It can also be based on rank-standardized forms, weighted synthesis forms, etc.

[0071] Furthermore, in some embodiments of this application, in step S4, the assembly offset correlation coefficient of each candidate key successional species... Calculate using the following formula:

[0072] in, Indicates the first Assembly offset correlation coefficients of candidate key successional species; Indicates the total number of grouping conditions; Indicates the first The average value of the comprehensive change characteristics of the candidate key successional species; This represents the average value of the assembly offset characteristic value.

[0073] Specifically, This is used to measure the synchronicity between the changing trends of species response intensity and community assembly offset intensity under multiple different grouping conditions.

[0074] If a species Value at High values ​​are also high under the grouping condition. The value is low under the grouping condition of low value. A value close to +1 indicates that the response strength of this species is highly synchronized with the deterministic assembly offset strength. If a species Value at The value is lower when it is high. A value close to -1 indicates that the species' response is inversely related to the deterministic assembly bias and may be more active in stochastic processes; If a species and There is no regular synchronization relationship. A value close to 0 indicates that the species' response is not significantly associated with community assembly mechanisms.

[0075] The denominator of the formula is sequence standard deviation and The product of the standard deviations of the series. By dividing by this product, It is standardized to a dimensionless pure numerical value and is unaffected by... and The influence of the original dimensions or variance. Thus, the different candidate species... Values ​​can be directly sorted and compared. Researchers can set a uniform AOC threshold to screen all candidate species without needing to adjust the threshold individually based on the data characteristics of each species. This greatly improves the automation and reproducibility of the method.

[0076] Furthermore, in one embodiment of this application, step S4 further includes: verifying the significance of the assembly offset correlation coefficient through a permutation test; when the assembly offset correlation coefficient is not lower than a preset threshold and the significance test result meets the preset significance level, the corresponding candidate key succession species is retained as a key succession species.

[0077] Specifically, the permutation test does not rely on theoretical assumptions about data distribution. Instead, it generates a distribution of correlation coefficients in an uncorrelated state by randomly rearranging the data to disrupt the correspondence between the comprehensive change characteristic values ​​and the assembly offset characteristic values. Only species with sufficiently high correlation coefficients that significantly deviate from a random distribution are preserved.

[0078] In this implementation, the selection of key successional species has a dual threshold. First, the correlation coefficient must be no lower than a preset threshold to ensure a sufficiently strong association between the species response and the assembly offset. Second, the significance test must meet a preset significance level to ensure that the association strength is not due to random fluctuations. Both thresholds are indispensable: high but insignificant correlation coefficients, or significant but low absolute correlation coefficients, are not retained. This significantly improves the reproducibility and reliability of the screening results and further reduces the false positive rate.

[0079] It should be noted that the permutation test is a classic and well-established statistical method, which will not be elaborated upon here.

[0080] Furthermore, step S1) also includes obtaining a species classification annotation table based on the high-throughput sequencing results, and integrating the abundance table, species classification annotation table, sample metadata table, and phylogenetic tree data into a structurally unified and interconnected microbial community analysis data object.

[0081] Specifically, microbial community analysis involves various types of data objects: abundance tables, taxonomic annotation tables, metadata tables, and phylogenetic tree data. These data have different dimensions and structures, but they also have inherent relationships. If they are not integrated into a unified microbial community analysis data object, subsequent steps require repeated table joins, matching, and operations, which can easily lead to errors.

[0082] In step S1), by integrating the four data tables into a single, structurally unified, and interconnected object, a corresponding relationship between the tables is enforced. This allows any mismatched entries to be identified and corrected during the integration process, thereby preventing calculation errors caused by data mismatches in subsequent analyses and improving the accuracy and reliability of the entire methodology. Furthermore, the unified structure of the data object facilitates tracing the original data source corresponding to each analysis result.

[0083] Furthermore, step S1) also includes the steps of quality control and filtering of the microbial community analysis data objects.

[0084] Specifically, in this implementation, the role of quality control and filtering is to remove non-target sequences and invalid low-abundance species, ensuring that the data matrices input to S2 to S5 reflect real, effective, and high-quality target microbial community information, and preventing noise and erroneous data from spreading to downstream analysis.

[0085] In this application, the objects of quality control and filtering can be: 1) Sequences not classified as bacterial domains; 2) Non-target sequences classified as mitochondria or chloroplasts; 3) Species with a total abundance of zero in all samples.

[0086] Furthermore, in some embodiments of this application, step S3) includes, before calculating the assembly offset index between different samples, a step of matching common species and synchronizing the phylogenetic distance matrix with the relative abundance matrix.

[0087] Common species matching refers to identifying microbial communities that coexist in the phylogenetic distance matrix and the relative abundance matrix; synchronous rearrangement ensures that the order of ASVs in the phylogenetic distance matrix and the relative abundance matrix is ​​completely consistent. This guarantees the mathematical validity and logical correctness of subsequent calculations from a data perspective, avoiding systematic calculation errors caused by mismatched or inconsistent species sets.

[0088] for example:

[0089] Furthermore, the average distance between the nearest observed taxa among samples can be calculated based on the rearranged results, and the community assembly shift index can be obtained through null model randomization analysis. The type of community assembly process is determined according to the following rules: When betaNTI is greater than a preset positive threshold, it is determined to be a heterogeneous selection; When betaNTI is less than the preset negative threshold, it is determined to be a homogeneous selection; When betaNTI is between the positive and negative thresholds, it is determined to be a random process.

[0090] Among them, heterogeneous selection indicates that poor environmental conditions exacerbate community differences, while homogeneous selection indicates that environmental pressure reduces community differences. Random process community differences originate from random fluctuations. Preferably, the preset positive threshold is 2, and the preset negative threshold is -2.

[0091] See Figure 4 , Figure 4 This figure compares the relative contributions of different microplastic exposure concentrations to the assembly process of environmental microbial communities. The horizontal axis represents exposure time, and the vertical axis represents relative contribution. Red blocks represent heterogeneous selection, blue blocks represent homogeneous selection, and gray blocks represent random selection. High indicates high concentration, and Low indicates low concentration. This figure shows that different microplastic types exhibit different time-response patterns and peak response periods. This invention can identify differences in microbial community succession characteristics under different degradable microplastic stress conditions.

[0092] To verify the applicability of the method of the present invention under different degradable microplastic conditions, the optimized data processing flow of steps S1 to S5 was used to analyze the PLA microplastic exposure group and the PBAT microplastic exposure group samples, respectively.

[0093] The analysis results show that by using steps S1 to S5 of the present invention, not only can candidate key succession species of PLA group and PBAT group be extracted respectively, but after quantitative retention in step S4, the final sets of key succession species obtained by the two groups are significantly different.

[0094] Meanwhile, the time response type and peak response period output by step S5 are also different, indicating that the present invention can identify the differences in the succession characteristics of microbial communities under different degradable microplastic stress conditions.

[0095] Figure 5 This is a heatmap showing the temporal dynamics of key successor species under different microplastic exposure times. Each row on the vertical axis represents a key successor species retained after correlation coefficient screening, labeled with its genus name and ASV number. The four columns on the left of the horizontal axis represent the exposure times of PBAT at 30, 60, 90, and 120 days, and the four columns on the right represent the exposure times of PLA at 30, 60, 90, and 120 days. The color of the blocks represents the relative abundance: yellow indicates a relatively high abundance under this treatment condition (high value); green or cyan indicates a moderate or slightly high relative abundance; and blue or purple indicates a relatively low abundance under this treatment condition (low value). This figure illustrates that key microplastic-related taxa undergo phased replacement over time, and this replacement pattern is significantly influenced by the type of plastic.

[0096] To verify the anti-interference effect of the quantitative retention rule in step S4 of this invention, the optimized data processing flow of steps S1 to S5 described above was used to analyze the recalcitrant microplastic PE group and the blank control CK group.

[0097] When only steps S1 to S2 are performed, the classification model can still output several candidate groups with high importance scores. However, after further performing steps S3 and S4, only candidate groups whose assembly offset correlation coefficient (AOC) reaches a preset threshold and passes the significance test are retained as key successional species. This result shows that step S4 can effectively eliminate candidate groups that only have statistical differences but lack assembly mechanism support, thereby reducing the risk of false positives and improving the stability of the final identification results.

[0098] After completing the analysis using the method of this invention, at least one of the following results or a combination thereof will be output: 1) A list of candidate key succession species, including the taxonomic name of each candidate species, its sample group, average relative abundance, feature importance score and its ecological role category; 2) Assembly offset correlation coefficient result table, including candidate key successional species names, AOC values, significance test results, and retention determination results; 3) Community assembly process determination result table, including betaNTI values, assembly process types and their relative contributions under different sample grouping conditions; 4) Key successional species-assembly offset association results table, including the association direction, association strength or retention determination results between each key successional species and changes in community assembly process type under different microplastic types, exposure concentrations, exposure times or environmental site conditions; 5) Classification table of time response patterns of key successional species, including the early response type, sustained response type or delayed response type category of each key successional species and its peak response period; 6) Summary table of critical response windows, used to summarize the peak response time information of key successional species under different microplastic types, exposure concentrations and environmental site conditions; 7) One or more of the following: candidate key successional species screening map, assembly offset correlation coefficient test map, community assembly process type distribution map, and key successional species time response pattern clustering map.

[0099] It should be noted that the conventional bioinformatics analysis steps, reagent preparation steps, and sequencing library construction steps not detailed in the embodiments can all be implemented using conventional techniques well known to those skilled in the art. The above embodiments are merely preferred embodiments of the present invention and are not intended to limit the present invention. All equivalent substitutions, improvements, or modifications made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

[0100] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. These modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for identifying key successional species in microbial communities under microplastic stress, characterized in that, Includes the following steps: S1. Extract microbial DNA from environmental samples affected by microplastic stress and perform high-throughput sequencing. Generate abundance tables and phylogenetic tree data based on the high-throughput sequencing results and bioinformatics workflow, and obtain a sample metadata table recording microplastic type, exposure concentration, exposure time, and environmental site information. Standardize the abundance matrix in the abundance table into a relative abundance matrix. Group the environmental samples according to the microplastic type, exposure concentration, exposure time, and environmental site information in the sample metadata table. S2. Using the relative abundance of microbial groups as the input feature and the sample group as the target variable, construct a classification model and extract the importance score of each microbial group. Based on the relative abundance matrix, the average relative abundance of each microbial group was calculated; Microbial taxa were stratified and screened based on average relative abundance and importance scores, and low-abundance high-contribution taxa and high-abundance high-contribution taxa were extracted as candidate key succession species. S3. Construct a phylogenetic distance matrix between microbial groups based on the phylogenetic tree data. The phylogenetic distance matrix records the closeness of the evolutionary relationship between each pair of microbial groups. Based on the phylogenetic distance matrix and relative abundance matrix, the assembly offset index between different samples is calculated; the community assembly process type of each sample pair is determined based on the assembly offset index, and the determination results are statistically analyzed according to the grouping conditions. The relative contributions of deterministic assembly processes and stochastic processes under each grouping condition are obtained based on the statistical results. S4. Based on the relative abundance changes and importance score changes of each candidate key successional species under different grouping conditions, calculate the comprehensive change characteristic value of each candidate key successional species under different grouping conditions. Based on the relative contributions of the deterministic assembly process and the relative contributions of the stochastic process under each grouping condition, calculate the assembly offset characteristic value under the corresponding grouping condition. The assembly offset correlation coefficient of each candidate key successional species is calculated based on the comprehensive change feature value and the assembly offset feature value. The candidate key successional species are then screened a second time based on the assembly offset correlation coefficient to obtain a set of key successional species.

2. The method for identifying key successional species in microbial communities under microplastic stress according to claim 1, characterized in that, It also includes step S5: S5. Construct a time response curve based on at least one of the abundance changes, importance score changes, or assembly offset correlation value changes of the key successional species at different exposure time points, and determine the response type of the key successional species based on the time response curve.

3. The method for identifying key successional species in microbial communities under microplastic stress according to claim 2, characterized in that, Step S5 further includes: using Dynamic Time Warping (DTW) to calculate the temporal response distance between key successional species, and performing temporal clustering based on the time response distance.

4. The method for identifying key successional species in microbial communities under microplastic stress according to any one of claims 1 to 3, characterized in that: In step S4, the comprehensive change characteristic values ​​of each candidate key successional species under different grouping conditions are... Calculate using the following formula: in, Indicates the first The candidate key successor species in the The comprehensive change characteristic value under each grouping condition; The candidate key succession species is indicated in the first... Relative abundance changes under each grouping condition; The candidate key succession species is indicated in the first... Changes in importance scores under each grouping condition; Represents a standardized function; and These are the weighting coefficients.

5. The method for identifying key successional species in microbial communities under microplastic stress according to claim 4, characterized in that: In step S4, the assembly offset characteristic value corresponding to the grouping conditions is... Calculate using the following formula: in, Indicates the first Assembly offset characteristic values ​​under each grouping condition; Indicates the first The relative contribution of the deterministic assembly process under each grouping condition; Indicates the first The relative contribution of a stochastic process under each grouping condition; This represents the standardized function.

6. The method for identifying key successional species in microbial communities under microplastic stress according to claim 5, characterized in that: In step S4, the assembly offset correlation coefficients of each candidate key successional species are... Calculate using the following formula: in, Indicates the first Assembly offset correlation coefficients of candidate key successional species; Indicates the total number of grouping conditions; Indicates the first The average value of the comprehensive change characteristics of the candidate key successional species; This represents the average value of the assembly offset characteristic value.

7. The method for identifying key successional species in microbial communities under microplastic stress according to any one of claims 1 to 3, characterized in that: Step S4 further includes: verifying the significance of the assembly offset correlation coefficient through a permutation test; when the assembly offset correlation coefficient is not lower than a preset threshold and the significance test result meets the preset significance level, the corresponding candidate key succession species is retained as a key succession species.

8. The method for identifying key successional species in microbial communities under microplastic stress according to any one of claims 1 to 3, characterized in that: Step S1) also includes obtaining a species classification annotation table based on the high-throughput sequencing results, and integrating the abundance table, species classification annotation table, sample metadata table, and phylogenetic tree data into a structurally unified and interconnected microbial community analysis data object.

9. The method for identifying key successional species in microbial communities under microplastic stress according to claim 8, characterized in that: Step S1) also includes the steps of quality control and filtering of the microbial community analysis data objects.

10. The method for identifying key successional species in microbial communities under microplastic stress according to claim 9, characterized in that: Step S3) Before calculating the assembly offset index between different samples, it also includes the step of matching common species and synchronously rearranging the phylogenetic distance matrix and the relative abundance matrix.