A method and system for screening of ecological memory retention species from waste polypropylene mask

By combining high-throughput sequencing and random forest regression models with microbial co-occurrence networks and neutral community models, species that retain ecological memory in biofilms from discarded polypropylene masks were screened. This solved the problem that existing technologies could not quantify the degree of ecological memory retention and screen key species, and achieved quantitative characterization of ecological memory retention and screening of ecological reliability.

CN122337348APending Publication Date: 2026-07-03NORTHWEST A & F UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NORTHWEST A & F UNIV
Filing Date
2026-06-05
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing technologies cannot quantify the degree of ecological memory retention of biofilms after waste polypropylene masks are transferred across water bodies, and it is difficult to selectively screen key species directly related to ecological memory from high-dimensional microbial data.

Method used

By obtaining baseline and transfer samples of biofilm from discarded polypropylene masks, high-throughput sequencing and data quality control were performed to construct a relative abundance matrix. Key microorganisms were screened using a random forest regression model, and species that maintain ecological memory were identified by combining a microbial co-occurrence network and a Sloan neutral community model.

Benefits of technology

This technology enables quantitative characterization of ecological memory retention, allowing the screening of super core species with the highest ecological reliability. It solves the problem that existing technologies cannot quantify the degree of ecological memory retention and screen key species.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122337348A_ABST
    Figure CN122337348A_ABST
Patent Text Reader

Abstract

This invention relates to the field of ecological conservation technology, specifically to a method and system for screening species that preserve ecological memory from discarded polypropylene masks. The method includes: obtaining baseline and transfer samples of biofilm from discarded polypropylene masks; obtaining a relative abundance matrix after DNA extraction, high-throughput sequencing, and quality control standardization; and extracting the top-ranking microorganisms from the transfer group to construct an independent variable feature matrix. Ecological memory retention is calculated based on the Bray-Curtis distance between the transfer group and the homologous baseline group. A random forest model is trained using the feature matrix, and key microorganisms are screened based on the error changes after shuffling the abundance. Ecological memory-preserving species are then screened based on the Spearman correlation between their abundance and retention. Finally, the intersection of ecological memory-preserving species, network key species, and neutral community key species is defined as the super-core species. This invention improves the accuracy of screening ecological memory-preserving species.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of ecological conservation technology, and in particular to a method and system for screening species that preserve the ecological memory of discarded polypropylene masks. Background Technology

[0002] This invention addresses the lack of quantitative characterization methods for the "ecological memory" phenomenon of microbial communities on the surface of discarded polypropylene masks after they are transferred across water bodies, and the difficulty of traditional community analysis methods in identifying key species directly related to ecological memory from high-dimensional microbial data. It proposes a machine learning-based screening method. Existing research mainly relies on high-throughput sequencing combined with conventional statistical methods such as principal coordinate analysis and permutation multivariate analysis of variance to describe differences in community structure and succession trends. However, these methods cannot pinpoint the overall changes in the community to specific microbial groups, nor can they quantitatively analyze the degree of memory retention and screen key species accordingly. Therefore, there is an urgent need for an analytical process that can quantify the degree of ecological memory retention and selectively screen relevant species. Summary of the Invention

[0003] This invention provides a method and system for screening species that retain ecological memory in discarded polypropylene masks, which solves the problem that existing technologies cannot quantify the "degree of ecological memory retention" of biofilms after the transfer of discarded polypropylene masks across water bodies, nor can they selectively screen key species directly related to ecological memory from high-dimensional microbial data.

[0004] The objective of this invention can be achieved through the following technical solutions:

[0005] The first aspect of this invention is to provide a method for screening species that retain ecological memory from discarded polypropylene face masks, comprising: Several baseline samples and several transfer samples of biofilm from discarded polypropylene masks were obtained; Microbial DNA was extracted from all baseline and transfer samples and subjected to high-throughput sequencing, data quality control, and standardization to obtain a relative abundance matrix; the rows in the relative abundance matrix represent various microorganisms, and the columns represent all baseline and transfer samples. Based on the relative abundance matrix, the overall relative abundance of each microorganism is obtained. Then, the microorganisms with the highest overall relative abundance are extracted and an independent variable feature matrix is ​​constructed with all translocation group samples. The rows in the independent variable feature matrix represent the extracted microorganisms, and the columns represent all translocation group samples. The ecological memory retention of each transfer group sample is obtained by comparing the relative abundance differences between each transfer group sample and all corresponding homologous baseline group samples in the feature matrix of independent variables; a target variable data table is constructed using the transfer group sample number and ecological memory retention as column variables. A random forest regression model is trained using the feature matrix of independent variables and the data table of target variables to obtain the trained random forest regression model. The relative abundance of each microorganism in the feature matrix of independent variables is randomly shuffled and then input into the trained random forest regression model along with the relative abundance data of the feature matrix of independent variables to calculate the error. Key microorganisms are selected based on the error calculation results. By analyzing the relationship between the relative abundance and ecological memory retention of each key microorganism in samples from different transfer groups, species that retain ecological memory were screened from the key microorganisms. A microbial co-occurrence network was obtained, and key species of the network were identified through the co-occurrence network. Key species of the neutral community were identified by the relationship between the average relative abundance of all transfer group samples of each microorganism and the frequency of transfer group samples with relative abundance greater than 0. Then, the intersection species of ecological memory-maintaining species, key species of the network, and key species of the neutral community were selected as the super core species for final screening.

[0006] Furthermore, the process of obtaining the relative abundance matrix includes: Microbial DNA was extracted from all baseline and transfer samples and subjected to high-throughput sequencing to obtain the original abundance matrix. The original abundance matrix was used to identify different microorganisms and listed as all baseline and transfer samples. Data quality control and standardization were performed on the original abundance matrix to obtain the relative abundance matrix. The data quality control includes removing sequences not classified as bacterial domains; removing non-target sequences classified as mitochondria or chloroplasts; and removing species with a total abundance of zero in all samples.

[0007] Further, obtaining the overall relative abundance of each microorganism based on the relative abundance matrix includes: The overall relative abundance of each microorganism is the sum of the relative abundances of all baseline and transfer samples for each microorganism in the relative abundance matrix.

[0008] Furthermore, the ecological memory retention of each transfer group sample includes: The relative abundance of all microorganisms in each column of the relative abundance matrix is ​​used to form a relative abundance sequence; the BC distance between any two columns is calculated using the Bray-Curtis algorithm based on the relative abundance sequences between any two columns. For each transfer group sample, identify its corresponding homologous baseline group sample, and calculate the mean BC distance between each transfer group sample and all corresponding homologous baseline group samples; use this as the average distance of each transfer group sample; and use the difference between 1 and the average distance of each transfer group sample as the ecological memory retention of each transfer group sample.

[0009] Further, the step of randomly shuffling all relative abundances of each microorganism in the independent variable feature matrix, and then inputting them, along with the relative abundance data from the independent variable feature matrix, into the trained random forest regression model for error calculation includes: The feature matrix of the independent variables is input into the trained random forest regression model to predict the ecological memory retention. The mean square error is calculated by comparing the predicted ecological memory retention with the ecological memory retention in the target variable data table, and is denoted as the original mean square error. The relative abundance of all samples from the transition groups corresponding to a microorganism in the independent variable feature matrix is ​​randomly shuffled to obtain a new independent variable feature matrix for each microorganism. The new independent variable feature matrix for each microorganism is then input into the trained random forest regression model to predict the ecological memory retention. The mean squared error is calculated by comparing the predicted ecological memory retention with the ecological memory retention in the target variable data table, and is denoted as the shuffle mean squared error. Based on the difference between the original mean squared error and the shuffled mean squared error, the percentage increase in mean squared error for each microorganism is obtained, expressed by the formula:

[0010] In the formula, This represents the original mean square error. This indicates the scrambled mean square error. This represents the percentage increase in mean squared error for each microorganism.

[0011] Furthermore, the screening process for the key microorganisms includes: Key microorganisms are selected based on the percentage increase in mean square error of all microorganisms. The specific process for selecting key microorganisms is as follows: the maximum curvature method is used to detect the inflection point of the microbial sequences arranged in descending order of the percentage increase in mean square error. The percentage increase in mean square error at the inflection point is used as the percentage threshold, and microorganisms with a percentage increase in mean square error greater than or equal to the percentage threshold are selected as key microorganisms.

[0012] Furthermore, the selection process for the ecological memory-preserving species includes: The independent variable feature matrix is ​​transposed and then horizontally concatenated with the target variable data table to obtain the machine learning dataset; each row in the merged machine learning dataset corresponds to a translocated group sample, and the columns include the relative abundance of microorganisms and the retention of ecological memory. The Spearman correlation coefficient between the relative abundance and ecological memory retention of each key microorganism in each transfer group sample was calculated and denoted as the Spearman correlation coefficient of each key microorganism. Key microorganisms with a Spearman correlation coefficient greater than 0 were identified as species that retain ecological memory.

[0013] Furthermore, the acquisition of the microbial co-occurrence network, and the identification of key species in the network through the co-occurrence network, includes: The significance of Spearman correlation coefficient is used to determine whether there is a co-occurrence relationship between any two microorganisms; all microorganisms are treated as nodes and co-occurrence relationships as edges, thus forming a microbial co-occurrence network; The microbial co-occurrence network is automatically divided into several modules using a module detection algorithm; the average connection strength between each node and all other connected nodes within the same module is calculated. Calculate the average connection strength between each node and all connected nodes in other modules. The connection strength between two nodes refers to the absolute value of the Spearman correlation coefficient of the relative abundance of all metastasis samples of the two microorganisms. Filter out Greater than the first preset threshold, and Microorganisms corresponding to nodes that exceed a second preset threshold are designated as key species in the network.

[0014] Furthermore, the screening process for key species in the neutral community includes: The average relative abundance and observed frequency of each microorganism in all transfer group samples were obtained by using the feature matrix of independent variables; where the observed frequency is the ratio of the frequency of each microorganism in all transfer group samples to the total number of transfer group samples; where the frequency of each microorganism in all transfer group samples is the number of transfer group samples with a relative abundance greater than 0. Based on the average relative abundance and observed occurrence frequency of all transgenic samples of all microorganisms, a Sloan neutral community model was fitted using the least squares method to obtain the fitted Sloan neutral community model, and the 95% confidence interval of the occurrence frequency was determined. The average relative abundance of all transgenic samples of each microorganism in the feature matrix of independent variables was input into the fitted Sloan neutral community model, and the predicted occurrence frequency was output. The predicted occurrence frequency was compared with the upper bound of the 95% confidence interval. If it was higher than the upper bound, it was determined to be a key species in the neutral community.

[0015] A second aspect of the present invention is to provide a screening system for the preservation of ecological memory species in discarded polypropylene masks, comprising: Data acquisition and preprocessing module: used to acquire several baseline group samples and several transfer group samples of biofilm from discarded polypropylene masks; extract microbial DNA from all baseline group samples and transfer group samples, and perform high-throughput sequencing, data quality control and standardization to obtain a relative abundance matrix; based on the relative abundance matrix, obtain the overall relative abundance of each microorganism, then extract the microorganisms with the highest overall relative abundance, and construct an independent variable feature matrix with all transfer group samples; Ecological memory retention calculation module: used to obtain the ecological memory retention of each transfer group sample by the difference in relative abundance between each transfer group sample and all corresponding homologous baseline group samples in the independent variable feature matrix; construct the target variable data table with the transfer group sample number and ecological memory retention as column variables; Key microorganism extraction module: This module trains a random forest regression model using the independent variable feature matrix and the target variable data table to obtain the trained random forest regression model. It randomly shuffles all relative abundances of each microorganism in the independent variable feature matrix and inputs them, along with the relative abundance data from the independent variable feature matrix, into the trained random forest regression model to calculate the error. Key microorganisms are then selected based on the error calculation results. Ecological memory retention species screening module: used to screen ecological memory retention species from key microorganisms by analyzing the relationship between the relative abundance of each key microorganism in each transfer group sample and the degree of ecological memory retention; Multi-dimensional joint screening module: used to obtain the microbial co-occurrence network, and to identify key species in the network through the co-occurrence network; to identify key species in the neutral community by the relationship between the average relative abundance of all transfer group samples of each microorganism and the frequency of transfer group samples with relative abundance greater than 0; and then to select the super core species as the intersection of ecological memory-maintaining species, key species in the network and key species in the neutral community.

[0016] Compared with existing technologies, the beneficial effects of this invention are as follows: Several baseline samples and several transfer samples of biofilm from discarded polypropylene masks were obtained; a control experiment basis for the ecological memory retention phenomenon was established by setting up baseline and transfer groups; microbial DNA was extracted from all baseline and transfer samples, and high-throughput sequencing, data quality control, and standardization were performed to obtain a relative abundance matrix; based on the relative abundance matrix, the overall relative abundance of each microorganism was obtained, and then the microorganisms with the highest overall relative abundance were extracted and used to construct an independent variable feature matrix with all transfer samples; through quality control, standardization, and feature screening, a high-quality model was provided for subsequent models. Low-noise input data; the ecological memory retention of each transfer group sample is obtained by comparing the relative abundance differences between each transfer group sample and all corresponding homologous baseline group samples; a target variable data table is constructed using the transfer group sample number and ecological memory retention as column variables; the abstract ecological memory phenomenon is quantified into a calculable continuous index, realizing a quantitative characterization of the degree of memory retention; a random forest regression model is trained using the independent variable feature matrix and the target variable data table to obtain the trained random forest regression model; all relative abundances of each microorganism in the independent variable feature matrix are randomly shuffled and then compared with the relative abundances in the independent variable feature matrix. The high-dimensional microbial data were input into the trained random forest regression model for error calculation. Key microorganisms were screened based on the error calculation results. Using the feature importance assessment function of random forest, key microorganisms that contribute most to memory retention prediction were automatically screened from the high-dimensional microbial data. Ecological memory-retaining species were screened from the key microorganisms by analyzing the relationship between the relative abundance of each key microorganism in each transfer group sample and its ecological memory retention. Correlation analysis was used to distinguish between positively and negatively correlated species, providing a clear ecological interpretation of the screening results. Key network species were identified through the network connectivity relationships between microorganisms. Each... By analyzing the relationship between the average relative abundance of all transfer group samples of a microorganism and the frequency of transfer group samples with relative abundance greater than 0, key species in neutral communities were identified. Then, the intersection of ecological memory-retaining species, network key species, and key species in neutral communities was used as the final super core species for screening. Through multi-dimensional cross-validation using co-occurrence networks and neutral models, false positives were eliminated, and the super core species with the highest ecological reliability were finally screened. This approach solves the problem that existing technologies cannot quantify the "degree of ecological memory retention" of biofilms after the transfer of discarded polypropylene masks across water bodies, nor can they selectively screen key species directly related to ecological memory from high-dimensional microbial data. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This invention provides a step-by-step flowchart of a method for screening species that preserve the ecological memory of discarded polypropylene masks. Figure 2 A schematic diagram of the module flow of a screening system for preserving species in the ecological memory of discarded polypropylene masks provided by the present invention; Figure 3 This is a schematic diagram illustrating the process of screening microorganisms by summing their relative abundance. Figure 4 A schematic diagram of the Spearman correlation coefficient between relative abundance and ecological memory retention; Figure 5 A schematic diagram showing the screening results of species whose ecological memory has been damaged; Figure 6 This is a schematic diagram illustrating the screening process for key species in a neutral community. Detailed Implementation

[0019] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0020] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0021] To address the problems existing in the background technology, a screening method and system for preserving species through ecological memory in discarded polypropylene masks has been designed, which has important practical significance.

[0022] This embodiment uses discarded polypropylene medical masks as the target substrate to simulate the process of ecological memory retention after colonization in different original aquatic environments and subsequent transfer to receiving aquatic environments. The original colonization environments include the flowing Wei River affected by sewage and the still artificial lake Xiaoxi Lake; the receiving water body is preferably an oligotrophic receiving water body, which in this embodiment is an ultrapure water system. The specific implementation steps are as follows: like Figure 1 As shown, the first aspect of this invention is to provide a method for screening species that retain ecological memory from discarded polypropylene masks, comprising the following steps: Step S1: Obtain several baseline samples and several transfer samples from the biofilm of discarded polypropylene masks; extract microbial DNA from all baseline and transfer samples, and perform high-throughput sequencing, data quality control, and standardization to obtain a relative abundance matrix; the rows in the relative abundance matrix represent various microorganisms, and the columns represent all baseline and transfer samples; based on the relative abundance matrix, obtain the overall relative abundance of each microorganism, then extract the microorganisms with the highest overall relative abundance, and construct an independent variable feature matrix with all transfer samples; the rows in the independent variable feature matrix represent various extracted microorganisms, and the columns represent all transfer samples.

[0023] It should be noted that by obtaining baseline samples colonized in the original water body and samples of the transferred group after cross-water transfer, a comparative basis of "control-change" is established. At the same time, the raw abundance matrix obtained by high-throughput sequencing is subjected to quality control and standardization processing to remove invalid or non-target sequences, eliminate sequencing depth differences, and extract the top few high-abundance groups as feature matrices, thereby reducing data noise and dimensionality, and providing high-quality and comparable input data for subsequent quantitative calculation of ecological memory retention and construction of machine learning models.

[0024] Specifically, waste polypropylene mask biofilm samples colonized in the original natural water environment were obtained as the baseline group; waste polypropylene mask biofilm samples transferred from the original natural water body to the receiving water body and incubated for different times were obtained as the transfer group; thus, several baseline group samples and several transfer group samples were obtained.

[0025] The process of obtaining the baseline and transfer group samples is illustrated as follows: First, microbial biofilms from the surface of discarded polypropylene masks are colonized in different original natural water bodies. After one week, these biofilms are transferred to a receiving water body for incubation. Samples are collected after two weeks, four weeks, and six weeks of incubation, which are the transfer group samples at different time points. Simultaneously, baseline samples from the original natural water body are collected at each time point. The number of transfer group samples is not specifically limited and can be determined by the implementer based on specific circumstances. The baseline and transfer group samples collected at the same time point are considered to be from the same source.

[0026] In this embodiment, the number of samples collected from the original natural water body is the baseline sample, and the number of baseline and transfer samples corresponding to each original natural water body at each time point is three; however, no specific limitation is made in this embodiment, and the implementer can determine it according to the specific situation.

[0027] Microbial DNA was extracted from each sample and high-throughput sequencing (such as 16S rRNA gene sequencing) was performed to obtain the abundance matrix of amplicon sequence variants (ASV) or operational taxonomic units (OTU), which is denoted as the original abundance matrix.

[0028] Among them, the ASV / OTU abundance matrix represents two abundance matrices with different resolutions: the ASV abundance matrix (100% similarity division) and the OTU abundance matrix (97% similarity division).

[0029] High-throughput sequencing is an advanced sequencing technology that can read hundreds of thousands to millions of DNA sequences in parallel at once, enabling rapid and low-cost acquisition of the genetic information of all microorganisms in a sample.

[0030] Amplicon sequence variants are unique DNA tags that are derived from sequencing sequences after rigorous noise reduction and are classified according to 100% sequence similarity. They can distinguish microbial genetic units with single-base differences and have a higher resolution than traditional OTUs.

[0031] Operational taxonomic units are traditionally grouped into a class based on a 97% sequence similarity threshold. Each class is considered an "operational taxonomic unit," representing a hypothetical microbial taxonomic unit.

[0032] The original abundance matrix is ​​a tabular data structure where rows represent various microorganisms, columns represent all baseline and transition group samples, and the values ​​in the table represent the number of sequences (i.e., abundance) of each microorganism detected in each sample. This matrix is ​​the foundational data format for all subsequent microbial community analyses.

[0033] The original abundance matrix undergoes data quality control to obtain a quality-controlled abundance matrix. This quality-controlled abundance matrix is ​​then standardized into a relative abundance matrix to eliminate differences in sequencing depth between different samples. Based on the relative abundance matrix, the overall relative abundance of each microorganism is obtained. Microorganisms with the highest overall relative abundance are then extracted and used to construct an independent variable feature matrix with all transfer group samples. In this independent variable feature matrix, rows represent the extracted microorganisms, and columns represent all transfer group samples. In this embodiment, the number of top-ranking microorganisms is 500; however, this is not specifically limited and can be determined by the implementer based on specific circumstances. The overall relative abundance of each microorganism is the sum of the relative abundances of all baseline and transfer group samples corresponding to each microorganism in the relative abundance matrix. A schematic diagram illustrating the process of screening microorganisms by summing relative abundances is shown below. Figure 3 As shown.

[0034] The original abundance matrix was subjected to data quality control to obtain a quality-controlled abundance matrix, which included removing sequences not classified as bacterial domains; removing non-target sequences classified as mitochondria or chloroplasts; and removing species with a total abundance of zero in all samples.

[0035] Step S2: Obtain the ecological memory retention of each transfer group sample by the difference in relative abundance between each transfer group sample and all corresponding homologous baseline group samples in the feature matrix of independent variables; construct the target variable data table with the transfer group sample number and ecological memory retention as column variables.

[0036] It should be noted that by quantifying the degree of dissimilarity in community structure between the transferred group samples and the original baseline group samples, the abstract phenomenon of "ecological memory" is transformed into a calculable continuous indicator (ecological memory retention). This indicator can quantitatively describe the degree to which the biofilm microbial community on the surface of discarded polypropylene masks retains the characteristics of its original environment after being transferred across water bodies. This provides a clear and regressible target variable for subsequent machine learning models, making it possible to screen key species related to memory retention from high-dimensional microbial data.

[0037] Specifically, the ecological memory retention of each transfer group sample is obtained by measuring the relative abundance difference between each transfer group sample and all corresponding homologous baseline group samples; wherein, the ecological memory retention of each transfer group sample includes: The relative abundance of all microorganisms in each column of the relative abundance matrix is ​​used to form a relative abundance sequence; the BC distance between any two columns is calculated using the Bray-Curtis algorithm based on the relative abundance sequences between any two columns. For each transfer group sample, identify its corresponding homologous baseline group sample, and calculate the mean of the BC distance between each transfer group sample and all corresponding homologous baseline group samples; use this as the average distance of each transfer group sample; and take the difference between 1 and the average distance of each transfer group sample as the ecological memory retention of each transfer group sample. The BC distance is calculated using the Bray-Curtis algorithm, which is a well-known technique and will not be described in detail here.

[0038] Thus, the ecological memory retention rate of each transfer group sample was obtained.

[0039] Using the sample number of the transfer group and the degree of ecological memory retention as column variables, construct a target variable data table (or dataset).

[0040] Step S3: Train a random forest regression model using the independent variable feature matrix and the target variable data table to obtain the trained random forest regression model; randomly shuffle all relative abundances of each microorganism in the independent variable feature matrix, and then input them, along with the relative abundance data from the independent variable feature matrix, into the trained random forest regression model to calculate the error; screen out key microorganisms based on the error calculation results.

[0041] It should be noted that a mapping relationship is established between the high-dimensional feature matrix of independent variables (independent variables) and the degree of ecological memory retention (dependent variable). The random forest regression model can effectively handle high-dimensional, nonlinear, and collinear microbial data, and improve prediction accuracy through ensemble learning. At the same time, the model's built-in feature importance assessment function can automatically calculate the contribution of each microbial group (each microorganism) to the prediction of ecological memory retention, thereby initially screening potential key predictive factors from massive amounts of microorganisms and providing a quantitative basis for subsequent screening and species identification.

[0042] It should be further explained that, since the behavior of the independent variable feature matrix is ​​microorganism, and the sample is listed, in order to achieve the concatenation of the independent variable feature matrix and the target variable data table, the independent variable feature matrix is ​​transposed to achieve the concatenation of the two.

[0043] Specifically, the feature matrix of the independent variables is transposed and then horizontally concatenated with the target variable data table to obtain the machine learning dataset. In the merged machine learning dataset, each row corresponds to a transfer group sample, and the columns include the relative abundance of microorganisms and the retention of ecological memory.

[0044] A random forest regression model was trained using a machine learning dataset to obtain the trained random forest regression model. The independent variable in the random forest regression model is the relative abundance of each microorganism; the dependent variable is the ecological memory retention of each transfer group sample.

[0045] A random seed is set before training to ensure that the results are repeatable. The number of decision trees is 1000.

[0046] It should be noted that, to determine the impact of each microbial species on the random forest regression model, the abundance value of a particular microorganism was intentionally randomly matched with the sample label to see if the model's predictions would worsen. If the predictions worsened significantly, it indicates that the microorganism was originally important; if there was no significant change, it suggests that it was dispensable.

[0047] Specifically, the feature matrix of the independent variables is input into the trained random forest regression model to predict the ecological memory retention. The mean square error is calculated by comparing the predicted ecological memory retention with the ecological memory retention in the target variable data table, and is denoted as the original mean square error. The relative abundance of all samples from the transition groups corresponding to a microorganism in the independent variable feature matrix is ​​randomly shuffled to obtain a new independent variable feature matrix for each microorganism. The new independent variable feature matrix for each microorganism is then input into the trained random forest regression model to predict the ecological memory retention. The mean squared error is calculated by comparing the predicted ecological memory retention with the ecological memory retention in the target variable data table, and is denoted as the shuffle mean squared error. Based on the difference between the original mean squared error and the shuffled mean squared error, the percentage increase in mean squared error for each microorganism is obtained; the specific formula for the percentage increase in mean squared error for each microorganism is as follows:

[0048] In the formula, This represents the original mean square error. This indicates the scrambled mean square error. This represents the percentage increase in mean squared error for each microorganism.

[0049] The calculation process of the mean square error is a well-known technique and will not be described in detail here.

[0050] Key microorganisms are selected based on the percentage increase in mean square error of all microorganisms. The specific process for selecting key microorganisms is as follows: the maximum curvature method is used to detect the inflection point of the microbial sequences arranged in descending order of the percentage increase in mean square error. The percentage increase in mean square error at the inflection point is used as the percentage threshold, and microorganisms with a percentage increase in mean square error greater than or equal to the percentage threshold are selected as key microorganisms.

[0051] The maximum curvature method is a well-known technique and will not be elaborated upon here. It should be noted that at least 10 key microorganisms must be screened; if fewer than 10 are selected, the top 10 microorganisms will be considered as key microorganisms.

[0052] Thus, the key microorganism was obtained.

[0053] Step S4: Screen for species that retain ecological memory from key microorganisms by analyzing the relationship between the relative abundance and ecological memory retention of each key microorganism in each transfer group sample.

[0054] It should be noted that the key microorganisms screened only indicate that these microorganisms have a high predictive contribution to the retention of ecological memory, but cannot distinguish whether the contribution is positive or negative. That is, an increase in the abundance of a certain microorganism may lead to an increase in the retention of ecological memory (positive correlation), or it may lead to a decrease (negative correlation), and the ecological significance of the two directions is completely different. Therefore, it is necessary to further calculate the Spearman correlation coefficient between the abundance of each key predictor and the retention of ecological memory, and to determine the direction based on the sign of the correlation coefficient. Positively correlated species are screened as species that retain ecological memory (reflecting the original environmental memory), and negatively correlated species are screened as species that disrupt ecological memory (reflecting community reconstruction or invasive alien species), thereby improving the ecological interpretability of the screening results.

[0055] Specifically, the relative abundance and ecological memory retention of all key microorganisms in the machine learning dataset for each transfer group sample are extracted; the Spearman correlation coefficient between the relative abundance and ecological memory retention of each key microorganism in each transfer group sample is calculated, denoted as the Spearman correlation coefficient for each key microorganism; the calculation process of the Spearman correlation coefficient is a well-known technique and will not be elaborated here. A schematic diagram of the Spearman correlation coefficient between relative abundance and ecological memory retention is shown below. Figure 4 As shown.

[0056] Key microorganisms with a Spearman correlation coefficient greater than 0 were classified as species preserving ecological memory; those with a Spearman correlation coefficient less than 0 were classified as species disrupting ecological memory. The screening results for species disrupting ecological memory are illustrated in the diagram below. Figure 5 As shown.

[0057] Thus, species that preserve ecological memory and species that destroy ecological memory have been identified.

[0058] Step S5: Obtain the microbial co-occurrence network, identify key species in the network through the co-occurrence network; identify key species in the neutral community by the relationship between the average relative abundance of all transfer group samples of each microorganism and the frequency of transfer group samples with relative abundance greater than 0; then, the intersection species of ecological memory-maintaining species, key species in the network, and key species in the neutral community are used as the final super core species for screening.

[0059] It is important to note that the selected species with ecological memory retention are based on statistical correlation, but correlation does not equate to ecological importance or certainty. Relying solely on correlation for screening may include species with accidental co-occurrence or indirect associations, lacking validation of the species' central role in the network topology and deterministic selection during community assembly. Therefore, co-occurrence network analysis is introduced to identify species playing a key pivotal role in the microbial interaction network, while a Sloan neutral community model is used to screen for species dominated by deterministic selection (rather than random drift). The intersection of these three methods is then used for cross-validation. Only microorganisms that simultaneously meet the three criteria of being "positively correlated with memory retention," "key species in the network," and "species selected by deterministic selection" are identified as super-core species, thus significantly improving the ecological reliability, robustness, and biological significance of the screening results.

[0060] Step S51: Co-occurrence network analysis to screen key species in the network.

[0061] It should be noted that, in order to independently verify the ecological memory-preserving species screened in the above steps from the perspective of the topological structure of the microbial ecological network, and to determine whether these species are truly in a core position in the community (such as intra-module hubs or inter-module connectors), rather than false positive species that are only statistically correlated with memory retention but have marginalized ecological functions, multi-dimensional cross-validation is used to enhance the ecological reliability and robustness of the screening results.

[0062] Specifically, the significance test of the Spearman correlation coefficient is used to determine whether there is a co-occurrence relationship between any two microorganisms; by treating all microorganisms as nodes and co-occurrence relationships as edges, a microbial co-occurrence network is constructed. The significance test of the Spearman correlation coefficient is a well-known technique and will not be elaborated upon here.

[0063] The microbial co-occurrence network is automatically divided into several modules (a module is a subgroup of nodes with tight connections within the network and sparse connections between modules) using a module detection algorithm; the average connection strength between each node and all other connected nodes within the same module is calculated. Calculate the average connection strength between each node and all connected nodes in other modules. The connection strength between two nodes refers to the absolute value of the Spearman correlation coefficient of the relative abundance of all translocation samples of the two microorganisms. In this embodiment, the module detection algorithm is implemented using either the Louvain algorithm or the Shared Minimum Distance (SMD) algorithm from SCNIC. SCNIC (Sparse Cooccurrence Network Investigation for Compositional data) is a software tool specifically designed for analyzing compositional data such as microbiomes.

[0064] Filter out Greater than the first preset threshold, and Microorganisms corresponding to nodes with values ​​greater than a second preset threshold are designated as key species in the network. In this embodiment, the first preset threshold is 0.8 and the second preset threshold is 0.6. However, in this embodiment, the first and second preset thresholds are not specifically limited and can be determined by the implementer according to specific circumstances.

[0065] Step S52: Fitting and screening key species for neutral community models.

[0066] It should be noted that, in order to independently validate the ecological memory-retaining species screened in the above steps from the perspective of microbial community assembly mechanisms, the Sloan neutral community model can distinguish whether microbial colonization in receiving water bodies is dominated by stochastic processes or deterministic processes (such as environmental selection and competitive exclusion). Using this model, species with a significantly higher occurrence frequency than the neutral prediction interval can be screened. The distribution of these species is controlled by deterministic mechanisms, making them more ecologically important. Cross-validating these deterministically selected species with ecological memory-retaining species ensures that the screened species are not only statistically positively correlated with memory retention but also driven by non-stochastic deterministic mechanisms during community assembly, thereby eliminating false positives caused by random drift and improving the ecological reliability of the screening results.

[0067] Specifically, the average relative abundance and observed frequency of each microorganism in all transfer group samples are obtained through the feature matrix of independent variables; wherein, the observed frequency is the ratio of the frequency of each microorganism in all transfer group samples to the total number of transfer group samples; wherein, the frequency of each microorganism in all transfer group samples is the number of transfer group samples with a relative abundance greater than 0.

[0068] Based on the average relative abundance and observed occurrence frequency of all transgenic samples of all microorganisms, the Sloan Neutral Community Model was fitted using the least squares method to obtain the fitted Sloan Neutral Community Model, and the 95% confidence interval for the occurrence frequency was determined. The average relative abundance of all transgenic samples of each microorganism in the feature matrix of independent variables was input into the fitted Sloan Neutral Community Model, and the predicted occurrence frequency was output. The predicted occurrence frequency was compared with the upper bound of the 95% confidence interval; if it was higher than the upper bound, it was identified as a key species in the neutral community. The screening process for key species in the neutral community is illustrated in the diagram below. Figure 6 As shown.

[0069] Step S53: Conduct final screening of species.

[0070] The intersection of ecological memory-preserving species, network key species, and neutral community key species was used as the final selection of super core species.

[0071] Thus, the final selection of super core species was obtained.

[0072] This concludes the embodiment.

[0073] like Figure 2 As shown, a second aspect of the present invention is to provide a screening system for the preservation of ecological memory species in discarded polypropylene masks, comprising: Data acquisition and preprocessing module 101: used to acquire several baseline group samples and several transfer group samples of biofilm from discarded polypropylene masks; extract microbial DNA from all baseline group samples and transfer group samples, and perform high-throughput sequencing, data quality control and standardization to obtain a relative abundance matrix; based on the relative abundance matrix, obtain the overall relative abundance of each microorganism, then extract the microorganisms with the highest overall relative abundance, and construct an independent variable feature matrix with all transfer group samples; Ecological memory retention calculation module 102: used to obtain the ecological memory retention of each transfer group sample by the difference in relative abundance between each transfer group sample and all corresponding homologous reference group samples in the independent variable feature matrix; and to construct a target variable data table with the transfer group sample number and ecological memory retention as column variables. Key microorganism extraction module 103: used to train a random forest regression model using the independent variable feature matrix and the target variable data table, and obtain the trained random forest regression model; randomly shuffle all the relative abundance of each microorganism in the independent variable feature matrix, and then input them and the relative abundance data of the independent variable feature matrix into the trained random forest regression model to calculate the error; and screen out key microorganisms based on the error calculation results. Ecological memory retention species screening module 104: used to screen ecological memory retention species from key microorganisms by the relationship between the relative abundance of each key microorganism in each transfer group sample and the degree of ecological memory retention; Multi-dimensional joint screening module 105: used to obtain the microbial co-occurrence network, identify key species in the network through the co-occurrence network; identify key species in the neutral community by the relationship between the average relative abundance of all transfer group samples of each microorganism and the frequency of transfer group samples with relative abundance greater than 0; and then use the intersection species of ecological memory-maintaining species, key species in the network and key species in the neutral community as the super core species for final screening.

[0074] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, optical storage, etc.) containing computer-usable program code.

[0075] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, systems, and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0076] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0077] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0078] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the present invention.

Claims

1. A method for screening species that retain ecological memory from discarded polypropylene masks, characterized in that, include: Several baseline samples and several transfer samples of biofilm from discarded polypropylene masks were obtained; Microbial DNA was extracted from all baseline and transfer samples and subjected to high-throughput sequencing, data quality control, and standardization to obtain a relative abundance matrix; the rows in the relative abundance matrix represent various microorganisms, and the columns represent all baseline and transfer samples. Based on the relative abundance matrix, the overall relative abundance of each microorganism is obtained. Then, the microorganisms with the highest overall relative abundance are extracted and an independent variable feature matrix is ​​constructed with all translocation group samples. The rows in the independent variable feature matrix represent the extracted microorganisms, and the columns represent all translocation group samples. The ecological memory retention of each transfer group sample is obtained by comparing the relative abundance differences between each transfer group sample and all corresponding homologous baseline group samples in the feature matrix of independent variables. A target variable data table was constructed using the transfer group sample number and ecological memory retention as column variables; A random forest regression model is trained using the feature matrix of independent variables and the data table of target variables to obtain the trained random forest regression model. The relative abundance of each microorganism in the feature matrix of independent variables is randomly shuffled and then input into the trained random forest regression model along with the relative abundance data of the feature matrix of independent variables to calculate the error. Key microorganisms are selected based on the error calculation results. By analyzing the relationship between the relative abundance and ecological memory retention of each key microorganism in samples from different transfer groups, species that retain ecological memory were screened from the key microorganisms. A microbial co-occurrence network was obtained, and key species of the network were identified through the co-occurrence network. Key species of the neutral community were identified by the relationship between the average relative abundance of all transfer group samples of each microorganism and the frequency of transfer group samples with relative abundance greater than 0. Then, the intersection species of ecological memory-maintaining species, key species of the network, and key species of the neutral community were selected as the super core species for final screening.

2. The method for screening species that preserve the ecological memory of discarded polypropylene masks according to claim 1, characterized in that, The process of obtaining the relative abundance matrix includes: Microbial DNA was extracted from all baseline and transfer samples and subjected to high-throughput sequencing to obtain the original abundance matrix. The original abundance matrix was used to identify different microorganisms and listed as all baseline and transfer samples. Data quality control and standardization were performed on the original abundance matrix to obtain the relative abundance matrix. The data quality control includes removing sequences not classified as bacterial domains; removing non-target sequences classified as mitochondria or chloroplasts; and removing species with a total abundance of zero in all samples.

3. The method for screening species that preserve the ecological memory of discarded polypropylene masks according to claim 1, characterized in that, The step of obtaining the overall relative abundance of each microorganism based on the relative abundance matrix includes: The overall relative abundance of each microorganism is the sum of the relative abundances of all baseline and transfer samples for each microorganism in the relative abundance matrix.

4. The method for screening species that preserve the ecological memory of discarded polypropylene masks according to claim 1, characterized in that, The ecological memory retention of each transfer group sample includes: The relative abundance of all microorganisms in each column of the relative abundance matrix is ​​used to form a relative abundance sequence; the BC distance between any two columns is calculated using the Bray-Curtis algorithm based on the relative abundance sequences between any two columns. For each transfer group sample, identify its corresponding homologous baseline group sample, and calculate the mean BC distance between each transfer group sample and all corresponding homologous baseline group samples; use this as the average distance of each transfer group sample; and use the difference between 1 and the average distance of each transfer group sample as the ecological memory retention of each transfer group sample.

5. The method for screening species that preserve the ecological memory of discarded polypropylene masks according to claim 1, characterized in that, The step of randomly shuffling all relative abundances of each microorganism in the independent variable feature matrix, and then inputting them, along with the relative abundance data from the independent variable feature matrix, into the trained random forest regression model for error calculation includes: The feature matrix of the independent variables is input into the trained random forest regression model to predict the ecological memory retention. The mean square error is calculated by comparing the predicted ecological memory retention with the ecological memory retention in the target variable data table, and is denoted as the original mean square error. The relative abundance of all samples from the transition groups corresponding to a microorganism in the independent variable feature matrix is ​​randomly shuffled to obtain a new independent variable feature matrix for each microorganism. The new independent variable feature matrix for each microorganism is then input into the trained random forest regression model to predict the ecological memory retention. The mean squared error is calculated by comparing the predicted ecological memory retention with the ecological memory retention in the target variable data table, and is denoted as the shuffle mean squared error. Based on the difference between the original mean squared error and the shuffled mean squared error, the percentage increase in mean squared error for each microorganism is obtained, expressed by the formula: In the formula, This represents the original mean square error. This indicates the scrambled mean square error. This represents the percentage increase in mean squared error for each microorganism.

6. The method for screening species that preserve the ecological memory of discarded polypropylene masks according to claim 5, characterized in that, The screening process for the key microorganisms includes: Key microorganisms are selected based on the percentage increase in mean square error of all microorganisms. The specific process for selecting key microorganisms is as follows: the maximum curvature method is used to detect the inflection point of the microbial sequences arranged in descending order of the percentage increase in mean square error. The percentage increase in mean square error at the inflection point is used as the percentage threshold, and microorganisms with a percentage increase in mean square error greater than or equal to the percentage threshold are selected as key microorganisms.

7. The method for screening species that preserve the ecological memory of discarded polypropylene masks according to claim 1, characterized in that, The selection process for species that preserve ecological memory includes: The independent variable feature matrix is ​​transposed and then horizontally concatenated with the target variable data table to obtain the machine learning dataset; each row in the merged machine learning dataset corresponds to a translocated group sample, and the columns include the relative abundance of microorganisms and the retention of ecological memory. The Spearman correlation coefficient between the relative abundance and ecological memory retention of each key microorganism in each transfer group sample was calculated and denoted as the Spearman correlation coefficient of each key microorganism. Key microorganisms with a Spearman correlation coefficient greater than 0 were identified as species that retain ecological memory.

8. The method for screening species that preserve the ecological memory of discarded polypropylene masks according to claim 1, characterized in that, The acquisition of the microbial co-occurrence network, and the identification of key species in the network through the co-occurrence network, includes: The significance of Spearman correlation coefficient is used to determine whether there is a co-occurrence relationship between any two microorganisms; all microorganisms are treated as nodes and co-occurrence relationships as edges, thus forming a microbial co-occurrence network; The microbial co-occurrence network is automatically divided into several modules using a module detection algorithm; the average connection strength between each node and all other connected nodes within the same module is calculated. Calculate the average connection strength between each node and all connected nodes in other modules. The connection strength between two nodes refers to the absolute value of the Spearman correlation coefficient of the relative abundance of all metastasis samples of the two microorganisms. Filter out Greater than the first preset threshold, and Microorganisms corresponding to nodes that exceed a second preset threshold are designated as key species in the network.

9. The method for screening species that preserve the ecological memory of discarded polypropylene masks according to claim 1, characterized in that, The screening process for key species in the neutral community includes: The average relative abundance and observed frequency of each microorganism in all transfer group samples were obtained by using the feature matrix of independent variables; where the observed frequency is the ratio of the frequency of each microorganism in all transfer group samples to the total number of transfer group samples; where the frequency of each microorganism in all transfer group samples is the number of transfer group samples with a relative abundance greater than 0. Based on the average relative abundance and observed occurrence frequency of all transgenic samples of all microorganisms, a Sloan neutral community model was fitted using the least squares method to obtain the fitted Sloan neutral community model, and the 95% confidence interval of the occurrence frequency was determined. The average relative abundance of all transgenic samples of each microorganism in the feature matrix of independent variables was input into the fitted Sloan neutral community model, and the predicted occurrence frequency was output. The predicted occurrence frequency was compared with the upper bound of the 95% confidence interval. If it was higher than the upper bound, it was determined to be a key species in the neutral community.

10. A screening system for species preserving ecological memory in discarded polypropylene masks, used to implement the screening method for species preserving ecological memory in discarded polypropylene masks as described in any one of claims 1-9, characterized in that, include: Data acquisition and preprocessing module: used to acquire several baseline group samples and several transfer group samples of biofilm from discarded polypropylene masks; Microbial DNA was extracted from all baseline and transgenic samples and subjected to high-throughput sequencing, data quality control, and standardization to obtain a relative abundance matrix. Based on the relative abundance matrix, the overall relative abundance of each microorganism was obtained. Then, the microorganisms with the highest overall relative abundance were extracted and used to construct an independent variable feature matrix with all transgenic samples. Ecological memory retention calculation module: used to obtain the ecological memory retention of each transfer group sample by the difference in relative abundance between each transfer group sample in the independent variable feature matrix and all corresponding homologous baseline group samples; A target variable data table was constructed using the transfer group sample number and ecological memory retention as column variables; Key microorganism extraction module: This module trains a random forest regression model using the independent variable feature matrix and the target variable data table to obtain the trained random forest regression model. It randomly shuffles all relative abundances of each microorganism in the independent variable feature matrix and inputs them, along with the relative abundance data from the independent variable feature matrix, into the trained random forest regression model to calculate the error. Key microorganisms are then selected based on the error calculation results. Ecological memory retention species screening module: used to screen ecological memory retention species from key microorganisms by analyzing the relationship between the relative abundance of each key microorganism in each transfer group sample and the degree of ecological memory retention; Multi-dimensional joint screening module: used to obtain the microbial co-occurrence network, and to identify key species in the network through the co-occurrence network; to identify key species in the neutral community by the relationship between the average relative abundance of all transfer group samples of each microorganism and the frequency of transfer group samples with relative abundance greater than 0; and then to select the super core species as the intersection of ecological memory-maintaining species, key species in the network and key species in the neutral community.