Mangrove forest habitat core fish community prediction method

By constructing an ecological network model and using eLSA dynamic programming algorithm, combining ecological network analysis and extended local similarity analysis, the problem that the existing technology cannot comprehensively analyze the dynamics of mangrove core fish communities is solved, and accurate identification of core fish communities and effective prediction of the future dynamics of the ecosystem are achieved.

CN120046783AInactive Publication Date: 2025-05-27SOUTH CHINA AGRICULTURAL UNIVERSITY

Patent Information

Application Number
CN202510116556.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology cannot comprehensively and accurately grasp the dynamic changes of core fish communities in mangrove habitats, and lacks methods to integrate ecological networks and correlation analysis.

Method used

A method of predicting core fish community in mangrove habitats was adopted to obtain fish eDNA data, build ecological network models, calculate characteristic indicators, perform standardized processing, comprehensive scoring, and sort important species. The eLSA dynamic programming algorithm was used to calculate the highest similarity score between species, and build an association network to determine core fish populations.

Benefits of technology

A comprehensive analysis of the dynamics of core fish communities has been achieved, and the core fish populations that are key to ecosystem functions have been accurately identified, which has enhanced the ability to predict the future dynamics of the ecosystem.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046783A_ABST
    Figure CN120046783A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of ecological environment science, in particular to a mangrove forest habitat core fish community prediction method, which comprises the following steps: acquiring fish eDNA data of a mangrove forest water area in different periods and performing data analysis to obtain composition structure data of a fish school; determining a food web; constructing an ecological network model; calculating a characteristic index of each species node in the ecological network model; carrying out standardization processing on the characteristic indexes; calculating a comprehensive score of the standardized feature indexes; important species and non-important species are distinguished according to the comprehensive score; important species data in different periods are preprocessed; calculating the highest similarity score between any two important species; evaluating the statistical significance; constructing an association network by using a species association result with statistical significance; and analyzing the association network to determine a core fish school. According to the method, the ecological network and correlation analysis are integrated, comprehensive analysis of the core fish community dynamics is realized, and the core fish school of the ecological system is accurately identified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of ecological environment science, and more specifically, to a method for predicting the core fish community in mangrove habitats. Background Art

[0002] Mangroves are important ecosystems in tropical and subtropical coastal areas, providing habitats, breeding grounds, and food resources for various organisms such as fish. The ecological service functions of mangroves are significant, including protecting the coast, reducing pollution, and maintaining biodiversity. In mangrove habitats, the core fish community is an important part of the ecosystem's energy flow and material cycle. These fish often have ecological indicator significance, and their dynamic changes can reflect the health status of the habitat. Usually, ENA (Ecological Network Analysis) and eLSA (Extended Local Similarity Analysis) are used for ecosystem research.

[0003] ENA (Ecological Network Analysis) is a method widely used in ecosystem research, mainly for analyzing the material and energy flow relationships between different components in an ecosystem. By establishing an ecological network model, the structural characteristics and dynamic behaviors of the ecosystem can be quantified. Specifically, ENA can identify key species, functional groups, and energy flow paths in the food web, providing theoretical support for ecosystem management and protection. However, traditional ENA has limitations. It is usually based on a static ecological network structure and is difficult to effectively reflect the spatio-temporal dynamic characteristics of the ecosystem. Moreover, traditional ENA focuses on direct energy and material flows and fails to fully consider indirect interactions and their impact on the ecosystem structure.

[0004] eLSA (Extended Local Similarity Analysis) is a method for exploring the dynamic associations between species in an ecosystem, especially suitable for dealing with time series data. Based on the principle of local similarity, eLSA can reveal the non-linear, asymmetric, and time-delay association relationships between species in a complex ecosystem by calculating the correlations and delay effects between species. However, eLSA mainly analyzes the statistical correlations between species and is difficult to directly infer causal relationships. It has high requirements for the quality and length of time series data and may be significantly affected by noise and missing data. Moreover, eLSA is independent of the overall framework of the ecological network and is difficult to be applied synergistically with other ecological analysis methods.

[0005] Although the existing ecological network analysis method and extended local similarity analysis method have achieved certain results in their respective fields, when applied separately to predict the core fish community in mangrove habitats, neither of them can comprehensively and accurately grasp the dynamic change law of the fish community. The two lack effective integration and fail to fully incorporate the special ecological environment factors of mangroves.

[0006] The prior art discloses a method for screening functional indicator species of fish in a river ecosystem. Sampling sites are determined in the river ecosystem, and fish organisms at the sampling sites are collected and identified. Grouping is carried out according to environmental types: sampling sites with similar habitats are grouped together. Screening of indicator species: Calculate the average abundance and relative occurrence frequency of each fish in each group, and then obtain the index value of each fish within each group based on the average abundance and relative occurrence frequency. Select the species with the largest index value in each group, and select the fish with an index value > 55% as indicator species. Determination of indicator species: Judge the selected species in terms of their species habits, habitats, and whether they are exotic species to establish indicator species. However, although this scheme can screen indicator species, it does not integrate ecological networks and correlation analysis and cannot comprehensively analyze the dynamics of indicator species. Summary of the Invention

[0007] The object of the present invention is to overcome the deficiency of the prior art in comprehensively analyzing the dynamics of indicator species, and provide a method for predicting the core fish community in mangrove habitats, which integrates ecological networks and correlation analysis to achieve a comprehensive analysis of the dynamics of the core fish community.

[0008] To solve the above technical problems, the technical solution adopted by the present invention is:

[0009] Provide a method for predicting the core fish community in mangrove habitats, including the following steps:

[0010] S1: Obtain the fish eDNA data in mangrove waters at different times and conduct data analysis to obtain the composition structure data of the fish population;

[0011] S2: Collect the feeding data of each species in the fish population, determine the functional positioning of the species and the ecological role of the species in energy flow, and determine the food web;

[0012] S3: Construct an ecological network model by combining biomass and the food web;

[0013] S4: Calculate the characteristic indicators of each species node in the ecological network model;

[0014] S5: Standardize the characteristic indicators;

[0015] S6: Calculate the comprehensive score of the standardized characteristic indicators;

[0016] S7: Sort the species according to the comprehensive score, compare the comprehensive score with the set threshold value to distinguish important species and unimportant species;

[0017] S8: Preprocess the data of important species in different periods;

[0018] S9: Use the eLSA dynamic programming algorithm to calculate the highest similarity score between any two important species;

[0019] S10: Evaluate statistical significance;

[0020] S11: Use the species association results with statistical significance to construct an association network;

[0021] S12: Analyze the association network, calculate the network metrics of important species, and determine the core fish stocks based on the time-lag pattern.

[0022] The method for predicting the core fish community in the mangrove habitat of the present invention includes ecological network analysis and extended local similarity analysis. Among them, ecological network analysis distinguishes important species and unimportant species by constructing an ecological network model and calculating model metrics; extended local similarity analysis further analyzes important species to construct an association network to determine the core fish stocks. The present invention integrates ecological network and association analysis to achieve a comprehensive analysis of the dynamics of the core fish community and accurately identify the core fish stocks that are crucial for ecosystem functions.

[0023] Further, in step S3, the process of constructing the ecological network model is as follows:

[0024] S31: Construct an adjacency matrix according to the food web; the rows and columns of the adjacency matrix represent different fish species. When species i preys on species j, the matrix element a ij = 1, otherwise a ij = 0;

[0025] S32: Determine the flux matrix by combining the biomass and the predation relationship of the food web;

[0026] S33: Draw a topological network relationship diagram between fish species.

[0027] Further, in step S4, the characteristic metrics include degree, in-degree, out-degree, betweenness centrality, closeness centrality, information centrality, topological importance index, criticality index, upward critical index, downward critical index, dispersion, distance-weighted dispersion, and distance-weighted reachability.

[0028] Further, the calculation expression of the characteristic metrics is:

[0029] D i = D in,i + D out,i ;

[0030]

[0031] i ≠ k, i ≠ j and j < k;

[0032]

[0033] K(i) = K b (i) + K t (i);

[0034]

[0035]

[0036] Wherein, D i represents the degree of vertex of species i, which is the total number of prey and predators of species i; D in,i represents the in-degree, which is the number of prey of species i; D out,i represents the out-degree, which is the number of predators of species i;

[0037] BC i represents the betweenness centrality of species i, which is used to quantitatively analyze the control ability of species over information exchange within the community; CC i represents the closeness centrality of species i, which is used to quantitatively measure the degree of advantage of species in transmitting information; IC i represents the information centrality of species i, which is used to quantitatively measure the information transmission ability between any two connected species; N represents the number of species appearing in the survey; g jk represents the number of shortest paths existing between species j and species k; g jk(i) represents the number of shortcuts passing through species i between species j and species k; d ij represents the shortcut distance between species i and species j; I ij represents the reciprocal of the shortcut distance between species i and species j;

[0038] n represents the maximum step length of the analysis; represents the importance index of the influence of species i on the topological structure of the fish community after n steps; α m,j represents the influence of species i on species j when species i reaches species j after m steps; a m,ji represents the influence of species i on species j at the m-th step;

[0039] K b (i) represents the upward criticality index of species i; K b (j) represents the upward criticality index of species j; Species j is the direct predator of species i; e(i)(j) is the number of species of the direct prey of species j; Kt (i) is the downward criticality index of species i; K t (j) is the downward criticality index of species j; f(i)(j) is the number of species that are direct predators of species j; K(i) represents the criticality index of species i;

[0040] F represents the dispersion, taking values in (0, 1); s i represents the number of species of species i in the i-th discrete population; D F represents the distance-weighted dispersion, taking values in (0, 1); D R represents the distance-weighted accessibility, taking values in (0, 1); d Mj represents the distance from a series of points M to any point j.

[0041] Furthermore, in step S5, the process of standardizing the characteristic indicators is as follows:

[0042] First, check the indicator distribution. Analyze whether there are extreme values through a histogram or box plot, and check the mean, standard deviation, maximum, and minimum of the indicators; according to the type of data conformity, select the following methods to standardize the data:

[0043] If the data distribution is uneven, there are extreme values, but there is no skewed distribution, select Z-score to standardize the data:

[0044]

[0045] In the formula, x represents the original indicator value; x' represents the standardized indicator value; μ represents the mean of the indicator; σ represents the standard deviation of the indicator;

[0046] Standardize the indicator value to a standard normal distribution. The standardized data fluctuates around 0, and it can be intuitively seen the position of each species relative to the average value on each topological index;

[0047] If the data has a power-law distribution or extreme skewed distribution, and there are no negative or zero values, select log transformation:

[0048] x′ = log(x + 1);

[0049] If the data distribution is uniform without extreme values, and the physical meaning of the indicator is clear, select the Min-Max normalization method:

[0050]

[0051] In the formula, x min represents the minimum value of the indicator; x max represents the maximum value of the indicator.

[0052] Furthermore, in step S6, the calculation process of the comprehensive score is as follows:

[0053] Using the principal component analysis method, calculate the variance contribution rate and assign weights to each index;

[0054] Perform weighted summation on the standardized index values:

[0055]

[0056] In the formula, S represents the comprehensive score; q represents the total number of indexes; w r represents the weight of the r-th index; x' r represents the standardized value of the r-th index.

[0057] Furthermore, in step S8, the preprocessing includes data matrix input and processing of missing values, processing of duplicate data, and data standardization processing;

[0058] The process of data matrix input and processing of missing values is as follows: Organize the important species data at different times into a matrix form, where each row represents the data of a species factor at different time points, and organize the data into a time series matrix X:

[0059] X = X [1:T][1:C] ;

[0060] In the formula, T represents the number of time points; C represents the number of species types;

[0061] If there are missing values in the data, fill them according to the method provided by the eLSA software; if the zero-order spline method is selected, directly assign the missing position a value of 0; if the third-order spline method or the nearest neighbor method is selected, calculate the missing values according to the corresponding interpolation algorithm;

[0062] The process of processing duplicate data is as follows: Summarize the duplicate data using the function F(X i ), and the summary method used is the MAD weighted median method;

[0063]

[0064] MAD(X i ) = Median(|X i - Median(X i )|);

[0065] In the formula, F(X i ) represents the standardized summary result; MAD(X i ) represents the absolute median difference; Median(X i ) represents the median;

[0066] The process of data standardization processing is as follows: Apply the selected function to transform the important species data at each time point to obtain F(Xi ) Subsequently, normalization is achieved by using the rank-based normal score transformation method.

[0067] Further, in step S9, let the time series data of important species A and B and D is the set time delay limit; P represents the cumulative local similarity score matrix; L represents the cumulative similarity count matrix;

[0068] The calculation process of the highest similarity score between two important species is as follows:

[0069] S91: Initialization; P 0,ω = 0; L 0,ω = 0;

[0070] In the formula, P 0,ω represents the similarity score in the 0th row of the matrix; represents the similarity score in the 0th column of the matrix; L 0,ω represents the cumulative number of similarities in the 0th row of the matrix; represents the cumulative number of similarities in the 0th column of the matrix;

[0071] S92: Dynamic programming calculation:

[0072]

[0073] In the formula; represents the value of the cumulative local similarity score matrix at position ; represents the value of the cumulative local similarity count matrix at position ; represents the similarity metric function of species A and species B at time points and ω;

[0074] S93: Calculate the maximum score;

[0075]

[0076] In the formula, P max (A,B) is the maximum local similarity score of species A and B; L max (A,B) represents the maximum similarity count of species A and B;

[0077] S94: Calculate the final score:

[0078]

[0079] S sgn (A,B) =sgn [P max (A,B)-L max (A,B)];

[0080] In the formula, S max (A,B) represents the standardized maximum similarity score of species A and B; S sgn (A,B) represents the final LS score.

[0081] Furthermore, in step S10, the process of evaluating statistical significance is as follows:

[0082] Permutation test: Fix B and rearrange all columns of A multiple times. The number of permutations is H. For each permuted dataset A (k) , calculate S max (A (k) ,B):

[0083]

[0084] In the formula, Q represents the probability of the current local similarity score when the two species are not associated; I is the indicator function;

[0085] FDR correction: For the Q values of all species factor pairs, calculate according to the Storey method; assume there are E species factors in total, then there are eLSA pairs calculated; adjust the Q value of each species factor pair through the algorithm of the Storey method to control the false positive rate in multiple tests.

[0086] Furthermore, in step S11, the process of constructing the association network is as follows:

[0087] For two species A and B, if within the time intervals [s 1 , t 1 of A and [s 2 , t 2 of B, they have the highest LS score, and s 1 < s 2 , then draw an arrow from A to B in the network, indicating that A may lead B in time, that is, A may have an impact on B; if s 1 > s 2 , then the arrow direction is opposite, and the solid and dashed lines are used to represent the positive and negative of the association. The solid line represents positive correlation, and the dashed line represents negative correlation.

[0088] Compared with the background technology, the beneficial effects of the present invention are as follows:

[0089] 1. Integrate the ecological network and association analysis to achieve a comprehensive analysis of the dynamics of the core fish community and accurately identify the core fish groups that are key to the ecosystem function;

[0090] 2. Introduce time series analysis and time lag effect to capture the spatio-temporal variation characteristics of community structure;

[0091] 3. Based on a multi-level mathematical model framework, enhance the prediction ability of the future dynamics of the ecosystem. Description of the Drawings

[0092] Figure 1 It is a flowchart of the prediction method for the core fish community in the mangrove habitat in the embodiment of the present invention. Detailed Embodiments

[0093] The present invention will be further described below in conjunction with the detailed embodiments.

[0094] Embodiment 1

[0095] This embodiment is the first embodiment of the prediction method for the core fish community in the mangrove habitat. As Figure 1 shown, it includes the following steps:

[0096] S1: Obtain the fish eDNA data in the mangrove waters at different times and conduct data analysis to obtain the composition structure data of the fish population;

[0097] S2: Collect the feeding data of each species in the fish population, determine the functional positioning of the species and the ecological role of the species in the energy flow, and determine the food web;

[0098] S3: Construct an ecological network model by combining biomass and food web;

[0099] S4: Calculate the characteristic indexes of each species node in the ecological network model;

[0100] S5: Standardize the characteristic indexes;

[0101] S6: Calculate the comprehensive score of the standardized characteristic indexes;

[0102] S7: Rank the species according to the comprehensive score, compare the comprehensive score with the set threshold, and distinguish important species and non-important species;

[0103] S8: Preprocess the important species data at different times;

[0104] S9: Use the eLSA dynamic programming algorithm to calculate the highest similarity score between any two important species;

[0105] S10: Evaluate statistical significance;

[0106] S11: Construct an association network using the species association results with statistical significance;

[0107] S12: Analyze association networks, calculate network metrics for important species, and identify core fish populations based on time-lagged patterns.

[0108] The above-mentioned method for predicting the core fish community in mangrove habitats includes ecological network analysis and extended local similarity analysis, wherein the ecological network analysis distinguishes important species from non-important species by constructing an ecological network model and calculating model indicators; the extended local similarity analysis further analyzes important species and constructs an association network to determine the core fish community. In this embodiment, the ecological network and association analysis are integrated to achieve a comprehensive analysis of the dynamics of the core fish community and accurately identify the core fish community that is critical to the function of the ecosystem.

[0109] In step S3, the process of constructing the ecological network model is:

[0110] S31: Construct an adjacency matrix based on the food web; the rows and columns of the adjacency matrix represent different fish species. When species i preys on species j, the matrix element a ij =1, otherwise a ij =0;

[0111] S32: Determine the flow matrix by combining biomass and predator-prey relationships in the food web; the flow matrix represents the flow of energy or matter in the ecological network. If species i preys on species j, the flow can be calculated based on the ratio of prey to prey biomass;

[0112] S33: Draw a topological network relationship diagram among fish species.

[0113] In step S4, the characteristic indicators include point degree, point in-degree, point out-degree, betweenness centrality, closeness centrality, information centrality, topological importance index, criticality index, upstream criticality index, downstream criticality index, dispersion, distance weight dispersion and distance weight reachability.

[0114] The calculation expression of the characteristic index is:

[0115] D i =D in,i +D out,i ;

[0116]

[0117]

[0118] i≠k,i≠j and j <k;

[0119]

[0120] K(i)=K b (i)+K t (i);

[0121]

[0122] Wherein, D i represents the degree of species i, which is the total number of prey and predators of species i; D in,i represents the in-degree, which is the number of prey of species i; D out,i represents the out-degree, which is the number of predators of species i;

[0123] BC i represents the betweenness centrality of species i, which is used to quantitatively analyze the control ability of species over information exchange within the community; CC i represents the closeness centrality of species i, which is used to quantitatively measure the degree of advantage of species in transmitting information; IC i represents the information centrality of species i, which is used to quantitatively measure the information transmission ability between any two connected species; N represents the number of species appearing in the survey; g jk represents the number of shortest paths existing between species j and species k; g jk(i) represents the number of shortcuts passing through species i existing between species j and species k; d ij represents the shortcut distance between species i and species j; I ij represents the reciprocal of the shortcut distance between species i and species j;

[0124] n represents the maximum step length of the analysis; represents the importance index of the influence of species i on the topological structure of the fish community after n steps; α m,j represents the influence of species i on species j when species i reaches species j after m steps; a m,ji represents the influence of species i on species j at the mth step;

[0125] K b (i) represents the upward criticality index of species i; K b (j) represents the upward criticality index of species j; Species j is the direct predator of species i; e(i)(j) is the number of species of the direct prey of species j; K t (i) is the downward criticality index of species i; K t (j) is the downward criticality index of species j; f(i)(j) is the number of species of the direct predator of species j; K(i) represents the criticality index of species i;

[0126] F represents the dispersion, with a value range of (0, 1); s i represents the number of species of species i in the ith discrete population; D F represents the distance-weighted dispersion, with a value range of (0, 1); D R represents the distance-weighted reachability, with a value range of (0, 1); dMj Denote the distances from a series of points M to any point j.

[0127] In step S5, the process of standardizing the characteristic indicators is as follows:

[0128] First, check the indicator distribution, analyze whether there are extreme values through histograms or box plots, and examine the mean, standard deviation, maximum, and minimum values of the indicators; according to the type of data conformity, select the following methods to standardize the data:

[0129] If the data distribution is uneven, there are extreme values, but there is no skewed distribution, select Z-score to standardize the data:

[0130]

[0131] In the formula, x represents the original indicator value; x' represents the standardized indicator value; μ represents the mean of the indicator; σ represents the standard deviation of the indicator;

[0132] Standardize the indicator values to a standard normal distribution. The standardized data fluctuates around 0, and intuitively shows the position of each species relative to the average value on each topological index;

[0133] If the data has a power-law distribution or extreme skewed distribution, and there are no negative or zero values, select the log transformation:

[0134] x′ = log(x + 1);

[0135] If the data distribution is uniform without extreme values and the physical meaning of the indicator is clear, select the Min-Max normalization method:

[0136]

[0137] In the formula, x min represents the minimum value of the indicator; x max represents the maximum value of the indicator.

[0138] In step S6, the calculation process of the comprehensive score is as follows:

[0139] Use the principal component analysis method to calculate the variance contribution rate and assign weights to each indicator;

[0140] Perform weighted summation on the standardized indicator values:

[0141]

[0142] In the formula, S represents the comprehensive score; q represents the total number of indicators; w r represents the weight of the r-th indicator; x' r represents the standardized value of the r-th indicator.

[0143] Example 2

[0144] This example is the first example of the prediction method for the core fish community in the mangrove habitat. This example is similar to Example 1, except that in step S8, the preprocessing includes the processing of data matrix input and missing values, the processing of duplicate data, and the data standardization processing;

[0145] The process of data matrix input and missing value processing is as follows: Organize the important species data at different times into a matrix form, where each row represents the data of a species factor at different time points, and organize the data into a time series matrix X:

[0146] X = X [1:T][1:C] ;

[0147] In the formula, T represents the number of time points; C represents the number of species types;

[0148] If there are missing values in the data, fill them according to the method provided by the eLSA software; if the zero-order spline method is selected, directly assign the missing position as 0; if the third-order spline method or the nearest neighbor method is selected, calculate the missing values according to the corresponding interpolation algorithm;

[0149] The process of duplicate data processing is as follows: Summarize the duplicate data using the function F(X i ), and the summary method used is the MAD weighted median method;

[0150]

[0151] MAD(X i ) = Median(|X i - Median(X i )|);

[0152] In the formula, F(X i ) represents the standardized summary result; MAD(X i ) represents the absolute median difference; Median(X i ) represents the median;

[0153] The process of data standardization processing is as follows: Apply the selected function to transform the important species data at each time point to obtain F(X i ), and then use the rank-based normal score transformation method to achieve normalization.

[0154] In step S9, assume the time series data of important species A and B, and D is the set time delay limit; P represents the cumulative local similarity score matrix; L represents the cumulative similarity quantity matrix;

[0155] The calculation process of the highest similarity score between two important species is as follows:

[0156] S91: Initialization; P 0,ω = 0; L 0,ω = 0;

[0157] In the formula, P 0,ω represents the similarity score in the 0th row of the matrix; represents the similarity score in the 0th column of the matrix; L 0,ω represents the cumulative number of similarities in the 0th row of the matrix; represents the cumulative number of similarities in the 0th column of the matrix;

[0158] S92: Dynamic programming calculation:

[0159]

[0160] In the formula; represents the value of the cumulative local similarity score matrix at position ; represents the value of the cumulative local similarity count matrix at position ; represents the similarity metric function of species A and species B at time points and ω;

[0161] S93: Calculate the maximum score;

[0162]

[0163] In the formula, P max (A, B) is the maximum local similarity score of species A and B; L max (A, B) represents the maximum number of similarities of species A and B;

[0164] S94: Calculate the final score:

[0165]

[0166] S sgn (A, B) = sgn [P max (A, B) - L max (A, B)];

[0167] In the formula, S max (A, B) represents the normalized maximum similarity score of species A and B; S sgn (A, B) represents the final LS score.

[0168] In step S10, the process of evaluating statistical significance is as follows:

[0169] Permutation test: Fix B and rearrange all columns of A multiple times. The number of permutations is H. For each permuted dataset A (k) , calculate S max (A (k) , B):

[0170]

[0171] where Q represents the probability of the current local similarity score under the condition that the two species are not associated; I is the indicator function;

[0172] FDR correction: For the Q values of all species factor pairs, calculate according to the Storey method; Suppose there are E species factors in total, then there are eLSA pairs calculated; Adjust the Q value of each species factor pair through the algorithm of the Storey method to control the false positive rate in multiple tests.

[0173] In step S11, the process of constructing the association network is as follows:

[0174] For two species A and B, if in the time interval [s 1 , t 1 of A and the time interval [s 2 , t 2 of B, they have the highest LS score, and s 1 < s 2 , then draw an arrow from A to B in the network, indicating that A may lead B in time, that is, A may have an impact on B; if s 1 > s 2 , then the arrow direction is opposite, and the solid and dashed lines are used to represent the positive and negative of the association. The solid line represents positive correlation, and the dashed line represents negative correlation.

[0175] In the specific content of the above specific implementation manner, each technical feature can be combined arbitrarily without contradiction. For the sake of concise description, not all possible combinations of the above technical features are described. However, as long as the combinations of these technical features do not exist in contradiction, they should all be considered as the scope described in this specification.

[0176] Obviously, the above embodiments of the present invention are only examples for clearly explaining the present invention, rather than limiting the implementation manner of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the claims of the present invention.

Claims

1. A method for predicting core fish communities in mangrove habitats, characterized in that: The following steps are involved: S1: Obtain eDNA data of fish in mangrove waters at different times and perform data analysis to obtain the composition and structure data of fish schools; S2: Collect feeding data of each species in the fish school, determine the functional positioning of the species and the ecological role of the species in energy flow, and determine the food web; S3: Combine biomass and food web to construct ecological network models; S4: Calculate the characteristic index of each species node in the ecological network model; S5: Standardize the characteristic indicators; S6: Calculate the comprehensive score of the standardized characteristic indicators; S7: Rank the species according to the comprehensive scores, compare the comprehensive scores with the set thresholds, and distinguish between important species and non-important species; S8: Preprocessing of important species data from different periods; S9: Calculate the highest similarity score between any two important species using the eLSA dynamic programming algorithm; S10: Assessing statistical significance; S11: Use statistically significant species association results to construct association networks; S12: Analyze association networks, calculate network metrics for important species, and identify core fish populations based on time-lagged patterns.

2. The method for predicting core fish communities in mangrove habitats according to claim 1, characterized in that: In step S3, the process of constructing the ecological network model is: S31: Construct an adjacency matrix based on the food web; the rows and columns of the adjacency matrix represent different fish species. When species i preys on species j, the matrix element a ij =1, otherwise a ij =0; S32: Determine the flow matrix by combining biomass and predator-prey relationships in the food web; S33: Draw a topological network relationship diagram among fish species.

3. The method for predicting core fish communities in mangrove habitats according to claim 1, characterized in that: In step S4, the characteristic indicators include point degree, point in-degree, point out-degree, betweenness centrality, closeness centrality, information centrality, topological importance index, criticality index, upstream criticality index, downstream criticality index, dispersion, distance weight dispersion and distance weight reachability.

4. The method for predicting core fish communities in mangrove habitats according to claim 3, characterized in that: The calculation expression of the characteristic index is: D i =D in,i +D out,i ; i≠k,i≠j and j <k; K(i)=K b (i)+K t (i); Where D i represents the point degree of species i, which is the total number of prey and predators of species i; D in,i represents the in-degree, which is the number of prey of species i; D out,i represents the point-out degree, which is the number of predators of species i; BC i represents the betweenness centrality of species i, which is used to quantitatively analyze the species’ ability to control information exchange within the community; CC i It represents the closeness centrality of species i, which is used to quantitatively measure the degree of advantage of species in transmitting information; IC i represents the information centrality of species i, which is used to quantitatively measure the information transmission ability of species to any two connected species; N represents the number of species appearing in the survey; g jk represents the number of shortest paths between species j and species k; g jk(i) represents the number of shortcuts between species j and species k that pass through species i; d ij represents the shortcut distance between species i and species j; I ij represents the inverse of the shortcut distance between species i and species j; n represents the maximum step size of the analysis; Represents the importance index of species i on the topological structure of fish community after n steps; α m,j represents the impact of species i on species j when species i reaches species j after m steps; a m,ji represents the impact of species i on species j at step m; K b (i) represents the upward criticality index of species i; K b (j) represents the upward criticality index of species j; species j is the direct predator of species i; e(i)(j) is the number of direct prey species of species j; K t (i) is the downward criticality index of species i; K t (j) is the downward criticality index of species j; f(i)(j) is the number of direct predators of species j; K(i) represents the criticality index of species i; F represents the discreteness, with a value of (0, 1); s i represents the number of species i in the i-th discrete population; D F represents the distance weight discreteness, with a value of (0, 1); D R represents the distance-weighted reachability, with a value of (0, 1); d Mj Represents the distance from a series of points M to any point j.

5. The method for predicting core fish communities in mangrove habitats according to claim 4, characterized in that: In step S5, the process of standardizing the characteristic index is as follows: First check the indicator distribution, analyze whether there are extreme values ​​through histogram or box plot, check the mean, standard deviation, maximum and minimum values ​​of the indicator; according to the type of data, use the following methods to standardize the data: If the data is unevenly distributed and has extreme values, but does not have a skewed distribution, choose Z-score to standardize the data: In the formula, x represents the original indicator value; x' represents the standardized indicator value; μ represents the mean value of the indicator; σ represents the standard deviation of the indicator; The index values ​​are standardized to a standard normal distribution. The standardized data fluctuates around 0, and the position of each species relative to the average value in each topological index can be intuitively seen; If the data has a power distribution or an extremely skewed distribution and there are no negative or zero values, choose the log transformation: x′=log(x+1); If the data distribution is uniform and has no extreme values, and the physical meaning of the indicator is clear, choose the Min-Max normalization method: In the formula, x min Indicates the minimum value of the indicator; x max Indicates the maximum value of the indicator.

6. The method for predicting core fish communities in mangrove habitats according to claim 5, characterized in that: In step S6, the calculation process of the comprehensive score is: Use principal component analysis to calculate variance contribution and assign weights to each indicator; Perform weighted summation on the standardized index values: In the formula, S represents the comprehensive score; q represents the total number of indicators; w r represents the weight of the rth indicator; x' r Represents the standardized value of the rth indicator.

7. The method for predicting core fish communities in mangrove habitats according to claim 6, characterized in that: In step S8, preprocessing includes processing of data matrix input and missing values, processing of duplicate data, and data standardization; The processing procedure for data matrix input and missing values is as follows: The important species data for different periods are organized into a matrix form, where each row represents the data of a species factor at different time points, and the data are organized into a time series matrix X: X=X [1:T][1:C] ; In the formula, T represents the number of time points; C represents the number of species types; If there are missing values in the data, they are filled according to the method provided by the eLSA software; if the zero-order spline method is selected, the missing positions are directly assigned a value of 0; if the third-order spline method or the nearest neighbor method is selected, the missing values need to be calculated according to the corresponding interpolation algorithm; The process of processing repeated data is as follows: use the function F(X i ) to summarize, the summary method used is the MAD weighted median method; MAD(X i )=Median(|X i -Median(X i )|); In the formula, F(X i ) represents the standardized summary result; MAD(X i ) represents the absolute median difference; Median(X i ) represents the median; The data standardization process is: apply the selected function to transform the important species data at each time point to obtain F(X i ), and then normalized using a rank-based normal score transformation method.

8. The method for predicting core fish communities in mangrove habitats according to claim 7, characterized in that: In step S9, the time series data of important species A and B are assumed, and D is the set time delay limit; P represents the accumulated local similarity score matrix; L represents the accumulated similarity quantity matrix; The calculation process for the highest similarity score between two important species is as follows: S91: Initialization; L 0,ω =0; Where P 0,ω represents the similarity score in the 0th row of the matrix; represents the similarity score in the 0th column of the matrix; L 0,ω Represents the cumulative number of similarities in the 0th row of the matrix; Represents the cumulative number of similarities in the 0th column of the matrix; S92: Dynamic programming calculation: Where; Represents the cumulative local similarity score matrix at position The value of Represents the cumulative local similarity matrix at position The value of Indicates species A and species B at time point and ω’s similarity measure function; S93: Calculate the maximum score; Where P max (A, B) Maximum local similarity score between species A and B; L max (A,B) represents the maximum similarity between species A and B; S94: Calculate the final score: S sgn (A,B)=sgn[P max (A,B)-L max (A,B)]; In the formula, S max (A, B) represents the standardized maximum similarity score of species A and B; S sgn (A, B) show the final LS scores.

9. The method for predicting core fish communities in mangrove habitats according to claim 8, characterized in that: In step S10, the process for evaluating statistical significance is as follows: Permutation test: fix B and rearrange all columns of A multiple times, the number of permutations is H, for each permuted data set A (k) , calculate S max (A (k) ,B): In the formula, Q represents the probability of the current local similarity score when the two species are not associated; I is the indicator function; FDR correction: For all pairs of species factors, the Q value is calculated according to the Storey method; assuming there are E species factors, then The eLSA was calculated pairwise; the Q value of each species factor pair was adjusted by the Storey method algorithm to control the false positive rate in multiple tests.

10. The method for predicting core fish communities in mangrove habitats according to claim 9, characterized in that: In step S11, the process for constructing the association network is as follows: For two species A and B, if they have the highest LS score within the time intervals [s1, t1] of A and [s2, t2] of B, and s1 < s2, then an arrow is drawn from A to B in the network, indicating that A may be temporally ahead of B, that is, A may have an impact on B; if s1 > s2, the arrow direction is opposite, and the solid or dashed nature of the line is used to represent the positive or negative nature of the association, with a solid line representing a positive correlation and a dashed line representing a negative correlation.

Citation Information

Patent Citations

  • Quantitative calculation method for lake wetland ecological water requirement in arid region

    CN103530530A

  • Graph matching based cross-species biological access discovery method

    CN107832583A

  • Screening method of fish indicator species

    CN118711677A

  • Method for evaluating health condition of water body by fish integrity index based on machine learning and environmental DNA (Deoxyribose Nucleic Acid) technology

    CN118866124A

  • Method and system for identification of key driver organisms from microbiome / metagenomics studies

    EP3276516B1

Cited By

  • Key habitat identification method, device and equipment based on species-habitat network

    CN121210931A

  • Method and system for evaluating influence of construction disturbance on biodiversity of natural reserve

    CN121329190A