Biodiversity prediction methods
By screening ecological data and employing a random forest model, the method addresses fragmented biodiversity data to predict comprehensive spatial distribution, improving conservation planning.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- INST OF ZOOLOGY GUANGDONG ACAD OF SCI
- Filing Date
- 2025-03-31
- Publication Date
- 2026-05-29
AI Technical Summary
Existing biodiversity survey data are scattered and localized, preventing the determination of a comprehensive spatial distribution pattern, which hinders conservation planning.
A method involving ecological information screening based on species integrity indicators and accumulation rates, followed by a biodiversity prediction model using a random forest model to predict biodiversity in unobserved areas, integrating preliminary predictions with observed data to achieve a comprehensive biodiversity assessment.
Enables the prediction of biodiversity across an entire region, enhancing conservation planning by providing accurate spatial distribution patterns.
Smart Images

Figure 2026088996000040 
Figure 2026088996000041 
Figure 2026088996000042
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biotechnology, and more particularly to a method for predicting biodiversity.
Background Art
[0002] The spatial distribution pattern of biodiversity is one of the core scientific issues in ecology, biogeography, and conservation biology, and is also the basis for conservation plans such as conservation priority areas, ecological corridors, and ecosystem restoration. Currently, in areas where biodiversity is not clear (also called areas to be predicted, such as areas where it is necessary to predict biodiversity as a whole), many biodiversity surveys have been carried out, and detailed and reliable site distribution data of biodiversity have been accumulated. However, these biodiversity survey data are often in a scattered and local distribution, and cannot provide the spatial distribution pattern of biodiversity as a whole for the area to be predicted, and thus cannot contribute to the conservation plan of the area to be predicted.
[0003]
Summary of the Invention
[0003] A method for predicting biodiversity, comprising the following steps: Obtaining ecological information of the area to be predicted within a certain period, where the ecological information of the area to be predicted within a certain period includes multiple sets of ecological data, and each set of ecological data in the multiple sets of ecological data includes environmental variable data and biological site distribution data, the environmental variable data includes climate factors, terrain factors, habitat factors, and disturbance factors, the climate factors include rainfall data and temperature data, the terrain factors include elevation data, the habitat factors include vegetation data, the disturbance factors include population density, and the biological site distribution data includes species name, number of individuals, and discovery location, Based on species integrity indicators and species stock rates, ecological information for areas to be predicted over a certain period of time is screened. Lean and obtain screened ecological information, and screen the area to be predicted. The region corresponding to the observed ecological information is determined as the observed region, and the region to be predicted is Regions other than those already observed are designated as unobserved regions. Based on biodiversity prediction models and environmental data of unobserved areas within a certain period, We determined preliminary predicted values for the distribution data of biological sites in the unobserved areas within the region. Here, the biodiversity prediction model is a random forest model, and the biodiversity prediction model The lure is trained using screened ecological information, Biological site distribution data for observed areas within a certain period and biological site distribution data for unobserved areas within a certain period Based on preliminary predictions of thread distribution data, predict the biodiversity of the area to be predicted within a certain period. A method for predicting biodiversity, characterized by determining the results. As one aspect of the present invention, based on the species integrity index and the species accumulation rate, predictions can be made for a certain period of time. The ecological information of the region is screened, the screened ecological information is obtained, and predictions are made. The areas corresponding to the screened ecological information in the region are determined as observed areas. The steps for determining the unobserved region as the area outside the observed region in the area to be predicted are as follows: Includes: The area to be predicted is divided into multiple blocks, Multiple blocks are screened for multiple first blocks, and the first block is determined by the following conditions. The following conditions must be met: The discovery location falls within the geographical area of the first block. At least one set of ecological data exists within the ecological information. For each of the multiple first blocks, the discovery location of each first block is Associating ecological data within the geographical range of the rock, and the ecological data associated with each first block From the biological site distribution data in the state data, the species completeness index and species accumulation rate of each first block were determined. Calculate, Multiple first blocks are screened for multiple second blocks, and the second blocks are as follows: Conditions met: The species integrity index of the second block is greater than the first threshold, and the species of the second block The accumulation rate is less than the second threshold, where the range of the first threshold value is [0.85, 1]. The range of the second threshold value is [0, 1]. Ecological data associated with multiple second blocks are screened as ecological information. Use this method to merge multiple second blocks, obtain the observed area within the region to be predicted, and make the prediction. Merge blocks within the region that are not already observed to obtain unobserved areas. As one aspect of the present invention, The formula for calculating the species completeness index for each first block is as follows: TIFF2026088996000001.tif1853 Here, TIFF2026088996000002.tif33 shows the species completeness index for each first block. TIFF2026088996000003.tif36 shows the number of species in each first block, TIFF2026088996000004.tif47 shows the species richness index. TIFF2026088996000005.tif43 shows the number of species with a population of 1 in each first block. TIFF2026088996000006.tif43 shows the number of species with a population of 2 in each first block. The species accumulation rate is the slope of the end of the cumulative curve of the number of species in a block over a certain period of time. The line represents the accumulation of the number of species within a certain period as the amount of biological site distribution data increases. It is a curve formed by [something]. As one aspect of the present invention, a biodiversity prediction model and environmental variables of unobserved areas within a certain period Based on the data, determine a preliminary prediction value of the biological site distribution data of unobserved areas within a certain period The step includes the following: Input the environmental variable data of unobserved areas within a certain period into the biodiversity prediction model, and the biodiversity prediction model outputs a preliminary prediction value of the biological site distribution data of unobserved areas within a certain period. As one aspect of the present invention, the step of determining the biodiversity prediction result of the area to be predicted within a certain period includes the following: The sum of the preliminary prediction value of the biological site distribution data of the unobserved area and the biodiversity prediction residual of the unobserved area is used as the final prediction value of the biological site distribution data of unobserved areas within a certain period, Next, combine the biological site distribution data of the observed areas within a certain period and the final prediction value of the biological site distribution data of unobserved areas within a certain period, and obtain the biodiversity prediction result of the area to be predicted within a certain period. As one aspect of the present invention, the method for obtaining the biodiversity prediction residual of unobserved areas is as follows , Randomly select multiple sets of ecological data from the screened ecological information as the residual test set Each set of ecological data in the residual test set includes environmental variable data and biological site distribution data corresponding to the environmental variable data, Input the environmental variable data in the residual test set into the biodiversity prediction model, and the biodiversity prediction model outputs a preliminary prediction value of the biological site distribution data, Next, the difference between the biological site distribution data corresponding to the environmental variable data in the residual test set and the preliminary prediction value of the biological site distribution data output from the biodiversity prediction model is used as the prediction residual of the observed area, The prediction residuals of the observed areas are interpolated by kriging to obtain the prediction residuals of the areas to be predicted. The prediction residuals of the observed areas in the prediction residuals of the areas to be predicted are removed, and the biodiversity prediction residuals of the unobserved areas are obtained. As one aspect of the present invention, the fixed period is 3 to 10 years, and the area to be predicted is an ecological protection area is.
Advantages of the Invention
[0004] In the biodiversity prediction method provided by the present invention, ecological information of the areas to be predicted is screened from the species completeness index and the species accumulation rate, and the screened ecological information is used to divide the areas to be predicted to obtain observed areas and unobserved areas. Next, a biodiversity prediction model (obtained by training based on the screened ecological information) and environmental variable data of the unobserved areas within a fixed period are used to obtain preliminary predicted values of the biological site distribution data of the unobserved areas within a fixed period. Then, based on the preliminary predicted values of the biological site distribution data of the unobserved areas within a fixed period and the biological site distribution data of the observed areas within a fixed period, the biodiversity prediction results of the areas to be predicted within a fixed period are determined. As can be seen from the above content, the above method first screens fragmented ecological information within the areas to be predicted, obtains ecological information with high data completeness (i.e., the screened ecological information), and then predicts the areas to be predicted corresponding to the ecological information with low data completeness according to the biodiversity prediction model to obtain preliminary predicted values of the biological site distribution data of the unobserved areas. Finally, the preliminary predicted values of the biological site distribution data of the unobserved areas and the biological site distribution data of the observed areas are used to obtain the spatial distribution pattern of the biodiversity of the entire area to be predicted, thereby contributing to the protection plan of the area to be predicted. [Brief explanation of the drawing]
[0005] [Figure 1] This is a schematic diagram (part 1) of the biodiversity prediction method. [Figure 2] This is the second schematic diagram of the biodiversity prediction method. [Figure 3] This is the third schematic diagram of the biodiversity prediction method. [Figure 4] This is a schematic diagram of the biodiversity distribution grid. [Modes for carrying out the invention]
[0006] The embodiments of this application predict biological site distribution data observed in the area to be predicted, By predicting the distribution of biological sites that have not been observed in the area to be predicted, Obtain the spatial distribution pattern of biodiversity across the entire region to be predicted. Naturally, biodiversity refers to the diversity of all living species, genes, and ecosystems on Earth. This refers to biodiversity, which includes genetic diversity, species diversity, and ecosystem diversity. In the example given, biodiversity refers to the diversity of species. In conventional techniques, biodiversity survey data tend to show a fragmented and localized distribution. Therefore, it is not possible to predict the spatial distribution pattern of biodiversity across the entire region, This does not contribute to conservation planning of the area to be measured. In order to solve the above problem, the practical application of this application The case study involves analyzing fragmented ecological information within the area to be predicted to predict the overall biodiversity of the area. By predicting the spatial distribution patterns of sex, we can contribute to conservation planning for areas that need to be predicted. This provides a method for predicting variability. As shown in Figure 1, the biodiversity prediction method provided by the embodiment of this application is S101~S Includes 104. S101 acquires ecological information for the area to be predicted within a certain period. Here, the ecological information of the region to be predicted within a certain period includes multiple sets of ecological data. Each set of ecological data within the ecological data consists of environment variable data and biological site distribution data. This includes environmental data such as climate factors, topographic factors, habitat factors, and disturbance factors. Climate factors include data such as rainfall and temperature. Topographic factors include elevation data. This includes data such as habitat factors, including vegetation data. Disturbance factors include data such as population density. Includes data. Biological site distribution data includes species name, number of individuals, and discovery location. Naturally, the above-mentioned period could be 3 years, or 10 years (for example, 2012 1 It could be from [month] to December 2022, or any fixed period between 3 and 10 years. The area to be predicted above may be an ecological protected area or an ecological restoration area. Alternatively, it could be in other areas where it is necessary to determine the spatial distribution patterns of biodiversity. In the embodiments of this application, the above-mentioned period and the area to be predicted are further limited It is not something that can be done. The organism may be a bird, mammal, reptile or amphibian, as can be selected. Taking birds as an example, the above species name is Passerdomesticus (house sparrow). ), the tree sparrow (Passermontanus) and the stone sparrow (Petronia p This includes etronia, and therefore, in the case of the tree sparrow, the biological site distribution data is This includes the species name of the tree sparrow, the number of tree sparrows, and the location where the tree sparrows were found. In a particular application scenario, the specific parameters of the above environment variable data are shown in Table 1. In Table 1, precipitation heterogeneity refers to the non-uniform distribution of precipitation over time and space. Altitude heterogeneity refers to the characteristics of environmental resources, ecological processes, and biological communities exhibited at different elevations. It refers to differences in characteristics, and the Normalized Density Vegetation Index (NDVI) is used to quantify vegetation cover and health. This refers to the remote sensing index, and crown height refers to the height of the top of the tree crown relative to the ground. It tastes good. Table 1: Environment Variable Data Table TIFF2026088996000007.tif103145 S102, based on species integrity indicators and species accumulation rates, predict the ecology of the area to be predicted over a certain period. The information is screened, the screened ecological information is obtained, and the area to be predicted is... The areas corresponding to the screened ecological information are determined as observed areas, and the areas to be predicted are determined. Regions within a given area that have not been observed are designated as unobserved regions. Naturally, after dividing the region and obtaining the observed area and the area to be predicted, the prediction within a certain period of time The ecological information of the area to be measured includes ecological information of observed areas within a certain period and unobserved areas within a certain period. The region is divided into ecological information. Ecological information of observed regions within a certain period consists of multiple sets of ecological data. This includes ecological information for unobserved areas within a certain period, and the ecological data includes multiple sets of ecological data. Selectively, as shown in Figure 2 in conjunction with Figure 1, S102 includes S1021 to S1024. . S1021. Divide the area to be predicted into multiple blocks. In one application scenario, the area to be predicted is divided into multiple grids (corresponding to the multiple blocks mentioned above). Divide the area into sections and obtain a grid map of the area to be predicted. Then, predict the ecology of the area to be predicted. Preprocessing of environment variable data in the information (projection to the same coordinate system, resampling to the same resolution) (including cutting out the area to be predicted) to generate auxiliary data, and the auxiliary data Associate with the corresponding grid location and retrieve environment variable data corresponding to the location of each grid. It's advantageous. S1022 screens multiple first blocks from multiple blocks. The first block described above satisfies the following conditions: the discovery location falls within the geographical area of the first block. There is at least one set of ecological data within the ecological information of the area to be predicted for a certain period. S1023, for each of the multiple first blocks, the location of each first block The location is associated with ecological data that falls within the geographical range of each first block, and is associated with each first block. From the biological site distribution data in the attached ecological data, the species integrity indicators for each first block Obtain the seed accumulation rate. In one embodiment, the formula for calculating the species integrity index for each first block is as follows: TIFF2026088996000008.tif1853 Here, TIFF2026088996000009.tif33 shows the species completeness index for each first block. TIFF2026088996000010.tif36 shows the number of species in each first block, TIFF2026088996000011.tif47 shows the species richness index. TIFF2026088996000012.tif43 shows the number of species with a population of 1 in each first block. TIFF2026088996000013.tif43 shows the number of species with a population of 2 in each first block. Taking one first block as an example, if the organism is a bird, it would be associated with this first block. The ecological data collected is {Species name: House sparrow, Number of individuals: 1, Location of discovery: Left side within the first block}. Placed}, {Species name: Stone sparrow, Number of individuals: 3, Location found: Center of the first block}, {Species name: Tree Sparrow, Individual: 1, Location: Right side of Block 1}, {Species: Domestic Sparrow, Individual Number: 1, Discovery location: Right side within the first block}. And the seed in this first block When calculating the completeness index, the species within this first block (house sparrow, stone sparrow, tree sparrow) number TIFF2026088996000014.tif36=3, and this represents the number of species (tree sparrows) with a population of 1 within this first block. TIFF2026088996000015.tif43=1, and the number of individuals of a species (house sparrow) with a population of 2 within this first block. TIFF2026088996000016.tif43=1, and this is the species completeness index for the first block. TIFF2026088996000017.tif33=1. The above species accumulation rate is the slope of the end of the cumulative curve of the number of species in one block over a certain period of time. The cumulative curve shows the accumulation of the number of species over a certain period as the amount of biological site distribution data increases. This is a curve formed by [the following]. Selectively, the horizontal coordinate of this cumulative curve is related to the first block. This represents the number of linked biological site distribution data points, with the vertical coordinate being the number of species within the first block. Since this cumulative curve is a commonly used technical means in the art, this application The specific process for constructing the cumulative curve is not described in detail in this embodiment. S1024 screens multiple second blocks from multiple first blocks. The second block above satisfies the following conditions: The species integrity index of the second block is greater than the first threshold. Larger than the second threshold, the seed accumulation rate in the second block is smaller than the second threshold. Here, the range of the first threshold value is... The value is [0.85, 1], and the range of the second threshold value is [0, 1]. Selectively, the first threshold may be any value in [0.85, 1], and the second threshold The value may be any value in the range [0, 1], and the values of the first and second thresholds described above are within a reasonable range. The options within the box may be arbitrarily selected, and are not limited to the embodiments of this application. Furthermore, for one second block, the species integrity index of this second block is greater than the first threshold. If the species accumulation rate is large and below the second threshold, the ecosystem associated with this second block The data can be considered to have high data integrity, in which case it can be associated with this second block. The reliability of the collected ecological data is high, and it better describes the biodiversity of this block. Please note that this means it is possible. S1025, Ecological data associated with multiple second blocks were screened. As state information, multiple second blocks are merged to obtain the observed area within the region to be predicted, and then The blocks within the region to be predicted, excluding the observed areas, are merged to obtain the unobserved areas. Naturally, the data integrity of the ecological data associated with multiple second blocks is high. If not, ecological data associated with multiple second blocks (i.e., screened) (Ecological information) is highly reliable and can better describe the biodiversity of this block. This means that it is necessary to observe, predict, or supplement the distribution data of multiple second-block biological sites. Since there is no longer a need for it, the region merged by multiple second blocks is defined as the observed region. Please understand that blocks outside the observed area within the region to be predicted are reliable biological regions. Due to a lack of site distribution data, areas outside the observed regions within the area to be predicted remain unobserved. This is defined as the measurement area. S103, based on a biodiversity prediction model and environmental variable data of unobserved areas over a certain period. This determines preliminary predictions of biological site distribution data for unobserved areas within a certain period. Here, the biodiversity prediction model is a random forest model, and The model includes multiple decision trees. The biodiversity prediction model uses screened ecological information. It is obtained through training using [a specific method / tool]. Specifically, S103 above includes: collecting environmental data of unobserved regions within a certain period of time. The biodiversity prediction model is input, and the model predicts biodiversity in unobserved areas over a certain period of time. Outputs preliminary predicted values for the distribution data. The training process for the biodiversity prediction model described above is as follows: Step 1: Multiple sets of ecological data are used as training sets from the ecological information obtained through screening. Select randomly as the target. The ecological data for each set in the above training set all include environment variable data and environment variable data. Includes corresponding biological site distribution data. Step 2, build a random forest model. The input to the random forest model is This is environment variable data, and the output of the random forest model is a preliminary version of the biological site distribution data. This is a predicted value. Step 3: Train the random forest model with the training set, and obtain the results obtained from the training. The Common Forest model will be used as a biodiversity prediction model. In the above application scenario, the training process of the random forest model (as opposed to step 3 above) The responses are as follows: Step 3.1, randomly select a subsample set. Using putbacks, a subset of samples is randomly selected from the training set several times, and multiple new A new set of subsamples is formed, and each sample set in the multiple subsample sets is The discrepancy is the biological site distribution data for the observed area within a certain period, and the above biological site distribution Includes environment variable data corresponding to the data. Step 3.2: Randomly select a subset of features. Using putback, a feature subset is randomized from the total number of features in a subsample set. Select "M". The size of this subset is smaller than the total number of features. For example, for one feature subset, multiple environment variables (e.g., temperature, precipitation) Select this option, and then the data corresponding to multiple environment variables in one subsample set. Select a sample. Then, build and train a decision tree based on that sample. Step 3.3: Train the decision tree. The subsample sets and feature subsets selected in Steps 3.1 and 3.2 This is used to train one decision tree in the random forest model. This decision tree is one The root node, multiple internal nodes connected to the root node, and multiple internal nodes Includes multiple connected subnodes. Here, the root node is the entire subsample set. The subnodes contain the data, and each subnode is a subset of data selected based on the root node. The leaf subnodes include the prediction results of the decision tree. During the subsequent growth process of the decision tree, each sub The node, until it reaches a stopping condition, divides the features already selected by the parent node. Therefore, we continue to select the best features from the remaining features. Step 3.4: Repeat steps 3.2 and 3.3, using a random forest. Complete the training of all decision trees in the model to obtain a biodiversity prediction model. Step 3.5: Regression prediction. Biodiversity prediction models obtain prediction results by averaging the prediction results of each decision tree. . For example, the above prediction results TIFF2026088996000018.tif43 satisfies the following equation. TIFF2026088996000019.tif1326 Here, TIFF2026088996000020.tif33 shows the number of decision trees in a random forest model. TIFF2026088996000021.tif32 indicates the decision tree number. TIFF2026088996000022.tif46 is a number The output of the decision tree for TIFF2026088996000023.tif32 is shown. Taking the case of birds as an example, the environmental data of unobserved regions within a certain period of time can be run Multiple decision trees in the dam forest model are input, and these multiple decision trees process the input environment changes. Based on several data points, leaf subnodes are matched, and multiple decision tree output results are obtained. The output results of multiple decision trees are averaged, and finally the prediction result of the biodiversity prediction model is obtained. This also allows us to obtain preliminary predictions of the distribution of biological sites in unobserved areas within a certain period. S104, biological site distribution data for observed areas within a certain period and unobserved areas within a certain period. Based on preliminary predictions of the distribution of biological sites in the region, the biological information to be predicted for the region within a certain period of time. Determine the diversity prediction results. As shown in Figure 3, in conjunction with Figure 2, the above S104 is selected to perform S1041 to S1042. include. S1041, Preliminary predictions of biological site distribution data in unobserved areas and biodiversity in unobserved areas The sum of the predicted residuals is used as the final predicted value for the distribution of biological sites in unobserved areas within a certain period. In S1041 described above, the method for obtaining the predicted biodiversity residuals for unobserved areas is as follows: Step 1: From the ecological information obtained through screening, multiple sets of ecological data are used as residual test results. Select randomly as a set. The ecological data for each set in the above residual test set are all environment variable data and environment variable data Includes biological site distribution data corresponding to the data. Step 2: Input the environment variable data from the residual test set into the biodiversity prediction model. The biodiversity prediction model outputs preliminary predictions for the distribution of biological sites. Step 3: Next, the biological site distribution data corresponding to the environment variable data in the residual test set. , and the difference between the preliminary predicted values of the biological site distribution data output from the biodiversity prediction model. The minutes are taken as the predicted residual for the observed region. Step 4: Perform kriging interpolation on the predicted residuals of the observed region to obtain the predicted residuals of the region to be predicted. ru. Furthermore, kriging interpolation, as an interpolation method that takes spatial correlation into consideration, is used in the embodiments of this application. Note that the requirements for spatial correlation during the interpolation of predicted residuals in the measured region can be met. I want to be treated that way. Selectable, predicted residuals interpolated by kriging interpolation TIFF2026088996000024.tif47 satisfies the following equation. TIFF2026088996000025.tif1332 Here, TIFF2026088996000026.tif32 shows the index variable for traversing all sampling points. TIFF2026088996000027.tif33 indicates the number of sampling points. TIFF2026088996000028.tif43 shows the weighting coefficients and the predicted residuals at the point (x0, y0). TIFF2026088996000029.tif47 and residual observations This is the optimal coefficient that satisfies the minimum difference in TIFF2026088996000030.tif46. The above point (x0, y0) is the prediction... This is a single coordinate point in a 2D coordinate system defined based on a grid map of the region. The 2D coordinate system includes the x and y axes, and the origin of the 2D coordinate system is the grid of the region to be predicted. It could be the center position of the map, or any other position on the grid map of the area to be predicted. This may also be the case, and the embodiments of this application do not further limit the two-dimensional coordinate system. Step 5: Remove the observed region's prediction residual from the prediction residual of the region to be predicted, and remove the unobserved region. Obtain the predicted residuals for biodiversity. S1042, biological site distribution data for observed areas within a certain period and unobserved areas within a certain period. By combining the final predicted values of the biological site distribution data for a given region, the region to be predicted for a certain period of time... Obtain biodiversity prediction results. In one application scenario of the above method, taking the case where the species is a bird as an example, Figure 4 shows the final result. This shows a grid map of the spatial distribution pattern of the overall biodiversity within the region to be predicted. In Figure 4, each grid corresponds to one color, and each color indicates the number of bird species found in that grid. vinegar. In one application scenario, a grid map of the area to be predicted with three resolutions (grid size) is used. The distances are 10km x 10km, 50km x 50km, and 100km x 100km respectively. In total, a random forest model was trained, and the random forest obtained from the training was... Test the model (i.e., the biodiversity prediction model). The training and testing processes for the random forest model in each situation are as follows: The screened ecological information obtained in S102 above was divided into a training set and a test set in an 8:2 ratio. The random forest model obtained by dividing it into sets and training with the training set (i.e.) After obtaining the biodiversity prediction model, the biological site distribution data (number of species, etc.) in the test set is obtained. The difference between the observed region and the prediction results of the biodiversity prediction model is used as the predicted residual, and kriging interpolation is performed. Through this process, the predicted residual for the unobserved area is obtained, and from this, the predicted residual for the entire region to be predicted is obtained. Then, the environment variable data in the test set is input into the biodiversity prediction model, and randomized... The sum of the prediction results of the Forest model and the corresponding prediction residuals, and the environment variables in the test set. Pearson correlation analysis is performed on the biological site distribution data (such as the number of species) corresponding to the data. At the same time, the above three situations (Situation 1: Grid size is 10km x 10km, Situation 2) Situation 1: Grid size is 50km x 50km, Situation 3: Grid size is 100km x 1 For each situation within 00km, the Maxent distribution model and the conventional exponential Using the distribution surface superposition method, biodiversity estimation is performed according to the same resolution and area. The process involves predicting the number of species of unobserved regions in the test set from the data in the training set. ), Pearson correlation of corresponding biological site distribution data (i.e., number of species) in the test set Perform an analysis. Furthermore, the above Maxent species distribution model and the conventional expert distribution surface superposition method are conventional This relates to the technology, and in the embodiments of this application, the specific implementation process of the two algorithms described above is described. I will not elaborate further. The combination of a random forest model and kriging interpolation provided by the embodiments of this application Estimated using a combination of the Maxent distribution model and the conventional expert distribution surface superposition method. The results of the correlation between the obtained biological site distribution data and the actual situation are shown in Tables 2, 3 and 4 below. This is shown in 4. Table 2: Comparison Table of Correlation Test Results for Situation 1 TIFF2026088996000031.tif97156 Table 3: Comparison Table of Correlation Test Results for Situation 2 TIFF2026088996000032.tif98160 Table 4: Comparison Table of Correlation Test Results for the Third Situation TIFF2026088996000033.tif108156 Naturally, the higher the correlation obtained from Pearson correlation analysis, the more accurate the predicted value is. The closer the value is to , the higher the model accuracy is proven. Generally, correlation |r|>0.7 is A strong correlation is indicated, a moderate correlation is indicated by 0.4 < |r| ≤ 0.7, and a weak correlation is indicated by |r| ≤ 0.4. As can be seen from Tables 2, 3, and 4 above, the biodiversity predictions provided by this application The correlation between the final prediction result and the true value obtained by combining the model and kriging interpolation is, in all cases, It is higher than the conventional expert distribution surface superposition method, and in the first situation, the Maxent distribution model The correlation was higher than that of the biodiversity prediction model and the kriging interpolation method, and the second situation And in the third situation, the correlation of the Maxent distribution model is with the random forest model and the crit Because it is lower than the combination of Ging interpolation methods, the biodiversity prediction provided by this application The combination of measurement models and kriging interpolation predicts the biodiversity of the entire region to be predicted. It is closer to the actual situation and more accurate.
Claims
1. A method for predicting biodiversity, Steps include obtaining ecological information for the area to be predicted within a certain period, Here, the ecological information of the area to be predicted within the aforementioned period includes multiple sets of ecological data, Each set of ecological data within the multiple sets includes environment variable data and biological site data. The environment variable data includes climatic factors, topographic factors, habitat factors, and disturbance factors. The climate factor includes rainfall data and temperature data, and the topographic factor includes elevation data. The organism includes, the habitat factor includes vegetation data, the disturbance factor includes population density, and the organism Site distribution data includes species name, population size, and discovery location. Based on species integrity indicators and species accumulation rates, the ecological information of the area to be predicted within the aforementioned period is used. Screening, obtaining the screened ecological information, and in the area to be predicted The region corresponding to the screened ecological information is determined to be the observed region, and the region to be predicted is A step of determining the areas in the region other than the observed areas as unobserved areas, Based on a biodiversity prediction model and environmental data of the unobserved area over a certain period, The steps include determining preliminary predicted values for the distribution data of biological sites in the unobserved area within the period, Here, the biodiversity prediction model is a random forest model, and the biodiversity The prediction model is obtained by training with the aforementioned screened ecological information. The biological site distribution data of the observed area within the aforementioned period and the unobserved data within the aforementioned period Based on the preliminary predictions of the biological site distribution data in the observation area, the predictions for a certain period should be made. Steps to determine the results of the regional biodiversity prediction, A method for predicting biodiversity, characterized by including the following:
2. Based on the aforementioned species integrity index and species stock rate, the ecological conditions of the area to be predicted within the aforementioned period The report is screened, the screened ecological information is obtained, and the area to be predicted is... The region corresponding to the screened ecological information is determined as the observed region, and predictions are made. The step of determining the areas other than the observed areas in the region to be observed as unobserved areas is as follows: Includes: The area to be predicted is divided into multiple blocks, A plurality of first blocks are screened from the plurality of blocks, and the first block is The following conditions are met: the discovery location falls within the geographical area of the first block within the specified period. At least one set of ecological data exists within the ecological information of the area to be predicted. For each of the aforementioned first blocks, the discovery position of each first block is The ecological data that falls within the geographical area of each of the first blocks is associated with the said first blocks. From the biological site distribution data in the associated ecological data, the species completeness of each of the first blocks Calculate the sex index and species accumulation rate, Screening multiple second blocks from the multiple first blocks, the second block The following conditions are met: The species integrity index of the second block is greater than the first threshold, The species accumulation rate of the second block is smaller than the second threshold, where the value of the first threshold is The range is [0.85, 1], and the range of the second threshold value is [0, 1]. The ecological data associated with the plurality of second blocks is the screened ecological data To be used as a report, the multiple second blocks are merged, and the observed area within the region to be predicted is used. Obtain the region, merge the blocks within the region to be predicted that are not the observed region, and the unobserved region. The method according to claim 1, characterized by obtaining a region.
3. The formula for calculating the species integrity index for each of the first blocks is as follows: Here, The first block indicates the species completeness index of each of the first blocks mentioned above. This indicates the number of species within each of the first blocks. This indicates the species richness index, This indicates the number of individuals of a species where the number of individuals in each of the first blocks is 1. This indicates the number of species where the number of individuals in each of the first blocks is 2. The species accumulation rate is the slope of the end of the cumulative curve of the number of species in one block within the aforementioned fixed period. Yes, and the cumulative curve increases as the number of biological site distribution data increases within the aforementioned period. The curve is formed by the accumulation of the number of the aforementioned species, as described in claim 2. method.
4. Based on the aforementioned biodiversity prediction model and the environmental variable data of the unobserved area over a certain period, Then, a preliminary prediction of the biological site distribution data for the unobserved area within a certain period is determined. The package includes: The environmental data of the unobserved region within the aforementioned period is input into the biodiversity prediction model. The biodiversity prediction model uses the distribution data of biological sites in the unobserved area within the specified period. The method according to claim 1, characterized in that it outputs a preliminary predicted value.
5. The step of determining the biodiversity prediction results for the area to be predicted within the aforementioned period is as follows: Including the following: Preliminary predicted values of biological site distribution data in the unobserved area and biodiversity predictions for the unobserved area. The sum of the residuals is taken as the final predicted value of the biological site distribution data for the unobserved area within the aforementioned period. 、 Next, the biological site distribution data of the observed area within the aforementioned period and the previous data within the aforementioned period By combining the final predicted values of biological site distribution data for the unobserved area, the aforementioned period The method according to claim 1, characterized by obtaining biodiversity prediction results for the region to be predicted.
6. The method for obtaining the predicted biodiversity residuals for the unobserved region is as follows: From the ecological information obtained through the screening process, multiple sets of ecological data were used as residual test sets. The ecological data for each set in the residual test set are randomly selected, and all of them change environmentally. This includes numerical data and biological site distribution data corresponding to the aforementioned environment variable data, The environmental variable data in the residual test set is input into the biodiversity prediction model, The biodiversity prediction model outputs preliminary predicted values for biological site distribution data. Next, the biological site distribution data corresponding to the environment variable data in the residual test set, and between the preliminary predicted values of the biological site distribution data output from the biodiversity prediction model The difference is taken as the predicted residual of the observed region. The predicted residuals of the observed region are interpolated using kriging to obtain the predicted residuals of the region to be predicted. Remove the predicted residual of the observed area from the predicted residual of the area to be predicted, and the unobserved area The method according to 5, characterized by obtaining the predicted residual of biodiversity.
7. The aforementioned period is 3 to 10 years, and the area to be predicted is an ecological protected area. The method according to claim 1, characterized by the present invention.