Biodiversity prediction methods
By screening ecological data using species completeness and accumulation rates, and employing a random forest model with Kriging interpolation, the method addresses fragmented biodiversity data issues, enabling accurate spatial predictions for conservation planning.
Patent Information
- Application Number
- JP2025057724
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2024-11-19
- Filing Date
- 2025-03-31
- Publication Date
- 2025-09-17
- Estimated Expiration
- 2045-03-31
AI Technical Summary
Current biodiversity surveys produce fragmented and localized data, making it difficult to predict spatial distribution patterns across regions, which hinders conservation planning and ecosystem restoration.
A method involving ecological information screening based on species completeness index and species accumulation rate to divide areas into observed and unobserved regions, using a biodiversity prediction model trained with random forest algorithms to estimate biological site distribution data in unobserved areas, and combining these estimates with Kriging interpolation to predict biodiversity across the entire region.
Enables the calculation of spatial distribution patterns of biodiversity, enhancing conservation planning by providing accurate biodiversity predictions across large areas.
Smart Images

Figure 0007739690000034 
Figure 0007739690000035 
Figure 0007739690000036
Abstract
Description
[Technical Field]
[0001] The present invention relates to the field of biotechnology, and in particular to a method for predicting biodiversity. [Background technology]
[0002] Spatial patterns of biodiversity are a core discipline in ecology, biogeography, and conservation biology. This is one of the key issues for conservation planning, such as priority conservation areas, ecological corridors, and ecosystem restoration. Currently, there are areas where biodiversity is unclear (where it is necessary to predict biodiversity as a whole). Many biodiversity surveys have been carried out in areas where predictions are needed, such as the Detailed and reliable site distribution data on biodiversity has been accumulated. However, these biodiversity survey data are often scattered and localized, and predictions are difficult to make. It is not possible to provide a spatial distribution pattern of biodiversity for the entire region to be measured, and it is difficult to predict. This did not contribute to conservation planning for the area. Summary of the Invention
[0003] 1. A method for predicting biodiversity, comprising the steps of: Obtain ecological information for the area to be predicted within a certain period of time, Here, the ecological information of the area to be predicted within a certain period includes multiple sets of ecological data. Each set of ecological data is made up of environmental variables and biological site distribution data. The environmental variable data includes climatic factors, topographical factors, habitat factors, and disturbance factors. The factors include rainfall data and temperature data, the topographic factors include elevation data, and the habitat factors include The data includes vegetation data, disturbance factors include population density, and biological site distribution data includes species names, numbers of individuals, and and the location of the discovery, Based on the species completeness index and species accumulation rate, ecological information of the area to be predicted within a certain period is screened. The ecological information obtained by the screening is then used to screen the area to be predicted. The area corresponding to the scanned ecological information is determined as the observed area, and the area to be predicted is determined as the observed area. The area other than the observed area is determined as an unobserved area, Based on the biodiversity prediction model and environmental variable data of unobserved areas within a certain period, determining preliminary estimates of biological site distribution data for unobserved areas within the Here, the biodiversity prediction model is a random forest model, and the biodiversity prediction model The model is trained using the screened biological information, Biological site distribution data for observed areas within a certain period and biological site distribution data for unobserved areas within a certain period Based on the preliminary estimates of the site distribution data, biodiversity prediction for the area to be forecast within a certain period of time is performed. and determining a result of the biodiversity prediction. In one aspect of the present invention, a species completeness index and a species accumulation rate are used to predict the species to be grown within a certain period of time. Screening of ecological information in the area, obtaining the screened ecological information, and making predictions determining an area corresponding to the screened ecological information in the area to be observed as an observed area; The step of determining the area other than the observed area in the area to be predicted as the unobserved area is as follows: Includes: Divide the area to be predicted into multiple blocks, Screen multiple first blocks from multiple blocks, and the first block meets the following conditions: The first block of the geographical area is within the range of the first block. There is at least one set of ecological data in the ecological information, For each of the plurality of first blocks, the first block is found at a position corresponding to the first block. The biological data associated with each first block is then associated with the biological data that falls within the geographical scope of the block. The species completeness index and species accumulation rate of each first block were calculated from the biological site distribution data in the morphological data. Calculate Screening of multiple second blocks from multiple first blocks, and the second blocks are as follows: Condition met: The species completeness index of the second block is greater than the first threshold, and the species of the second block The accumulation rate is smaller than the second threshold, where the value of the first threshold is in the range [0.85, 1]. , the value range of the second threshold is [0, 1], Biological data associated with multiple second blocks as screened biological information The second blocks are used to merge the observed areas within the area to be predicted. The blocks other than the observed area within the observation area are merged to obtain the unobserved area. As one aspect of the present invention, The calculation formula for the species completeness index of each first block is as follows: TIFF0007739690000001.tif1853 where, TIFF0007739690000002.tif33 shows the species completeness index for each first block, TIFF0007739690000003.tif36 shows the number of species in each first block, TIFF0007739690000004.tif47 shows the species richness index, TIFF0007739690000005.tif43 shows the number of species with an individual population of 1 in each first block, TIFF0007739690000006.tif43 shows the number of species with an individual population of 2 in each first block. The species accumulation rate is the slope of the end of the cumulative curve of the number of species in one block over a certain period of time. The line represents the accumulation of species numbers within a certain period as the number of biological site distribution data increases. The curve formed by As one aspect of the present invention, a biodiversity prediction model and environmental variables in an unobserved area within a certain period of time are used. Based on the data, a preliminary estimate of the distribution of biological sites in unobserved areas within a certain period is determined. The steps of: Environmental variable data from unobserved areas over a certain period of time is input into a biodiversity prediction model to estimate biodiversity. The prediction model outputs preliminary predictions of biological site distribution data in unobserved areas within a certain period of time. As one aspect of the present invention, a process for determining the biodiversity prediction results for a region to be predicted within a certain period of time is Steps include: The sum of the preliminary predicted values of biological site distribution data in unobserved areas and the residuals of biodiversity predictions in unobserved areas is the final predicted value of the biological site distribution data in the unobserved area within a certain period, Next, we will look at the biological site distribution data for observed areas over a certain period of time and the biological site distribution data for unobserved areas over a certain period of time. The final predicted values of the biodiversity site distribution data are combined to estimate the biodiversity of the area to be predicted within a certain period. Obtain gender prediction results. In one aspect of the present invention, a method for obtaining biodiversity prediction residuals in an unobserved area is as follows: , From the ecological information obtained through screening, multiple sets of ecological data were used as residual test sets. The ecological data in each group in the residual test set were randomly selected from the environmental variables. and biological site distribution data corresponding to environmental variable data; The environmental variable data in the residual test set were input into the biodiversity prediction model, and the biodiversity prediction was The model outputs preliminary predictions of biological site distribution data, Next, we used the biological site distribution data corresponding to the environmental variable data in the residual test set, and the biological Differences between preliminary predicted values of biological site distribution data output from biodiversity prediction models have been observed. Let the prediction residuals of the region be The prediction residuals of the observed area are interpolated using Kriging to obtain the prediction residuals of the area to be predicted. The predicted residuals for the observed areas in the area to be predicted are removed, and the biodiversity of the unobserved areas is estimated. Obtain the prediction residuals. In one aspect of the present invention, the predetermined period is 3 to 10 years, and the area to be predicted is an ecologically protected area. is. [Effects of the Invention]
[0004] In the biodiversity prediction method provided by the present invention, the following is calculated from the species completeness index and the species accumulation rate: Screening of ecological information for the area to be predicted, and using the screened ecological information The area to be predicted is divided into observed and unobserved areas. Next, a biodiversity prediction model is used. (Trained based on screened biological information) and untrained within a certain period Using environmental variable data from observed areas, we can analyze biological site distribution data from unobserved areas within a certain period. Obtain a preliminary forecast value. Then, make a preliminary forecast of the distribution data of biological sites in unobserved areas within a certain period. Based on the values and the biological site distribution data of the observed area within a certain period, a forecast for a certain period is made. As can be seen from the above, the above method determines the biodiversity prediction results for the area to be monitored. First, we screen the fragmented ecological information within the area to be predicted and check the data completeness. High-quality ecological information (i.e., screened ecological information) is obtained, and then biodiversity prediction is carried out. According to the observation model, the areas to be predicted corresponding to the ecological information with low data completeness are predicted, and the unobserved areas are predicted. Finally, we obtain a preliminary estimate of the biological site distribution data in the unobserved area. Using the preliminary predicted values of the distribution data of biological sites in the observed area, the predicted location is calculated. This will help to obtain spatial distribution patterns of biodiversity across the region, which will contribute to the conservation planning of the area to be predicted. do. [Brief explanation of the drawings]
[0005] [Figure 1] This is a schematic diagram of the biodiversity prediction method. [Figure 2] This is the second schematic diagram of the biodiversity prediction method. [Figure 3] This is the third schematic diagram of the biodiversity prediction method. [Figure 4] FIG. 1 is a schematic diagram of a biodiversity distribution grid. DETAILED DESCRIPTION OF THE INVENTION
[0006] The embodiments of the present application predict observed biological site distribution data in a region to be predicted; By predicting unobserved biological site distribution data in the area to be predicted, Obtain the spatial distribution pattern of biodiversity across the region to be predicted. Biodiversity, of course, refers to the diversity of all living species, genes, and ecosystems on Earth. Biodiversity includes genetic diversity, species diversity, and ecosystem diversity. In the present embodiment, biodiversity refers to species diversity. In the prior art, biodiversity survey data tend to be fragmented and localized, with a point-like distribution. Therefore, it is not possible to provide a spatial distribution pattern of biodiversity throughout the region to be predicted, and prediction is difficult. In order to solve the above problem, the implementation of this application In this example, the biodiversity of the entire area to be predicted is calculated from fragmented ecological information within the area to be predicted. By predicting the spatial distribution pattern of biodiversity, it is possible to estimate the biodiversity of the area, which will be useful for conservation planning. A method for predicting variety is provided. As shown in FIG. 1, the biodiversity prediction method provided by the embodiment of the present application includes steps S101 to S Including 104. S101, obtaining ecological information of a region to be predicted within a certain period of time; Here, the ecological information of the area to be predicted within a certain period includes multiple sets of ecological data. Each set of ecological data is made up of environmental variables and biological site distribution data. Environmental variable data includes climate factors, topographical factors, habitat factors, and disturbance factors. Climate factors include data such as precipitation data and temperature data. Topographic factors include data such as elevation data. Habitat factors include vegetation data, etc. Disturbance factors include population density, etc. Biosite distribution data includes species name, population size, and location of discovery. Of course, the period may be three years or ten years (for example, January 1, 2012). The period may be from January to December 2022, or any fixed period between 3 and 10 years. The area to be predicted may be an ecological protection area or an ecological restoration area. or any other area where spatial patterns of biodiversity need to be determined. In the examples of the present application, the period and the area to be predicted are further limited. It is not something that can be done. Optionally, the organism may be a bird, a mammal, a reptile, or an amphibian. For example, in the case of birds, the species name above is the house sparrow (Passer domesticus ), tree sparrow (Passermontanus) and stone sparrow (Petronia p etronia), etc. Therefore, in the case of tree sparrows, the biological site distribution data are Includes tree sparrow species name, tree sparrow population size and where to find tree sparrows. In a certain application scenario, the specific parameters of the above environmental variable data are shown in Table 1. In Table 1, precipitation heterogeneity refers to the uneven distribution of precipitation in time and space. Elevational heterogeneity refers to the differences in environmental resources, ecological processes, and community characteristics exhibited at different elevations. The Normalized Difference Vegetation Index (NDVI) is a method for quantifying vegetation cover and health. The canopy height is the height of the top of the tree crown relative to the ground. Taste. Table 1: Environmental Variable Data Table TIFF0007739690000007.tif103145 S102, based on the species completeness index and species accumulation rate, the ecology of the area to be predicted within a certain period Screening the information, obtaining the screened ecological information, and The area corresponding to the screened ecological information is determined as the observed area, and the area to be predicted is determined as the observed area. The area other than the observed area in the region is determined as an unobserved area. Naturally, after dividing the observed area and the area to be predicted, the prediction within a certain period is The ecological information of the area to be monitored is also the ecological information of the area that has been observed within a certain period and the ecological information of the area that has not been observed within a certain period. The ecological information of the area is divided into multiple sets of ecological data. The ecological information of the unobserved area within a certain period includes multiple sets of ecological data. Optionally, as shown in FIG. 2 in conjunction with FIG. 1, S102 includes S1021 to S1024. . S1021, divide the area to be predicted into multiple blocks. In one application scenario, the area to be forecasted is divided into multiple grids (corresponding to the above multiple blocks). The grid map of the area to be predicted is then obtained. Preprocessing of environmental variable data in the information (projection to the same coordinate system, resampling to the same resolution) , including cutting to the size of the area to be predicted) to generate auxiliary data, and The data is then linked to the grid at the corresponding location and environmental variables corresponding to the location of each grid are acquired. You'll benefit. S1022, screening a plurality of first blocks from a plurality of blocks. The first block above meets the following conditions: The discovery location falls within the geographical range of the first block. ,There is at least one set of ecological data in the ,ecological information of the area to be predicted within a certain period. S1023, for each of the first blocks among the plurality of first blocks, The device is associated with the ecological data that falls within the geographical range of each first block, and the data associated with each first block The species completeness index and the species distribution index of each first block are calculated from the biological site distribution data in the attached ecological data. Obtain the species and species accumulation rate. In one embodiment, the formula for calculating the seed completeness index for each first block is: TIFF0007739690000008.tif1853 where, TIFF0007739690000009.tif33 shows the species completeness index for each first block, TIFF0007739690000010.tif36 shows the number of species in each first block, TIFF0007739690000011.tif47 shows the species richness index, TIFF0007739690000012.tif43 shows the number of species with an individual population of 1 in each first block, TIFF0007739690000013.tif43 shows the number of species with an individual population of 2 in each first block. For example, if a creature is a bird, it will be associated with this first block. The ecological data obtained was: Species: House sparrow, Number of individuals: 1, Location of discovery: Left side of first block Placement}, {Species: Stone Sparrow, Number of individuals: 3, Found location: Center of the first block}, {Species: Tree Sparrow, Number of individuals: 1, Found location: Right side of the first block}, {Species name: House sparrow, Individual Number: 1, Discovery location: Right position in the first block}. And the seed of this first block When calculating the completeness index, species within this first block (house sparrow, stone sparrow, tree sparrow) Number of TIFF0007739690000014.tif36=3, and the number of species (tree sparrows) with a population of 1 in this first block TIFF0007739690000015.tif43=1, and the number of species (house sparrows) with a population of 2 in this first block TIFF0007739690000016.tif43=1, and the species completeness index of this first block TIFF0007739690000017.tif33=1. The species accumulation rate is the slope of the terminal cumulative curve of the number of species in one block over a certain period of time. The accumulation curve shows the accumulation of the number of species over a certain period of time as the number of biotic site distribution data increases. Optionally, the abscissa of this cumulative curve is the curve formed by the first block. The number of linked biological site distribution data, and the ordinate is the number of species in the first block. This cumulative curve is a technical means commonly used in the technical field, and therefore, In the examples, the specific process of constructing the cumulative curve is not detailed. S1024, screening the plurality of second blocks from the plurality of first blocks. The second block satisfies the following condition: the species completeness index of the second block is greater than the first threshold. The seed accumulation rate of the second block is smaller than the second threshold. Here, the range of the first threshold value is is [0.85, 1], and the range of values for the second threshold is [0, 1]. Optionally, the first threshold may be any value in [0.85, 1], and the second threshold The value can be any value in [0, 1], and the values of the first and second thresholds are within a reasonable range. The range may be arbitrarily selected, and is not further limited in the examples of the present application. For one second block, the species completeness index of this second block is higher than the first threshold. If the species accumulation rate is less than the second threshold, the ecosystem associated with this second block is The data integrity of the data can be considered high, and if so, the data associated with this second block The reliability of the ecological data collected is high, allowing a better description of the biodiversity of this block. Note that this means that S1025, the biological data associated with multiple second blocks are screened and The second blocks are merged to obtain the observed area within the area to be predicted. Blocks other than the observed area within the area to be predicted are merged to obtain the unobserved area. Naturally, the data integrity of the biometric data associated with multiple second blocks is high. If not, multiple secondary blocks of associated biodata (i.e., screened The ecological information obtained is highly reliable and allows for a better description of the biodiversity of this block. This means that it is necessary to observe, predict, or supplement biological site distribution data from multiple second blocks. Therefore, the area merged by multiple second blocks is defined as the observed area. Please understand that blocks outside the observed area within the area to be predicted are not reliable biological data. Due to a lack of site distribution data, areas outside the observed area within the region to be predicted are left unobserved. The measurement area is defined as S103, based on a biodiversity prediction model and environmental variable data for unobserved areas within a certain period. ,Determine a preliminary prediction of biological site distribution data in unobserved areas within a certain period. Here, the biodiversity prediction model is a random forest model, and the random forest The model includes multiple decision trees. The biodiversity prediction model uses screened ecological information. This is what is obtained by training using Specifically, the above step S103 includes the following: analyzing the environmental variable data of unobserved areas within a certain period of time; The biodiversity prediction model calculates the biodiversity of unobserved areas within a certain period of time. It outputs preliminary predictions for the distributed data. Optionally, the training process of the above biodiversity prediction model is as follows. Step 1: From the ecological information obtained through screening, multiple sets of ecological data are used as a training set. randomly selected as the target. The ecological data in each group in the training set are all environmental variable data and environmental variable data. Includes corresponding biological site distribution data. Step 2: Build a random forest model. The input for the random forest model is The environmental variable data and the output of the random forest model are preliminary biological site distribution data. This is a predicted value. Step 3: Train a random forest model using the training set. The random forest model is used as a biodiversity prediction model. In the above application scenario, the training process of the random forest model (step 3 above) The corresponding values are as follows: Step 3.1, randomly select a subsample set. Using putback, we randomly select some samples from the training set several times to generate multiple new samples. form new subsample sets, and each sample set in the multiple subsample sets is Both are biological site distribution data for the observed area within a certain period, and the above biological site distribution Contains environmental variable data corresponding to the data. Step 3.2, randomly select a feature subset. Use putback to randomly select one feature subset from the total number of features in the subsample set. The size of this subset is smaller than the total number of features. Illustratively, for one feature subset, multiple environmental variables (e.g., temperature, precipitation) Then, data corresponding to multiple environmental variables in one subsample set are Select them as samples, then build and train a decision tree based on the samples. Step 3.3, Train the decision tree. The subsample set and feature subset selected in step 3.1 and step 3.2 We use the data to train one decision tree in the random forest model. A root node, several internal nodes connected to the root node, and several internal nodes It contains multiple connected sub-nodes, where the root node is the sum of all the sub-samples in the sub-sample set. The subnodes contain all the data, and the subnodes contain selected subset data based on the root node. The leaf subnodes contain the prediction results of the decision tree. During the subsequent growth process of the decision tree, each subnode A node selects split features in addition to those already selected by its parent node until it reaches a stopping condition. Continue selecting the best features from the remaining features for Step 3.4: Repeat step 3.2 and step 3.3, random forest Training of all decision trees in the model is completed to obtain a biodiversity prediction model. Step 3.5: Regression prediction. The biodiversity prediction model obtains prediction results by averaging the prediction results of each decision tree. . For example, the above prediction results TIFF0007739690000018.tif43 satisfies the following formula: TIFF0007739690000019.tif1326 where, TIFF0007739690000020.tif33 shows the number of decision trees in the random forest model. TIFF0007739690000021.tif32 indicates the decision tree number, TIFF0007739690000022.tif46 is the number The output result of the decision tree for TIFF0007739690000023.tif32 is shown below. For example, if the biological species is a bird, environmental variables in an unobserved area within a certain period of time can be randomly collected. The inputs are fed into multiple decision trees in the dam forest model, and the multiple decision trees are then fed into the input environmental variables. Based on the data, leaf subnodes are matched and multiple decision tree output results are obtained. The output results of multiple decision trees are averaged to finally obtain the prediction result of the biodiversity prediction model, i.e. In other words, a preliminary prediction of biological site distribution data in unobserved areas within a certain period of time is obtained. S104, biological site distribution data for observed areas within a certain period and unobserved areas within a certain period Based on the preliminary predicted values of the biological site distribution data of the area, the biological Determine the diversity prediction results. As shown in FIG. 3 in conjunction with FIG. 2, the above S104 may be replaced with S1041 to S1042. include. S1041, Preliminary estimates of biological site distribution data in unobserved areas and biodiversity in unobserved areas The sum of the prediction residuals is taken as the final predicted value of the biological site distribution data for unobserved areas within a certain period. In S1041 above, the method for obtaining biodiversity prediction residuals for unobserved areas is as follows. Step 1: From the ecological information obtained through screening, multiple sets of ecological data are subjected to a residual test. Randomly select a set of The ecological data in each group of the residual test set are all environmental variable data and environmental variable data. The data includes biological site distribution data corresponding to the data. Step 2: Input the environmental variable data in the residual test set into the biodiversity prediction model and generate The biodiversity prediction model outputs preliminary predictions of biosite distribution data. Step 3: Next, the biological site distribution data corresponding to the environmental variable data in the residual test set , and the difference between the preliminary predictions of biological site distribution data output from the biodiversity prediction model Let min be the prediction residual for the observed region. Step 4: Kriging interpolation is performed on the predicted residuals of the observed area to obtain the predicted residuals for the area to be predicted. do. Note that Kriging interpolation, which is an interpolation method that takes spatial correlation into consideration, is used in the examples of this application. Note that the spatial correlation requirement can be met during the process of interpolating the predicted residuals of the surveyed area. I want to be done that. Optionally, the predicted residuals are interpolated using Kriging interpolation. TIFF0007739690000024.tif47 satisfies the following formula: TIFF0007739690000025.tif1332 where, TIFF0007739690000026.tif32 indicates the index variable for traversing all sampling points, TIFF0007739690000027.tif33 indicates the number of sampling points, TIFF0007739690000028.tif43 shows the weighting coefficients and the prediction residual at the point (x0, y0) TIFF0007739690000029.tif47 and residual observations The point (x0, y0) is the optimum coefficient that satisfies the requirement to minimize the difference between the above predictions. It is a coordinate point in a two-dimensional coordinate system established based on a grid map of the area. The two-dimensional coordinate system includes the x-axis and the y-axis, and the origin of the two-dimensional coordinate system is the grid map of the area to be predicted. This may be the center of the grid or any other location in the grid map of the area to be predicted. The embodiments of the present application are not limited to a two-dimensional coordinate system. Step 5: Remove the predicted residuals for the observed areas from the predicted residuals for the area to be predicted, and We obtain the biodiversity prediction residuals of S1042, biological site distribution data for observed areas within a certain period and unobserved areas within a certain period The final predicted values of the biological site distribution data of the area are combined to predict the area to be predicted within a certain period. Obtain biodiversity prediction results. In one application scenario of the above method, if the species is birds, for example, Figure 4 shows the final The spatial distribution pattern grid map of the entire biodiversity within the area to be predicted is shown in Fig. Each grid in Figure 4 corresponds to one color, and each color indicates the number of bird species found in the grid. vinegar. In one application scenario, grid maps of the area to be forecasted at three resolutions (grid size The sizes are 10km x 10km, 50km x 50km, and 100km x 100km. In each case, we trained a random forest model and the resulting random forest Test the model (i.e., biodiversity prediction model). For each situation, the training and testing process of the random forest model is as follows: The screened ecological information obtained in S102 above was used as the training set and the test set in a 8:2 ratio. The random forest model (i.e., After obtaining the biodiversity prediction model, we use the site distribution data (such as the number of species) in the test set. ) and the predicted results of the biodiversity prediction model are used as the prediction residual for the observed area, and kriging interpolation is performed. The prediction residuals for the unobserved areas are obtained through this, and the prediction residuals for the entire area to be predicted are thereby obtained. Then, the environmental variable data in the test set are input into the biodiversity prediction model, and random fields are used. The sum of the forest model prediction results and the corresponding prediction residuals, and the environmental variable data in the test set Perform Pearson correlation analysis on the biological site distribution data (such as number of species) corresponding to the data. At the same time, the above three situations (first situation: grid size is 10km x 10km, second situation: 1st situation: grid size is 50km x 50km, 2nd situation: grid size is 100km x 1 For each situation within 00km, the Maxent distribution model and the conventional expert Biodiversity estimates were calculated according to the same resolution and area using the map distribution surface overlay method. (i.e., predict the number of species in unobserved regions in the test set from the data in the training set) ), Pearson correlation of the corresponding biosite distribution data (i.e., number of species) in the test set Conduct analysis. The above Maxent species distribution model and the conventional expert distribution surface overlay method are The specific implementation process of the above two algorithms is described in the embodiments of this application. I won't elaborate further on this. The combination of the random forest model and Kriging interpolation provided by the examples of this application Estimation by combining the Maxent distribution model and the conventional expert distribution surface superposition method The correlation results between the biological site distribution data obtained and the actual situation are shown in Tables 2, 3 and 4 below. Shown in Figure 4. Table 2: Comparison of correlation test results for the first situation TIFF0007739690000031.tif97156 Table 3: Comparison of correlation test results for the second situation TIFF0007739690000032.tif98160 Table 4: Comparison of correlation test results for the third situation TIFF0007739690000033.tif108156 Naturally, the higher the correlation obtained from the Pearson correlation analysis, i.e., the more accurate the predicted value, the better the accuracy of the analysis. The closer the value is to the model, the higher the accuracy of the model. Generally, correlation |r|>0.7 A strong correlation is observed, 0.4<|r|≦0.7 is a moderate correlation, and |r|≦0.4 is a weak correlation. As can be seen from Tables 2, 3 and 4 above, the biodiversity forecasts provided by the present application The correlation between the final prediction results obtained by combining the model and the Kriging interpolation method and the true value was It is higher than the conventional expert distribution surface superposition method, and in the first situation, the Maxent distribution model The correlation of the model was higher than that of the combination of the biodiversity prediction model and the Kriging interpolation method, and the second situation And in the third situation, the correlation of the Maxent distribution model is higher than that of the random forest model. The biodiversity prediction provided by this application is therefore lower than the combined quantification interpolation method. The combination of the observation model and kriging interpolation predicts the biodiversity of the entire area to be predicted. It is closer to the actual situation and has higher accuracy.
Claims
1. A method for predicting biodiversity by a computer, comprising: A step of acquiring ecological information of an area to be predicted within a certain period of time; Here, the ecological information of the area to be predicted within a certain period includes a plurality of sets of ecological data, Each set of ecological data in the multiple sets of ecological data includes environmental variable data and biological site data. The environmental variable data includes climate factors, topographical factors, habitat factors, and disturbance factors. the climate factors include rainfall data and temperature data, and the topographical factors include elevation data. the habitat factors include vegetation data, the disturbance factors include population density, and the biological Site distribution data includes species name, number of individuals, and location of discovery. Based on the species completeness index and species accumulation rate, the ecological information of the area to be predicted within the certain period is calculated. Screening is performed, and the screened ecological information is acquired. The area corresponding to the screened biological information is determined as the observed area, and the area to be predicted is determined as the observed area. determining an area other than the observed area in the region as an unobserved area; Based on the biodiversity prediction model and environmental variable data of the unobserved area within a certain period, determining a preliminary prediction of biological site distribution data in the unobserved area within a time period; Here, the biodiversity prediction model is a random forest model, and the biodiversity a prediction model is trained using the screened biological information; The biological site distribution data of the observed area within the certain period and the biological site distribution data of the unobserved area within the certain period Based on the preliminary predicted value of biological site distribution data in the observation area, the predicted value within a certain period is determining a regional biodiversity prediction result; A method for predicting biodiversity by a computer, comprising:
2. Based on the species completeness index and species accumulation rate, the ecological situation of the area to be predicted within the certain period is calculated. The information is screened, the screened ecological information is acquired, and the information is applied to the area to be predicted. The area corresponding to the screened biological information in the The step of determining an area other than the observed area in the target area as an unobserved area includes the following steps: Includes: Dividing the area to be predicted into a plurality of blocks; screening a plurality of first blocks from the plurality of blocks, the first blocks being: The following condition is met: the discovery location falls within the geographical range of the first block within the given period There is at least one set of ecological data in the ecological information of the area to be predicted; For each of the plurality of first blocks, a discovery position of the first block is Associating each of the first blocks with biological data falling within the geographic range of the first block, From the biological site distribution data in the associated ecological data, the species completeness of each of the first blocks is determined. Calculate the sex index and species accumulation rate. screening a plurality of second blocks from the plurality of first blocks; satisfies the following condition: the seed completeness index of the second block is greater than a first threshold; The seed accumulation rate of the second block is less than a second threshold value, where the value of the first threshold value is The range is [0.85, 1], and the value of the second threshold is in the range [0, 1]; The biometric data associated with the plurality of second blocks is filtered out. The second blocks are used as information, and the second blocks are merged to obtain the observed area within the region to be predicted. and merging the blocks other than the observed area in the area to be predicted, and 2. The method of claim 1, further comprising obtaining a region.
3. Based on the biodiversity prediction model and environmental variable data for the unobserved area within a certain period, A step for determining a preliminary predicted value of biological site distribution data in the unobserved area within a certain period of time. The pack includes: inputting environmental variable data of the unobserved area within the certain period into the biodiversity prediction model; The biodiversity prediction model calculates biological site distribution data of the unobserved area within the certain period.
2. The method of claim 1, further comprising: outputting a preliminary estimate of
4. The step of determining the biodiversity prediction result for the region to be predicted within the certain period of time includes the following steps: Includes: Preliminary predicted values of biological site distribution data in the unobserved area and biodiversity prediction in the unobserved area The sum of the residuals is used as the final predicted value of the biological site distribution data of the unobserved area within the given period. 、 Next, the biological site distribution data of the observed area within the certain period and the previous data within the certain period The final predicted values of the biological site distribution data in the unobserved area are combined, and The method according to claim 1, characterized in that a biodiversity prediction result is obtained for an area to be predicted.
5. The method for obtaining the biodiversity prediction residuals in the unobserved area is as follows: From the biological information obtained by the screening, multiple sets of biological data are classified into a residual test set. The residual test set is randomly selected as the residual test set. and biosite distribution data corresponding to the number data and the environmental variable data; The environmental variable data in the residual test set is input to the biodiversity prediction model, and the resulting The biodiversity prediction model outputs preliminary predictions of biological site distribution data, Next, biological site distribution data corresponding to the environmental variable data in the residual test set; and between the preliminary predicted values of biological site distribution data output from the biodiversity prediction model. The difference is the prediction residual of the observed region, Kriging interpolation is performed on the predicted residuals of the observed region to obtain predicted residuals for the region to be predicted; The prediction residual of the observed region is removed from the prediction residual of the region to be predicted, and the prediction residual of the unobserved region is removed. The method according to claim 4, characterized in that a biodiversity prediction residual of
6. The predetermined period is 3 to 10 years, and the area to be predicted is an ecological reserve.
2. The method according to claim 1.
Citation Information
Patent Citations
Method for predicting proper habitat of ginkgo fruit forest based on climate and soil factors
CN112749834A
Historical cultivated land distribution reconstruction method based on random forest model
CN113902580A
Geochemical variable space prediction method based on geostatistical weighted random forest
CN114139819A
Modification method for environmental preservation area
JP2002247900A
Environmental assessment system
JP2004102606A