A method and apparatus for biodiversity prediction

By using random forest models and Kriging interpolation, the problem of scattered distribution of biodiversity data was solved, providing an overall spatial distribution pattern of biodiversity in the study area and supporting more effective conservation planning.

CN119903337BActive Publication Date: 2025-11-18INST OF ZOOLOGY GUANGDONG ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411651270.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-19
Publication Date
2025-11-18
Estimated Expiration
2044-11-19

AI Technical Summary

Technical Problem

The observation results of biodiversity data in existing technologies are scattered and locally distributed, which cannot provide the overall spatial distribution pattern of biodiversity in the study area, making conservation planning difficult.

Method used

By acquiring environmental variables and biodiversity data of the study area, a random forest model was used to predict the unobserved areas. The sample set was selected by combining species integrity and the slope of the end of the species accumulation curve, and the prediction residuals were processed by Kriging interpolation to obtain the biodiversity data of the unobserved areas.

Benefits of technology

It enables the prediction of the spatial distribution pattern of biodiversity in the study area, supporting more effective conservation planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119903337B_ABST
    Figure CN119903337B_ABST
Patent Text Reader

Abstract

The application provides a biodiversity prediction method and device, belongs to the field of biotechnology, can predict the biodiversity of unobserved areas in a research area, gives the spatial distribution pattern of the overall biodiversity of the research area, and thus facilitates the protection planning of the research area. The method screens biodiversity data in the research area, trains a random forest model using the screened biodiversity data (biodiversity data of observed areas), predicts environmental variable data of unobserved areas through the random forest model, obtains biodiversity prediction data of the unobserved areas, and then obtains biodiversity data of the unobserved areas according to the biodiversity prediction data of the unobserved areas and corresponding biodiversity prediction residuals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biotechnology, and more particularly to a method and apparatus for predicting biodiversity. Background Technology

[0002] The spatial distribution pattern of biodiversity is one of the core scientific questions in ecology, biogeography, and conservation biology; it is also the foundation for conservation planning such as priority areas, ecological corridors, and ecological restoration. Currently, my country is conducting extensive biodiversity monitoring surveys, using artificial transects and sampling points, infrared cameras, and equipment such as acoustic or video data acquisition, accumulating a wealth of detailed basic data on species distribution.

[0003] However, these survey data are often scattered and localized, failing to provide a spatial distribution pattern of biodiversity in the overall study area, which is detrimental to conservation planning in the study area. Summary of the Invention

[0004] This invention proposes a biodiversity prediction method and apparatus, which can predict the biodiversity of unobserved areas in a study area and provide the spatial distribution pattern of biodiversity in the overall study area, thereby facilitating conservation planning in the study area.

[0005] To achieve the above objectives, the present invention adopts the following technical solution:

[0006] Firstly, this invention provides a biodiversity prediction method, comprising: acquiring environmental variable data and biodiversity data of a study area over a period of time; the biodiversity data includes: species name, number of individuals, and location of discovery; the environmental variable data includes: climate factors, topographic factors, habitat factors, and disturbance factors. Based on species integrity and the slope of the end of the species accumulation curve, the biodiversity data is screened and a sample set is determined; within the study area, the portion of the selected biodiversity data containing the discovery locations is considered an observed area, otherwise it is considered an unobserved area. A random forest model is used to predict the environmental variable data of the unobserved areas over a period of time, obtaining the predicted biodiversity data for the unobserved areas over that period of time; the random forest model is trained using the sample set. Then, based on the predicted biodiversity data and the corresponding biodiversity prediction residuals, the biodiversity data for the unobserved areas over a period of time is determined; the biodiversity prediction residuals are obtained by interpolating the prediction residuals of the random forest model.

[0007] The biodiversity prediction method provided by this invention screens biodiversity data within a study area and trains a random forest model using the screened biodiversity data (data from observed areas). This random forest model then predicts environmental variables in unobserved areas, yielding predicted biodiversity data for these unobserved areas. Finally, based on the predicted biodiversity data and the corresponding residuals, the predicted biodiversity data for the unobserved areas is obtained. Thus, this method provides both the biodiversity data from observed and unobserved areas. Therefore, this invention can predict biodiversity in unobserved areas within a study area, providing a spatial distribution pattern of biodiversity across the entire study area, thereby facilitating conservation planning for the study area.

[0008] In one implementation of the first aspect, biodiversity data is screened and a sample set for the random forest model is determined based on species integrity and the slope of the end of the species accumulation curve, including:

[0009] The study area was divided into multiple blocks;

[0010] For each of the multiple blocks, biodiversity data of the geographic range where the location falls into the block is associated with the block; and the species integrity and the slope of the end of the species accumulation curve of the block are calculated using the biodiversity data associated with the block.

[0011] The sample set consists of biodiversity data from multiple blocks where species integrity is greater than the first threshold and the slope of the species accumulation curve at the end of the second threshold is less than the second threshold, as well as the environmental variable data corresponding to the biodiversity data.

[0012] Blocks associated with biodiversity data in the sample set are classified as observed areas, while remaining blocks in the study area are classified as unobserved areas.

[0013] In one implementation of the first aspect, the training method for the random forest model is as follows:

[0014] The random forest model is trained using a training set; the input to the random forest model is environmental variable data, and the output of the random forest model is biodiversity prediction data; the sample set includes the training set and the test set.

[0015] In one implementation of the first aspect, the method for determining the biodiversity prediction residuals is as follows:

[0016] The trained random forest model is tested on a test set, and the difference between the actual biodiversity data in the test set and the biodiversity prediction data output by the trained random forest model is used as the prediction residual for the observed area; the sample set includes the training set and the test set.

[0017] Kriging interpolation is performed on the prediction residuals of the observed areas to obtain the biodiversity prediction residuals of the study area.

[0018] Secondly, this invention provides a biodiversity prediction device, comprising an acquisition module, a screening module, a prediction module, and a determination module. The acquisition module acquires environmental variable data and biodiversity data for a study area over a period of time; the biodiversity data includes species names, individual numbers, and discovery locations. The screening module screens the biodiversity data and determines a sample set based on species integrity and the slope of the end of the species accumulation curve; within the study area, the portion of the discovery locations in the screened biodiversity data that fall within the observed area is considered an observed area; otherwise, it is considered an unobserved area. The prediction module uses a random forest model to predict the environmental variable data for unobserved areas over a period of time, obtaining predicted biodiversity data for unobserved areas over that period; the random forest model is trained using the sample set. The determination module determines the biodiversity data for unobserved areas over a period of time based on the predicted biodiversity data and the corresponding biodiversity prediction residuals; the biodiversity prediction residuals are obtained by interpolating the prediction residuals of the random forest model.

[0019] In one implementation of the second aspect, the screening module is specifically used to divide the study area into multiple blocks. For each block, biodiversity data of the geographical area where the discovery location falls within the block is associated with the block; and the species integrity and the slope of the end of the species accumulation curve for the block are calculated using the biodiversity data associated with the block. Blocks with species integrity greater than a first threshold and a slope of the end of the species accumulation curve less than a second threshold are classified as observed areas, and the remaining blocks are classified as unobserved areas. The biodiversity data associated with the blocks classified as observed areas and the corresponding environmental variable data are used as a sample set.

[0020] In one implementation of the first and second aspects, the organisms include birds.

[0021] In one implementation of the first and second aspects, the formula for calculating species integrity is as follows;

[0022]

[0023] Where SI represents species completeness, Sobs represents the number of species observed in the sample, Chao1 represents the richness index, n1 represents the number of species containing only one individual, and n2 represents the number of species containing only two individuals.

[0024] The slope at the end of the species accumulation curve is calculated from the cumulative curve of the number of species in the block over a period of time.

[0025] Thirdly, the present invention provides an electronic device including a processor and a memory coupled to the processor; the memory is used to store computer instructions, and when the electronic device is running, the processor executes the computer instructions stored in the memory to cause the electronic device to perform the method as described in the first aspect above or any implementation thereof.

[0026] Fourthly, the present invention provides a computer-readable storage medium including computer program instructions that, when executed by a computer, cause the computer to perform the method as described in the first aspect above or any implementation thereof.

[0027] Fifthly, the present invention provides a computer program product, including computer program instructions, which, when executed on a computer, cause the computer to perform the method as described in the first aspect above or any of its implementations.

[0028] The technical effects corresponding to the second to fifth aspects and their possible implementations can be referred to the above description of the technical effects of the first aspect and its possible implementations, and will not be repeated here. Attached Figure Description

[0029] Figure 1 This is one of the schematic diagrams of the biodiversity prediction method provided in the embodiments of this application;

[0030] Figure 2 This is the second schematic diagram of the biodiversity prediction method provided in the embodiments of this application;

[0031] Figure 3 This is the third schematic diagram of the biodiversity prediction method provided in the embodiments of this application;

[0032] Figure 4 This is the fourth schematic diagram of the biodiversity prediction method provided in the embodiments of this application;

[0033] Figure 5 This is the fifth schematic diagram of the biodiversity prediction method provided in the embodiments of this application;

[0034] Figure 6 This is a schematic diagram of the biodiversity distribution grid provided in an embodiment of this application;

[0035] Figure 7 This is one of the structural schematic diagrams of the biodiversity prediction device provided in the embodiments of this application. Detailed Implementation

[0036] In the specification and claims of this invention, the terms "first" and "second," etc., are used to distinguish different objects, rather than to describe a specific order of objects.

[0037] In the embodiments of this application, "and / or" indicates a relationship between objects. For example, A and / or B can represent the following three situations: A exists alone, B exists alone, and A and B exist simultaneously.

[0038] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0039] In the description of this invention, unless otherwise stated, "a plurality of" means two or more. For example, a plurality of blocks means two or more blocks.

[0040] The methods and apparatus provided in this application relate to the spatial distribution patterns of biodiversity. The embodiments of this application use observed environmental variable data and biodiversity data in the study area to predict unobserved biodiversity data in the study area, thereby obtaining the overall spatial distribution pattern of biodiversity in the study area.

[0041] As is understood, biodiversity refers to the diversity of all biological species, genes, and ecosystems on Earth, and protecting biodiversity is of great significance for maintaining the Earth's ecological balance. Biodiversity includes genetic diversity, species diversity, and ecosystem diversity; in this embodiment, biodiversity refers to species diversity.

[0042] In existing technologies, biodiversity data is often obtained through human observation. Human observation is characterized by randomness, high uncertainty, and large errors, resulting in scattered, localized point-like distributions of biodiversity data. This fails to provide a comprehensive spatial distribution pattern of biodiversity in the study area, hindering conservation planning. To address these issues, this application provides a biodiversity prediction method and apparatus. The method involves screening biodiversity data within the study area and training a random forest model using the screened data (biodiversity data from observed areas). The random forest model then predicts environmental variables in unobserved areas, yielding predicted biodiversity data for these unobserved areas. Finally, based on the predicted biodiversity data and the corresponding residuals, the predicted biodiversity data for the unobserved areas is obtained. Thus, both observed and unobserved biodiversity data are obtained. Therefore, this application can predict biodiversity in unobserved areas within a study area, providing a comprehensive spatial distribution pattern of biodiversity in the study area, which is beneficial for conservation planning.

[0043] For example, the biodiversity prediction method provided in this embodiment of the invention can be executed by an electronic device with processing capabilities, such as a computer or server. Taking a computer as an example, the hardware components of the computer may include: a processor, memory, a network interface, a user interface, a communication bus, etc.

[0044] The processor is used to control electronic devices to perform related processing and calculation tasks. The processor may include a central processing unit (CPU) or other processors. The processor may be single-core or multi-core. For example, the processor may include multiple CPUs.

[0045] Memory is used to store computer instructions and related data, such as environmental variable data, biodiversity data, sample sets, biodiversity prediction data, and biodiversity prediction residuals. Memory can be random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical storage, disk storage media, or other magnetic storage devices, or any other medium capable of storing program code or data accessible by a computer. Optionally, memory can be integrated into the processor, or it can be independent of the processor.

[0046] A network interface is used for communication between a computer and other devices or communication networks. A network interface can be a transceiver with transmit and receive capabilities. Optionally, a network interface may include standard wired interfaces or wireless interfaces (such as Wi-Fi interfaces, Bluetooth interfaces, and 5G interfaces).

[0047] The communication bus is used to enable communication between different components. For example, the processor, memory, network interface and user interface mentioned above can be interconnected through the communication bus.

[0048] The user interface may include a display screen and an input unit (such as a keyboard). Optionally, the user interface may also include a standard wired interface or a wireless interface.

[0049] Those skilled in the art will understand that the computer described above may include more or fewer components, or combine certain components, or have different component arrangements; the embodiments of this application do not limit this.

[0050] like Figure 1 As shown, the biodiversity prediction method provided in this application includes S101-S104.

[0051] S101. Obtain environmental variable data and biodiversity data for the study area over a period of time.

[0052] The biodiversity data includes: species name, number of individuals, and location of discovery; the environmental variable data includes: climate factors, topographic factors, habitat factors, and disturbance factors.

[0053] The aforementioned organisms can be birds, mammals, reptiles, or amphibians; this application does not limit the types of organisms described.

[0054] Optionally, the specific parameters and acquisition methods of the above environmental variable data are shown in Table 1.

[0055] Table 1

[0056]

[0057]

[0058] In one application scenario, taking birds as an example, biodiversity data spontaneously observed by bird enthusiasts / workers can be obtained from the China Birdwatching Record Center.

[0059] S102. Based on species integrity and the slope of the end of the species accumulation curve, screen biodiversity data and determine the sample set; in the study area, the part where the discovery location is found in the screened biodiversity data is the observed area, otherwise it is the unobserved area.

[0060] Optionally, combined Figure 1 ,like Figure 2 As shown, S102 includes S1021-S1024.

[0061] S1021. Divide the study area into multiple blocks.

[0062] S1022. For each of the multiple blocks, associate the biodiversity data of the geographic range where the discovery location falls into the block with the block; and calculate the species integrity and the slope of the end of the species accumulation curve of the block using the biodiversity data associated with the block.

[0063] In one implementation, the formula for calculating species integrity is as follows;

[0064]

[0065] Where SI represents species completeness, Sobs represents the number of species observed in the sample, Chao1 represents the richness index, n1 represents the number of species containing only one individual, and n2 represents the number of species containing only two individuals.

[0066] The slope at the end of the aforementioned species accumulation curve is calculated from the cumulative curve of the number of species in the block over a period of time. The specific calculation method is a common technique in this technical field, and will not be elaborated here in the embodiments of this application.

[0067] S1023. Collect biodiversity data from multiple blocks where species integrity is greater than the first threshold and the slope of the species accumulation curve at the end is less than the second threshold, as well as environmental variable data corresponding to the biodiversity data, as a sample set.

[0068] Optionally, the first threshold can be 0.85, and the second threshold can be 1. The values ​​of the first threshold and the second threshold can be selected within a reasonable range, and the embodiments of this application do not impose further limitations.

[0069] S1024. The blocks associated with the biodiversity data in the sample set are classified as observed areas, and the remaining blocks in the study area are classified as unobserved areas.

[0070] In one application scenario, taking birds as an example, the biodiversity data spontaneously observed by bird enthusiasts / workers may contain data from areas with insufficient observation data, resulting in a number of bird species recorded that is far less than the actual number of species in those areas. Therefore, the above-mentioned screening method is needed to exclude the insufficiently observed biodiversity data to ensure that the data in the sample set used to train the random forest model is closer to the real situation, thereby improving the predictive ability of the random forest model.

[0071] In the above application scenario, the study area is divided into multiple grids (i.e., the aforementioned blocks), resulting in a grid map of the study area. The collected environmental variable data is preprocessed to generate auxiliary data (projected to the same coordinate system, resampled to the same resolution, and cropped to the size of the study area). The processed environmental variable data is then extracted into the corresponding grids, obtaining the environmental variable data for each grid location. Biodiversity data within the geographic area where the location falls into a block are then associated with the grid. Finally, the sample set is obtained through the filtering process described in S1022-S1024.

[0072] In one implementation, combined with Figure 2 ,like Figure 3 As shown, before S103, the above method also includes: S105, training a random forest model.

[0073] Optionally, the training method for the random forest model is as follows: the random forest model is trained using a training set; the input of the random forest model is environmental variable data, and the output of the random forest model is biodiversity prediction data; the sample set includes a training set and a test set. For example, 20% of the data in the above sample set can be used as the test set, and 80% of the data in the above sample set can be used as the training set. Alternatively, the sample set can be divided according to other proportions; this embodiment of the application does not limit the scope of the classification.

[0074] In one application scenario, taking birds as an example, the training set includes biodiversity data and environmental variable data of the observed area over a period of time, and the biodiversity data and environmental variable data correspond one-to-one with the grid.

[0075] In the above application scenarios, the training process of the random forest model is as follows.

[0076] Step 1: Randomly select a subsample set.

[0077] Multiple new subsample sets are formed by randomly selecting a portion of samples with replacement from the training set (80% of the grid samples). Each of these subsample sets includes biodiversity data and environmental variable data of the observed area over a period of time.

[0078] Step 2: Randomly select a subset of features.

[0079] During the training of a decision tree, a subset of features is randomly selected with replacement from the total number of features in the subset. The size of this subset is smaller than the total number of features (for example, selecting eight features such as temperature, precipitation, and vegetation index from a total of nine environmental variables as a feature subset).

[0080] For example, for one feature subset, multiple environmental variables (e.g., temperature, precipitation) are selected, and then data corresponding to the multiple environmental variables in a subsample set are selected as samples (e.g., sample 1 has 100 species, temperature of 20, and precipitation of 50; sample 2 has 200 species, temperature of 10, and precipitation of 30). A decision tree is then constructed and trained based on the samples.

[0081] Step 3: Train the decision tree.

[0082] Using the subset of samples and features selected in steps 1 and 2, a decision tree model is trained. The decision tree model consists of a root node and multiple internal nodes (also called decision nodes or split nodes) connected to the root node, and multiple child nodes connected to the internal nodes. The root node contains all data from the subset of samples, the child nodes contain subsets of data selected based on the root node, and the leaf nodes contain the predictions of the decision tree. During the subsequent growth of the decision tree, each child node continues to select the best features from the remaining features beyond those already selected by its parent node (e.g., a second split: selecting seven features from eight as a new feature subset), until a stopping condition is met (e.g., all samples belong to the same category or the number of samples contained in a node is less than a certain threshold).

[0083] Specifically, during the construction of the decision tree, the algorithm recursively selects the optimal feature for splitting. Starting from the root node, based on the selected feature and the corresponding splitting criteria (such as information gain, Gini index, etc.), the dataset is divided into different subsets, each corresponding to a different child node. This process continues, constantly generating new child nodes, until a stopping condition is met (such as the number of samples in the node being less than a certain threshold, the node purity reaching a certain standard, or the tree depth reaching a preset maximum). Finally, this node is considered a leaf node, and it determines a prediction result based on the sample data it contains. The prediction result is based on the statistics (such as mean, median, etc.) of all the sample data in that node.

[0084] Step 4: Repeat steps 2 and 3.

[0085] Repeat steps 2 and 3 multiple times to generate multiple decision trees, thus obtaining the random forest model.

[0086] Step 5: Regression Prediction.

[0087] The random forest model obtains the final prediction result by averaging the prediction results of each tree.

[0088] For example, the above prediction results It satisfies the following formula.

[0089]

[0090] Where N represents the number of decision trees in the random forest model, i represents the decision tree number, and h i (x) represents the output of the decision tree numbered i.

[0091] In another implementation method, combined with Figure 3 ,like Figure 4 As shown, after S105 and before S103, the above method also includes: S106, determining the biodiversity prediction residuals of the study area.

[0092] For example, in combination Figure 4 ,like Figure 5 As shown, S106 includes S1061-S1062.

[0093] S1061. Test the trained random forest model using the test set, and use the difference between the actual biodiversity data in the test set and the biodiversity prediction data output by the trained random forest model as the prediction residual for the observed area.

[0094] S1062. Perform Kriging interpolation on the prediction residuals of the observed area to obtain the biodiversity prediction residuals of the study area.

[0095] It should be noted that interpolation refers to using the value of a function at a certain point to estimate the approximate value of the function at other points. There are various interpolation methods, including linear interpolation, polynomial interpolation, and spline interpolation, each with its specific application scenarios and advantages and disadvantages. Kriging interpolation is an interpolation method that considers the geographical location and spatial correlation of the data. Since this application's embodiment needs to consider spatial correlation when interpolating the prediction residuals of the observed area, the biodiversity prediction residuals of the study area obtained through Kriging interpolation can better reflect the actual situation of the study area, thereby improving the accuracy of the biodiversity prediction residuals of the study area.

[0096] As is understandable, Kriging interpolation is a spatial interpolation method that performs a weighted summation of a certain attribute (in this embodiment, the prediction residual of a random forest model) at several discrete points in space, thereby interpolating the prediction residual to an unobserved location at that location. The prediction residual Z(x0) obtained by the above interpolation satisfies the following formula.

[0097]

[0098] Where i represents the index variable used to iterate through all sampling points, n represents the number of sampling points, and λ iThe weighting coefficients are those that satisfy the condition that the residual estimate Z(x0) at point (x0, y0) is equal to the residual observation Z(x0). i The optimal coefficient with the smallest difference.

[0099] The point (x0, y0) mentioned above is a coordinate point in a two-dimensional coordinate system set based on the grid map of the study area mentioned above; the two-dimensional coordinate system includes x and y axes, and the origin of the two-dimensional coordinate system can be the center position of the grid map of the study area, or it can be another position in the grid map of the study area. The embodiments of this application do not further limit the two-dimensional coordinate system.

[0100] S103. The random forest model predicts environmental variable data for unobserved areas over a period of time, thus obtaining predicted biodiversity data for unobserved areas over a period of time.

[0101] The random forest model is trained using a sample set.

[0102] In one application scenario, taking birds as an example, environmental variable data of unobserved areas over a period of time are input into multiple decision trees in a random forest model. The leaf nodes of the multiple decision trees are matched according to the input environmental variable data, thereby obtaining multiple decision tree output results. The average of the multiple decision tree output results is then calculated to obtain the prediction result of the random forest model, that is, the biodiversity prediction data of unobserved areas over a period of time.

[0103] S104. Based on the biodiversity prediction data and the corresponding biodiversity prediction residuals, determine the biodiversity data of unobserved areas over a period of time.

[0104] The biodiversity prediction residuals are obtained by interpolating the prediction residuals of the random forest model.

[0105] In this embodiment of the application, the biodiversity data of the unobserved area over a period of time is the sum of the predicted biodiversity data of the unobserved area over a period of time and the predicted biodiversity residual.

[0106] In one application scenario, taking birds as an example, Figure 6 This results in a grid map showing the overall spatial distribution pattern of biodiversity within the study area. (Example:) Figure 6 As shown, each grid corresponds to a color, and each color represents the number of bird species found in the grid.

[0107] In summary, the biodiversity prediction method provided in this application involves screening biodiversity data within the study area, training a random forest model using the screened biodiversity data (biodiversity data from observed areas), predicting environmental variables in unobserved areas using the random forest model, obtaining predicted biodiversity data for unobserved areas, and then obtaining the actual biodiversity data for unobserved areas based on the predicted biodiversity data and the corresponding biodiversity prediction residuals. Thus, both the observed and unobserved biodiversity data are obtained. Therefore, this application embodiment can predict biodiversity in unobserved areas within a study area, providing the overall spatial distribution pattern of biodiversity in the study area, thereby facilitating conservation planning for the study area.

[0108] In one application scenario, a random forest model is trained and tested on three different resolution grid maps of the study area (grid sizes of 10km×10km, 50km×50km, and 100km×100km, respectively). The training and testing process of the random forest model is as follows for each case.

[0109] The sample set was divided into a training set and a test set in an 8:2 ratio. After training the random forest model using the training set, the difference between the biodiversity data (i.e., the number of species) in the training set and the prediction results of the random forest model was used as the prediction residual for the observed area. The prediction residual for the unobserved area was obtained by Kriging interpolation, thus obtaining the prediction residual for the entire study area.

[0110] Next, the environmental variable data in the test set were input into the random forest model, and the sum of the prediction results of the random forest model and the corresponding prediction residuals was subjected to Pearson correlation analysis with the biodiversity data (i.e. the number of species) corresponding to the environmental variable data in the test set.

[0111] Meanwhile, for each of the three scenarios mentioned above, the Maxent distribution model and the traditional expert distribution surface overlay method were used respectively to extrapolate biodiversity (i.e., predict the number of species in the unobserved areas of the test set based on the data in the training set) at the same resolution and in the same area, and Pearson correlation analysis was performed with the corresponding biodiversity data (i.e. the number of species) in the test set.

[0112] It's important to note that the Maxent species distribution model is a species distribution prediction model based on the maximum entropy principle. It analyzes known species distribution data and environmental variables, using the maximum entropy principle to infer the probability of a species' distribution in unknown areas. In this case, the Maxent species distribution model runs on Java software. Based on the biodiversity data of a single organism (e.g., a bird species) in a grid within the training set, and combined with environmental variable information related to its location, it predicts the distribution of that organism in unknown areas of the study area, ultimately obtaining a simulated distribution grid map of the species. Therefore, by traversing the biodiversity data of each organism in the training set, the number of species in each grid of the study area's grid map can be predicted. In contrast, traditional expert distribution surfaces are based on existing expert-defined bird distribution ranges, created as a grid, and the species count is directly obtained by overlaying these grids.

[0113] The results of comparing the biodiversity data derived from the random forest model combined with Kriging interpolation, the Maxent distribution model, and the traditional expert distribution surface overlay method with the actual situation are shown in Tables 1, 2, and 3 below.

[0114] Table 1

[0115]

[0116] Table 2

[0117]

[0118] Table 3

[0119]

[0120] Understandably, the higher the correlation obtained from Pearson correlation analysis, the closer the predicted value is to the true value, indicating higher model accuracy. Generally speaking, a correlation |r|>0.7 is considered strong, 0.4<|r|≤0.7 is moderate, and |r|≤0.4 is weak.

[0121] Tables 1, 2, and 3 above show the correlation results between the biodiversity data derived from the random forest model provided in this application, combined with Kriging interpolation, the Maxent distribution model, and the traditional expert distribution surface overlay method, and the actual situation when the resolution of the grid map of the study area is 10km, 50km, and 100km (i.e., grid sizes are 10km×10km, 50km×50km, and 100km×100km, respectively).

[0122] As shown in Tables 1, 2, and 3, when the grid map of the study area has the three resolutions mentioned above, the correlation between the final prediction results of the random forest model combined with Kriging interpolation provided in this application and the actual values ​​is higher than that of the traditional expert distribution surface overlay method. At a resolution of 10km×10km, the correlation of the Maxent distribution model is higher than that of the random forest model combined with Kriging interpolation. At resolutions of 50km×50km and 100km×100km, the correlation of the Maxent distribution model is lower than that of the random forest model combined with Kriging interpolation. It can be seen that at resolutions of 50km×50km and 100km×100km, the final prediction results of the random forest model combined with Kriging interpolation provided in this application are closer to the actual situation, thus proving that the random forest model combined with Kriging interpolation provided in this application has the highest accuracy.

[0123] Building upon this, the training process of the random forest model optimizes model performance by constructing multiple decision trees and integrating their predictions. Each decision tree is built independently, and the model's diversity is increased by randomly selecting features and sample data. During training, the random forest continuously tries different feature combinations and splitting methods to find the optimal prediction model. This training method allows the random forest to capture complex nonlinear relationships in the data. In contrast, linear regression methods describe the relationship between features and the target by fitting a straight line (or plane, hyperplane). If the relationship between biodiversity distribution and independent variables is nonlinear, or if there are complex interactions between independent variables, linear regression may fail to accurately fit this relationship, leading to inaccurate predictions.

[0124] Meanwhile, the random forest model effectively improves its generalization ability by constructing multiple decision trees and integrating their prediction results. In particular, compared with existing linear regression methods, the random forest model provided in this application, combined with Kriging interpolation, is more sensitive to outliers, thus providing relatively accurate predictions even when facing complex natural phenomena such as biodiversity distribution.

[0125] Accordingly, embodiments of this application provide a biodiversity prediction device, such as... Figure 7 As shown, it includes an acquisition module 501, a filtering module 502, a prediction module 503, and a determination module 504.

[0126] The acquisition module 501 is used to acquire environmental variable data and biodiversity data of the study area over a period of time; the biodiversity data includes: species name, number of individuals, and location of discovery. For example, the acquisition module 501 is used to implement S101 of the above-mentioned biodiversity prediction method.

[0127] The screening module 502 is used to screen biodiversity data and determine the sample set based on species integrity and the slope of the end of the species accumulation curve. Within the study area, the portion of the biodiversity data whose locations fall into the screening data is considered an observed area; otherwise, it is considered an unobserved area. For example, the screening module 502 is used to implement S102 of the above-mentioned biodiversity prediction method.

[0128] The prediction module 503 is used to use the random forest model to predict environmental variable data for unobserved areas over a period of time, thereby obtaining biodiversity prediction data for these unobserved areas. The random forest model is trained using a sample set. For example, the prediction module 503 is used to implement step S103 of the aforementioned biodiversity prediction method.

[0129] The determination module 504 is used to determine the biodiversity data of unobserved areas over a period of time based on the biodiversity prediction data and the corresponding biodiversity prediction residuals; the biodiversity prediction residuals are obtained by interpolating the prediction residuals of the random forest model. For example, the determination module 504 is used to implement S104 of the above-mentioned biodiversity prediction method.

[0130] Optionally, the screening module 502 is specifically used to: divide the study area into multiple blocks. For each block, biodiversity data of the geographical area where the discovery location falls within the block is associated with the block; and the species integrity and the slope of the end of the species accumulation curve of the block are calculated using the biodiversity data associated with the block. Blocks with species integrity greater than a first threshold and the slope of the end of the species accumulation curve less than a second threshold are classified as observed areas, and the remaining blocks are classified as unobserved areas. The biodiversity data associated with the blocks classified as observed areas and the environmental variable data corresponding to the biodiversity data are used as a sample set. For example, the screening module 502 is specifically used to implement S1021-S1024 of the above biodiversity prediction method.

[0131] Optionally, the biodiversity prediction device further includes a training module 505. The training module 505 is used to train a random forest model. For example, the training module 505 is used to implement step S105 of the biodiversity prediction method described above.

[0132] In one application scenario, the training module 505 is specifically used to: train a random forest model using a training set; the input of the random forest model is environmental variable data, and the output of the random forest model is biodiversity prediction data; the sample set includes a training set and a test set.

[0133] Optionally, the biodiversity prediction device further includes a residual determination module 506. The residual determination module 506 is used to determine the biodiversity prediction residual for the study area. For example, the residual determination module 506 is used to implement step S106 of the above-described biodiversity prediction method.

[0134] In one application scenario, the residual determination module 506 is specifically used to: test the trained random forest model using a test set, and use the difference between the actual biodiversity data in the test set and the biodiversity prediction data output by the trained random forest model as the prediction residual for the observed area; the sample set includes a training set and a test set. The prediction residual for the observed area is interpolated to obtain the biodiversity prediction residual for the study area. For example, the residual determination module 506 is specifically used to implement steps S1061-S1062 of the above biodiversity prediction method.

[0135] The various modules of the above-mentioned biodiversity prediction device can also be used to perform other steps in the above method embodiments. All relevant content involved in the above method embodiments can be referred to in the functional description of the corresponding functional module, and will not be repeated here.

[0136] This application also provides an electronic device, including: a processor and a memory coupled to the processor; the memory is used to store computer instructions, and when the electronic device is running, the processor executes the computer instructions stored in the memory to cause the electronic device to perform the methods in the above embodiments. The processor can implement the acquisition module 501, the filtering module 502, the prediction module 503, and the determination module 504; the memory can also be used to store environmental variable data, biodiversity data, sample sets, biodiversity prediction data, and biodiversity prediction residuals, etc.

[0137] This application also provides a computer-readable storage medium including a computer program that, when run on a computer, performs the methods described in the above embodiments.

[0138] This application also provides a computer program product, which includes computer program instructions that, when run on a computer, execute the methods described in the above embodiments.

[0139] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0140] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for predicting biodiversity, characterized in that, include: Acquire environmental variable data and biodiversity data for the study area over a period of time; The biodiversity data includes: species name, number of individuals, and location of discovery; the environmental variable data includes: climate factors, topographic factors, habitat factors, and disturbance factors. Based on species integrity and the slope of the end of the species accumulation curve, the biodiversity data are screened to determine the sample set for the random forest model; within the study area, the portion of the biodiversity data containing the discovered locations is considered the observed area, otherwise it is considered the unobserved area; including: The research area is divided into multiple blocks; For each of the plurality of blocks, biodiversity data of locations falling within the geographic range of the block are associated with the block; and the species integrity and the slope of the end of the species accumulation curve of the block are calculated using the biodiversity data associated with the block. The biodiversity data in the plurality of blocks where the species integrity is greater than a first threshold and the slope of the species accumulation curve at the end is less than a second threshold, along with the environmental variable data corresponding to the biodiversity data, are used as the sample set. Blocks associated with biodiversity data in the sample set are assigned to the observed area, and the remaining blocks in the study area are assigned to the unobserved area. The random forest model predicts environmental variable data for the unobserved area over a given period of time, thereby obtaining predicted biodiversity data for the unobserved area over that period of time; the random forest model is trained using the sample set. Based on the biodiversity prediction data and the corresponding biodiversity prediction residuals, biodiversity data for the unobserved area within the specified time period are determined; the biodiversity prediction residuals are obtained by interpolating the prediction residuals of the random forest model.

2. The method as described in claim 1, characterized in that, The formula for calculating species integrity is as follows; Where SI represents species completeness, Sobs represents the number of species observed in the sample, Chao1 represents the richness index, n1 represents the number of species containing only one individual, and n2 represents the number of species containing only two individuals. The slope of the end of the species accumulation curve is calculated from the cumulative curve of the number of species in the block over the specified period.

3. The method as described in claim 1, characterized in that, The training method for the random forest model is as follows: The random forest model is trained using a training set; the input to the random forest model is environmental variable data, and the output of the random forest model is biodiversity prediction data; the sample set includes the training set and the test set.

4. The method as described in claim 1, characterized in that, The method for determining the biodiversity prediction residuals is as follows: The trained random forest model is tested using a test set, and the difference between the actual biodiversity data in the test set and the biodiversity prediction data output by the trained random forest model is used as the prediction residual for the observed area; the sample set includes the training set and the test set. Kriging interpolation is performed on the prediction residuals of the observed area to obtain the biodiversity prediction residuals of the study area.

5. The method as described in claim 1, characterized in that, The organisms mentioned include birds.

6. A biodiversity prediction device, applied in the biodiversity prediction method according to claim 1, characterized in that, The device includes an acquisition module, a filtering module, a prediction module, and a determination module; The acquisition module is used to acquire environmental variable data and biodiversity data of the study area over a period of time; the biodiversity data includes: species name, number of individuals, and discovery location. The filtering module is used to filter the biodiversity data and determine the sample set based on species integrity and the slope of the end of the species accumulation curve; in the study area, the part of the discovery location in the filtered biodiversity data that falls into the observed area is the observed area, otherwise it is the unobserved area; The prediction module is used by the random forest model to predict environmental variable data of the unobserved area within the specified time period, thereby obtaining biodiversity prediction data of the unobserved area within the specified time period; the random forest model is trained using the sample set. The determining module is used to determine the biodiversity data of the unobserved area within the specified time period based on the biodiversity prediction data and the biodiversity prediction residuals corresponding to the biodiversity prediction data; the biodiversity prediction residuals are obtained by interpolating the prediction residuals of the random forest model.

7. An electronic device, characterized in that, The device includes a processor and a memory coupled to the processor; the memory is used to store computer instructions, and when the electronic device is running, the processor executes the computer instructions stored in the memory to cause the electronic device to perform a biodiversity prediction method as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, It includes computer program instructions that, when executed by a computer, cause the computer to perform a biodiversity prediction method as described in any one of claims 1 to 5.

9. A computer program product, characterized in that, It includes computer program instructions that, when executed on a computer, cause the computer to perform a biodiversity prediction method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Biodiversity evaluation method and system

    CN116485274A

  • Urban scale biodiversity intelligent monitoring method and device

    CN118395395A