A Method for Functional Trait Completion in Plant Databases Based on Dual Matching of Geographic Coordinates and Species Classification

By using a dual matching method of geographic coordinates and species classification, the spatial heterogeneity and nomenclature differences of plant functional trait data have been resolved, enabling the completion of functional trait data in plant databases and supporting ecological research and ecosystem management.

CN120929451BActive Publication Date: 2026-04-03NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-14
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing plant functional trait data exhibit spatial heterogeneity and nomenclature differences across different locations and data sources, leading to biased research results and inaccurate data matching, thus limiting their application in ecological research.

Method used

A dual matching method based on geographic coordinates and species classification was adopted. By standardizing species names, unifying data formats, converting data to POI data format, and using buffer matching and hierarchical Bayesian interpolation, the plant functional trait data was completed.

Benefits of technology

It enables precise matching and completion of plant functional trait data, supporting data needs in fields such as ecological research, plant community structure research, environmental monitoring and assessment, and ecological protection and restoration, and providing solid data support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929451B_ABST
    Figure CN120929451B_ABST
Patent Text Reader

Abstract

This invention provides a method for completing plant database functional traits based on dual matching of geographic coordinates and species classification, comprising: Step 1, acquiring species data for each sample plot in the study area, calculating the relative abundance of species in the sample plots, and standardizing species names using the Global Plant List (TPL) database; Step 2, acquiring plant functional trait data from different sources; Step 3, unifying the data format of the collected plant functional trait data; Step 4, converting the sample plot data in the study area into point-of-interest (POI) data in sample point format, and recording the buffer radius when a match is successful; Step 5, compiling plant functional trait data suitable for the species in the sample plots; Step 6, interpolating missing data; Step 7, obtaining the Shannon index for each sample plot. This method supplements the plant database with functional trait data for the study area, has a wide range of applications, and can provide different perspectives and data support for ecosystem management and ecological research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of plant ecology and plant database construction, and in particular relates to a method for completing the functional traits of plant databases based on dual matching of geographic coordinates and species classification. Background Technology

[0002] Traits are relatively stable and measurable external characteristics exhibited by living organisms such as plants, animals, and microorganisms after long-term adaptation and evolution to the external environment. They have significant ecological importance and are a crucial pathway for exploring the impact and adaptation of organisms on the environment. Among these, plant functional traits refer to relatively stable and measurable morphological, physiological, and phenological parameters that directly or indirectly affect plant growth, survival, and reproduction. By collecting and integrating plant functional trait data to construct a plant functional trait database, crucial data support can be provided for establishing empirical or mechanistic functional trait-environment relationships, optimizing and validating parameters in dynamic vegetation models, and conducting multi-dimensional trait research from new perspectives at the species, functional group, and community levels. This effectively promotes functional trait research towards complex natural ecosystems and serves the solution of regional ecological and environmental problems. Although plant functional trait data has attracted attention, the following problems still exist in its use:

[0003] The morphology, anatomy, and water use of the same plant species can vary in different locations. Conversely, different plant species may exhibit similar functional trait responses due to similar environmental selection processes. This means that functional trait data for the same plant species often exhibit significant spatial heterogeneity across different regions. Therefore, collected functional trait data cannot be arbitrarily applied to the same plant species, as this may lead to biased research results. Secondly, due to differences in disciplinary background and experimental conditions, different data sources often have variations in the nomenclature of plant names and functional traits. There is no unified standard for species names across different data sources, and data from different sources often cannot be precisely matched due to naming conventions, translation differences, or inconsistent formats. This results in a large amount of valuable information being discarded or unconnected during the integration process. These two problems hinder the research on functional traits and, to some extent, limit the application of plant functional traits at different research scales and in different fields of ecology.

[0004] To meet the data needs of ecological research, researchers have developed several trait databases, including global databases such as the TRY database and the BIEN database, as well as regional trait databases such as the China Plant Trait Database (Second Edition) and China Plant Trait V2. However, due to the massive amount of data being scattered across various literatures and databases, and the potential inconsistencies in definitions from different sources, there are still serious gaps in the available data for species and traits in many regions. Many trait-based statistical methods require complete datasets, which means that the functional traits of the species being studied cannot have missing data. Summary of the Invention

[0005] Purpose of the invention: The technical problem to be solved by this invention is to address the shortcomings of existing technologies by providing a method for completing the functional traits of plant databases based on dual matching of geographic coordinates and species classification, including the following steps:

[0006] Step 1: Obtain species data for each plot in the study area, calculate the relative abundance of species in the plots, and standardize species names using the TPL database of the global plant list.

[0007] Step 2: Obtain plant functional trait data from different sources (TRY global plant functional trait database, China plant trait database, BIEN plant information and ecological network trait database, etc.).

[0008] Step 3: Standardize the data format of the collected plant functional trait data and split the Latin names of species in the data by genus and species name; convert the plant functional trait data into point-of-interest (POI) data based on the geographic coordinate information in the data.

[0009] Step 4: Convert the sample plot data in the study area into POI data in sample point format. Extract POI data of plant functional traits by establishing a buffer for the POI data in sample point format of the sample plots. Obtain the matching relationship table between the sample plots and plant functional trait data in the study area, and record the buffer radius when a match is successful.

[0010] Step 5: Based on the matching relationship table between the sample plots and plant functional trait data in the study area, extract the successfully matched sample plot data and the corresponding plant functional trait data. Then, based on the species names of the species contained in the sample plots, perform a second screening on the species names in the matched plant functional trait data and compile the plant functional trait data of the species that are suitable for the sample plots.

[0011] Step 6: Fill the successfully matched plant functional trait data into the corresponding study area plot data, count the plant functional trait data of each plot, sort out the information of plant species that were not matched and the information of plant species that still have missing functional trait data, and select to remove data or use the hierarchical Bayesian method BHPMF to interpolate the missing data according to the proportion of the number in the plot.

[0012] Step 7: Extract the relative height, relative basal diameter, and relative density of the species, calculate the importance of the species, and use the importance of the species as a weight to calculate the weighted average trait value of the plant community. Substitute the relative abundance of the species in the plot into the Shannon index formula to obtain the Shannon index for each plot.

[0013] In step 1, the relative abundance of species in the sample plots is calculated using the following formula:

[0014] ,

[0015] in Let be the relative abundance of the i-th species; Let be the number of individuals of the i-th species in the sample plot; It represents the sum of the number of individuals of all species in the sample plot.

[0016] In step 1, the plant name lookup function of the plantlist package in R language is used to standardize the species names in the sample plot. If the plant name lookup function TPL cannot find the corresponding result, the Flora of China is used for supplementary lookup.

[0017] In step 2, plant functional trait data are obtained through the global plant functional trait database TRY and the Chinese Plant Trait Database (Second Edition). At the same time, relevant literature on plant functional traits is retrieved to supplement the plant functional trait data of regions and species that are not covered by the global plant functional trait database TRY and the Chinese Plant Trait Database (Second Edition).

[0018] Step 3 includes:

[0019] Step 3-1: Standardize the format of the collected plant functional trait data and remove data with empty values, missing geographic coordinates, or species names;

[0020] Step 3-2: Delete irrelevant characters (unexpected spaces, commas, default values, etc.) from the data.

[0021] Step 3-3: Assign a unique number to each plant functional trait data point;

[0022] Steps 3-4: Split the Latin names of species in the data and represent them separately in the form of genus name and species name;

[0023] The standardized plant functional trait data are converted into a spatial vector format based on the latitude and longitude information in the data, so that the plant functional trait data are represented in the form of point-of-intake (POI) data.

[0024] Steps 3-4 include:

[0025] Convert the plant functional trait data file format to worksheet format. Using the geographic information system software ArcMap, select the conversion tool in the ArcToolbox and use the Excel to convert the plant functional trait data into a table. Then, use the "Add XY Data" function to fill in the latitude and longitude information, select the corresponding coordinate system, and represent the plant functional trait data in the form of point of interest (POI) data.

[0026] In step 4, the sample plot data in the study area are converted into point-of-indication (POI) data in latitude and longitude format using the method in step 3. The coordinate system of the sample plot POI data and the sample plot POI data of the plant functional traits is the same. After the conversion, a buffer is established for the sample plot POI data. The sample plot POI data of the plant functional traits is captured in the buffer to obtain the matching relationship between the sample plot POI data and the sample plot POI data of the plant functional traits. The buffer radius when the matching is successful is recorded.

[0027] In step 4, the sf::st_read() function in the Rtudio software is used to read the POI data in sample plot format and the POI data in sample plot format of plant functional traits, respectively; a circular buffer is created for the POI data in sample plot format using the st_buffer() function in the simple element sf package, with three lengths of 5km, 20km and 50km as the buffer radius, respectively.

[0028] The `st_join()` function is used to perform spatial joins, filtering out the sample point format POI data of plant functional trait data that fall within the buffer; the `complete.cases()` function is used to filter out unmatched null records; the `as.data.frame()` function is used to convert the matching result sf object into a standard data frame format; and the `merge()` function is used to integrate and statistically analyze the sample point format POI data of successfully matched plots and the sample point format POI data of plant functional trait data, generating a geographic coordinate matching table of the sample point format POI data of plots and the sample point format POI data of plant functional trait data.

[0029] Step 5 includes: performing the following three-fold matching on each plant functional trait data by combining the Chinese species names and Latin plant names from the sample plot data:

[0030] The first match is to match the Chinese names of species in the plant functional trait data with the Chinese names of species in the sample plot data. Only Chinese characters are compared to check if they are completely consistent. The first match has two results: a successful match and a failed match.

[0031] The second matching involves matching the genus name portion of the species' Latin name in the plant functional trait data with the species' Latin name in the sample plot data to check for an inclusion relationship. That is, whether the species' Latin name in the sample plot data with a successful geographical match contains characters from the genus name portion of the species' Latin name in the plant functional trait data. The second matching has two possible results: inclusion and non-inclusion.

[0032] The third matching involves matching the species name part of the Latin name of the species in the plant functional trait data with the species Latin name in the sample plot data to check whether there is an inclusion relationship. That is, whether the species Latin name in the sample plot data with successful geographical matching contains the characters of the species name part of the Latin name of the species in the plant functional trait data. The third matching has two results: inclusion and non-inclusion.

[0033] Step 6 includes: statistically analyzing species for which no functional trait data was found or for which functional trait data was still missing. If the relative abundance of a species in the sample plot is below a threshold of 10%, the species is removed; otherwise, the hierarchical Bayesian method BHPMF is used for interpolation, with the formula as follows:

[0034] ,

[0035] Where E is the expected value, n is the row index, m is the column index, and L is the number of levels in the hierarchical structure. For the original object and property matrix, For row-side potential vectors, For column-side latent vectors, and For interlayer canonical strength, For nodes The latent vector of the parent node in the (l-1)th layer; , The standard deviation of the filler value; For indicator variables, mark the position. Are there any observations, when When (n,m) are non-missing values, take Otherwise, it is 0; for the bottom-up approach, replace l-1 with l+1 and set the parent node... Replace with child nodes ;

[0036] Step 7 includes: extracting the relative height, relative basal diameter, and relative density indices of the species; calculating the species importance index (IV); using species importance as a weight; combining it with the matched functional trait data; calculating the community weighted trait value (CWM); and performing a normality test (Shapiro.test) on the community weighted trait value CWM; and performing a log-log transformation on the community weighted trait value CWM that does not meet the normal distribution. The specific calculation formula is as follows:

[0037] ,

[0038] ,

[0039] Where S represents the total number of species in the sample plot. X1 represents the individual trait value of the i-th species in each quadrat; X2 represents the relative height, X3 represents the relative basal diameter, and X4 represents the relative density.

[0040] Finally, the relative abundance calculated in step 1 and the total number of species in each plot are substituted into the Shannon index formula to obtain the Shannon index for each plot:

[0041] ,

[0042] Where S is the number of species;

[0043] The successfully matched plant functional trait data, the calculated community weighted trait values, and the Shannon index are entered into the corresponding plot information in the database to improve the plant functional trait content in the database.

[0044] This method supplements existing plant databases with applicable plant functional trait data through a dual matching mechanism of geographic coordinates and species classification. This is of great significance for research in fields such as plant community structure, environmental monitoring and assessment, ecological protection and restoration, and construction of plant functional trait networks.

[0045] This technology can be applied to the following fields:

[0046] Supplementing plant functional trait data: The interactions among plant functional traits collectively determine plant function. Supplementing plant databases with existing plant functional trait data helps to study vegetation responses to environmental conditions at the individual to plant community scale.

[0047] Constructing plant functional trait networks: supplementing plant databases with functional trait data can provide data support for capturing and visualizing the relationships between multiple plant traits in multiple dimensions, and revealing plant response and adaptation strategies to environmental or resource changes.

[0048] Ecological research: By matching geographical coordinates with species classification, existing plant functional trait data can be fully utilized to supplement the plant database of the study area. This helps to reveal the relationship between individual traits and the entire functional trait system, predict the adaptive responses of plants under different environmental and resource conditions, and provide more solid data support for the study of ecosystem stability.

[0049] Biodiversity conservation and ecosystem management: In the context of current global climate change, integrating plant functional trait data can provide data support for guiding the formulation of biodiversity conservation and ecosystem management policies.

[0050] Beneficial Effects: This invention proposes a method for completing plant database functional trait data based on dual matching of geographic coordinates and species classification. It integrates existing plant functional trait data to complete the required functional traits for the plant database of the study area. This is of great significance for research in community and functional ecology, biodiversity conservation, ecosystem and landscape management, and land surface modeling. This method fully utilizes the existing massive amounts of plant functional trait data, supplementing the plant database of the study area with functional trait data through meticulous data matching and screening. It has a wide range of applications and can provide different perspectives and data support for ecosystem management and ecological research. Attached Figure Description

[0051] Figure 1 This is a schematic diagram of the geographic coordinate matching buffer.

[0052] Figure 2 This is a diagram illustrating species classification and matching.

[0053] Figure 3 This is a flowchart of the method of the present invention. Detailed Implementation

[0054] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, and the advantages of the present invention in the above and / or other aspects will become clearer.

[0055] like Figure 3 As shown, this embodiment of the invention provides a method for completing functional traits in a plant database based on dual matching of geographic coordinates and species classification, including the following steps:

[0056] Step 1: Obtain species data for each plot in the study area, calculate the relative abundance of species in the plots, and standardize species names using the TPL database.

[0057] Step 2: Obtain plant functional trait data from different sources, such as the TRY Global Plant Functional Trait Database, the Chinese Plant Trait Database, and the BIEN Trait Database.

[0058] Step 3: Standardize the data format of the collected plant functional trait data and separate the Latin names of species in the data by genus and species. After data preprocessing, convert the plant functional trait data into POI (point of interest) data based on the geographic coordinate information in the data.

[0059] Step 4: Convert the sample plot data in the study area into POI data. By establishing a buffer for the sample plot POIs, capture the POI data of plant functional traits multiple times to obtain the matching relationship between the sample plots and plant functional trait data in the study area, and record the buffer radius when the matching is successful.

[0060] Step 5: Extract the successfully matched sample plot data and the corresponding plant functional trait data. Based on the species names of the species contained in the sample plots, perform a secondary screening of the species names in the matched plant functional trait data, and compile the plant functional trait data of the species suitable for the sample plots.

[0061] Step 6: Fill the successfully matched plant functional trait data into the corresponding study area plot data, count the plant functional trait data of each plot, sort out the information of plant species that were not matched and the information of plant species that still have missing functional trait data, and select to remove the data or use the hierarchical Bayesian method BHPMF to interpolate the missing data according to their proportion in the plot.

[0062] Step 7: Extract the relative height, relative basal diameter, and relative density of the species to calculate the species importance. Then, use species importance as a weight to calculate the weighted average trait value of the plant community. Substitute the relative abundance of the species in the plot into the Shannon index formula to obtain the Shannon index for each plot.

[0063] This method is applicable to scenarios where the database already possesses complete sample plot information (including the geographical coordinates of the sample plots, the number of plants within each plot, and species names), but lacks plant functional trait data. The process includes: acquiring plant functional trait data from external sources and converting it into spatial points based on its geographical coordinates; setting a buffer range for each sample plot in the database and extracting functional trait data points within the buffer; matching the extracted data with species names based on the plant names recorded within the sample plots, and filling the corresponding plant information in the sample plots with the successfully matched functional trait data; finally, using the completed trait data, calculating the weighted average trait value and Shannon index of the plant community in each sample plot, thereby quantifying the community characteristics of plant functional traits.

[0064] In a specific embodiment of this scheme, the study area is selected at the Chinese scale. The sample plot data is obtained by combining data from field surveys and some literature, and includes information such as sample plot number, geographic coordinates, Chinese name of species, and Latin name of species.

[0065] The relative abundance of species in the sample plots was calculated as follows:

[0066] ,

[0067] in Let be the relative abundance of the i-th species; Let be the number of individuals of the i-th species in the sample plot; It represents the sum of the number of individuals of all species in the sample plot.

[0068] The standardization of species names in the sample plots is accomplished using the TPL() function of the plantlist package in R. If TPL() cannot find the corresponding results, the Flora of China can be used for supplementary queries.

[0069] In this specific embodiment, the plant functional trait data in step 2 mainly come from the TRY Global Plant Functional Trait Database and the second edition of the Chinese Plant Trait Database. At the same time, relevant plant functional trait literature is retrieved to supplement the plant functional trait data of regions and species that cannot be covered by the above two databases.

[0070] After data collection is completed, the format of plant functional trait data from different sources needs to be standardized, and the Latin names of plants need to be separated according to the classification of genus and species. The Latin names of plants are represented separately in the form of genus name and species name for subsequent species name matching.

[0071] In this specific embodiment, step 3 includes:

[0072] Step 3-1: Standardize the format of the collected functional trait data and remove data with null values, or without geographic coordinates or species names.

[0073] Step 3-2: Delete unexpected spaces, commas, default values, and other irrelevant characters from the data;

[0074] Step 3-3: Assign a unique number to each functional trait data.

[0075] Steps 3-4 involve splitting the Latin names of species in the data and representing them separately as genus and species names.

[0076] The standardized plant functional trait data are converted into a spatial vector format based on the latitude and longitude information in the data, so that the plant functional trait data can be represented in the form of POI sites.

[0077] In this specific embodiment, steps 3-4 include:

[0078] Change the file format of the plant functional trait data to .xls. Use ArcMap software, select the conversion tool in the ArcToolbox, convert the trait data into a table using the Excel to table conversion tool, and then fill in the latitude and longitude information using the add XY data function. Select the corresponding coordinate system to represent the functional trait data in the form of POI data.

[0079] like Figure 1 As shown, the red dot in the circle represents the Zijinshan site in Nanjing, Jiangsu Province, with coordinates (32.05, 118.833). The green dots represent points containing plant functional trait data. The green dots enclosed in the red circle are the points that can be added to the Zijinshan site for plant functional trait data. After the plant functional trait data is converted, the sample plot data in the study area is converted into POI data in the form of latitude and longitude information using the same method as in step 3, and the coordinate system of the sample plot POI data and the functional trait POI data is the same. After the conversion, a buffer is established for the sample plot POI data, and the functional trait POI data within the buffer is captured to obtain the matching relationship between the sample plot POI data and the functional trait POI data, and the buffer radius when a match is successful is recorded.

[0080] In this specific embodiment, step 4 includes:

[0081] The `sf::st_read()` function in the Rtudio software was used to read the plot POI data and functional trait POI data respectively. A circular buffer was created for the plot POI data using the `st_buffer()` function in the `sf` package, with buffer radii of 5km, 20km, and 50km respectively. Spatial join was performed using `st_join()` to filter out functional trait POI data falling within the buffer. `complete.cases()` was used to filter out unmatched null records. The matching result `sf` object was converted to a standard data frame format using `as.data.frame()`. The `merge()` function was used to integrate and statistically analyze the successfully matched plot POI and functional trait POI data, generating a geographic coordinate matching table for both.

[0082] like Figure 2 As shown, using the geographic coordinate matching table obtained in step 4, the corresponding sample plot data and plant functional trait data can be extracted respectively. The plant species names in the sample plot data are matched with the species names in the plant functional trait data. The successfully matched functional trait data are filled into the corresponding sample plot information to obtain the plant functional trait data of the sample plots in the study area whose geographic coordinates and species classifications both match.

[0083] In this embodiment, step 5 includes:

[0084] Typically, species name matching is performed directly by comparing characters. However, due to naming rules, transliteration, and other reasons, data from different sources may assign different species names to the same plant, leading to matching character failures. To reduce such errors, this method splits the Latin name of the species in the plant functional trait data into genus name and species name, and performs three matches on each functional trait data point in conjunction with the Chinese name of the species:

[0085] First match: Plant functional trait data [species Chinese name] matched plot data [species Chinese name];

[0086] Second matching: matching plant functional trait data [species Latin name - genus name part] with plot data [species Latin name];

[0087] Third matching: matching plant functional trait data [species Latin name - species name part] with plot data [species Latin name];

[0088] The first matching involves matching the Chinese names of species in the plant functional trait data with the Chinese names of species in the sample plot data. Only the Chinese characters on both sides are compared to check if they are completely consistent. This matching has two possible results: successful matching and failed matching.

[0089] The second matching involves matching the genus name portion of the species' Latin name in the plant functional trait data with the species' Latin name in the sample plot data to check if there is an inclusion relationship between the two. That is, whether the species' Latin name in the sample plot data with successful geographical matching contains characters from the genus name portion of the species' Latin name in the plant functional trait data. This matching has two possible results: inclusion and non-inclusion.

[0090] The third matching involves matching the species name part of the Latin name of the species in the plant functional trait data with the species Latin name in the sample plot data to check whether there is an inclusion relationship between the two. That is, whether the species Latin name in the sample plot data with successful geographical matching contains the characters of the species name part of the species Latin name in the plant functional trait data. This matching has two results: inclusion and non-inclusion.

[0091] The combinations of the three matching results are shown in Table 1.

[0092] Table 1. List of combinations of results from triple matching

[0093]

[0094] Results with serial numbers 1 and 5 are considered directly usable data and can be entered into the plant information of the corresponding plot. Results with serial numbers 2, 3, 4, 6, and 7 are considered questionable data and can be entered into the corresponding plant information after manual screening to prevent data loss due to naming problems or data recording errors.

[0095] After the data screening is completed, the plant functional trait data of each sample plot are counted, and the information of plant species that have not been matched or whose functional trait data are still missing is sorted out. Based on their proportion in the sample plot, the data is either removed or the missing values ​​are interpolated using the BHPMF hierarchical Bayesian method.

[0096] In this specific example summary, step 6 includes:

[0097] For species with no matched functional trait data or whose functional trait data are still missing, statistical analysis was performed. If the relative abundance of a species in its sample plot was less than 10%, the species was removed; if it was greater than 10%, the BHPMF stratified Bayesian method was used for interpolation. The formula is:

[0098] ,

[0099] Where E is the expected value, n is the row index (individual / species), m is the column index (trait), and L is the number of levels in the hierarchical structure. For indicator variables, mark the position. Are there any observations? The original object and trait matrix is ​​used. In BHPMF, the traits are typically log-normalized and z-score-normalized before being fed into the model. The observations used for training and evaluation come from the observed portion of X. For row-side potential vectors, For column-side latent vectors, and For interlayer canonical strength, For nodes The parent node at the 1st The latent vector of the layer; It is a collection of data at all levels; when When (n,m) are non-missing values, take Otherwise, it is 0. For bottom-up methods, use replace and the parent node Replace with child nodes .

[0100] After screening using a dual matching mechanism of geographic information and species classification, functional trait data of plants in the sample plots of the study area were obtained. Using this data, species importance was calculated, and the weighted average trait value of the plant community was calculated using species importance as a weight.

[0101] In this specific embodiment, step 7 includes:

[0102] The relative height, relative basal diameter, and relative density of extracted species were used to calculate species importance (IV). Species importance was then used as a weight and combined with matched functional trait data to calculate community-weighted trait values ​​(CWM). The calculated CWMs were then subjected to a Shapiro test for normality, and those not conforming to a normal distribution were transformed using a logarithmic transformation. The specific calculation formula is as follows:

[0103] ,

[0104] ,

[0105] Where s represents the total number of species in the sample plot. Let be the importance value of the i-th species in the sample plot. X1 represents the individual trait value of the i-th species in each quadrat. X2 represents the relative height, X3 represents the relative basal diameter, and X4 represents the relative density.

[0106] Finally, the relative abundance calculated in step 1 and the total number of species in the sample plots are substituted into the Shannon index formula to obtain the Shannon index for each sample plot.

[0107] In this specific embodiment, the Shannon index formula used is as follows:

[0108] ,

[0109] Where S is the number of species. It is the relative abundance of the i-th species.

[0110] The successfully matched plant functional trait data, the calculated community-weighted trait values, and the Shannon index are entered into the corresponding plot information in the database to improve the plant functional trait content in the database.

[0111] This invention provides a method for completing the functional traits of a plant database based on dual matching of geographic coordinates and species classification. Many methods and approaches exist for implementing this technical solution; the above description is merely a preferred embodiment of the invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this invention, and these improvements and modifications should also be considered within the scope of protection of this invention. All components not explicitly stated in this embodiment can be implemented using existing technologies.

Claims

1. A method for completing functional traits in plant databases based on dual matching of geographic coordinates and species classification, characterized in that, Includes the following steps: Step 1: Obtain species data for each plot in the study area, calculate the relative abundance of species in the plots, and standardize species names using the TPL database of the global plant list. Step 2: Obtain plant functional trait data from different sources; Step 3: Standardize the data format of the collected plant functional trait data and split the Latin names of species in the data by genus and species name; convert the plant functional trait data into point-of-interest (POI) data based on the geographic coordinate information in the data. Step 4: Convert the sample plot data in the study area into POI data in sample point format. Extract POI data of plant functional traits by establishing a buffer for the POI data in sample point format of the sample plots. Obtain the matching relationship table between the sample plots and plant functional trait data in the study area, and record the buffer radius when a match is successful. Step 5: Based on the matching relationship table between the sample plots and plant functional trait data in the study area, extract the successfully matched sample plot data and the corresponding plant functional trait data. Then, based on the species names of the species contained in the sample plots, perform a second screening on the species names in the matched plant functional trait data and compile the plant functional trait data of the species that are suitable for the sample plots. Step 6: Fill the successfully matched plant functional trait data into the corresponding study area plot data, count the plant functional trait data of each plot, sort out the information of plant species that were not matched and the information of plant species that still have missing functional trait data, and select to remove data or use the hierarchical Bayesian method BHPMF to interpolate the missing data according to the proportion of the number in the plot. Step 7: Extract the relative height, relative basal diameter, and relative density of the species, calculate the importance of the species, and use the importance of the species as a weight to calculate the weighted average trait value of the plant community. Substitute the relative abundance of the species in the plot into the Shannon index formula to obtain the Shannon index for each plot. Fill the successfully matched plant functional trait data, the calculated community weighted trait value, and the Shannon index into the corresponding plot information in the database to improve the plant functional trait content in the database.

2. The method according to claim 1, characterized in that, In step 1, the relative abundance of species in the sample plots is calculated using the following formula: Where S is the total number of species in the sample plot, and P i Let n be the relative abundance of the i-th species; i Let be the number of individuals of the i-th species in the sample plot; It represents the sum of the number of individuals of all species in the sample plot.

3. The method according to claim 2, characterized in that, In step 1, the plant name lookup function of the plantlist package in R language is used to standardize the species names in the sample plot. If the plant name lookup function TPL cannot find the corresponding result, the Flora of China is used for supplementary lookup.

4. The method according to claim 3, characterized in that, Step 3 includes: Step 3-1: Standardize the format of the collected plant functional trait data and remove data with empty values, missing geographic coordinates, or species names; Step 3-2: Delete irrelevant characters from the data; Step 3-3: Assign a unique number to each plant functional trait data point; Steps 3-4: Split the Latin names of species in the data and represent them separately in the form of genus name and species name; The standardized plant functional trait data are converted into a spatial vector format based on the latitude and longitude information in the data, so that the plant functional trait data are represented in the form of point-of-intake (POI) data.

5. The method according to claim 4, characterized in that, Steps 3-4 include: Convert the plant functional trait data file format to worksheet format. Using the geographic information system software ArcMap, select the conversion tool in the ArcToolbox and use the Excel to convert the plant functional trait data into a table. Then, use the "Add XY Data" function to fill in the latitude and longitude information, select the corresponding coordinate system, and represent the plant functional trait data in the form of point of interest (POI) data.

6. The method according to claim 5, characterized in that, In step 4, the sample plot data in the study area are converted into point-of-indication (POI) data in latitude and longitude format using the method in step 3. The coordinate system of the sample plot POI data and the sample plot POI data of the plant functional traits is the same. After the conversion, a buffer is established for the sample plot POI data. The sample plot POI data of the plant functional traits is captured in the buffer to obtain the matching relationship between the sample plot POI data and the sample plot POI data of the plant functional traits. The buffer radius when the matching is successful is recorded.

7. The method according to claim 6, characterized in that, In step 4, the sf::st_read() function in the Rtudio software is used to read the POI data in sample plot format and the POI data in sample plot format of plant functional traits, respectively; a circular buffer is created for the POI data in sample plot format using the st_buffer() function in the simple element sf package, with three lengths of 5km, 20km and 50km as the buffer radius, respectively. The `st_join()` function is used to perform spatial joins, filtering out the sample point format POI data of plant functional trait data that fall within the buffer; the `complete.cases()` function is used to filter out unmatched null records; the `as.data.frame()` function is used to convert the matching result sf object into a standard data frame format; and the `merge()` function is used to integrate and statistically analyze the sample point format POI data of successfully matched plots and the sample point format POI data of plant functional trait data, generating a geographic coordinate matching table of the sample point format POI data of plots and the sample point format POI data of plant functional trait data.

8. The method according to claim 7, characterized in that, Step 5 includes: performing the following three-fold matching on each plant functional trait data by combining the Chinese names of species and the Latin names of plants in the sample plot data: The first match is to match the Chinese names of species in the plant functional trait data with the Chinese names of species in the sample plot data. Only Chinese characters are compared to check if they are completely consistent. The first match has two results: a successful match and a failed match. The second matching involves matching the genus name part of the species Latin name in the plant functional trait data with the species Latin name in the sample plot data to check for the existence of an inclusion relationship. That is, whether the species Latin name in the sample plot data with successful geographical matching contains the characters of the genus name part of the species Latin name in the plant functional trait data. The second matching has two results: inclusion and non-inclusion. The third matching involves matching the species name part of the Latin name of the species in the plant functional trait data with the species Latin name in the sample plot data to check whether there is an inclusion relationship. That is, whether the species Latin name in the sample plot data with successful geographical matching contains the characters of the species name part of the Latin name of the species in the plant functional trait data. The third matching has two results: inclusion and non-inclusion.

9. The method according to claim 8, characterized in that, Step 6 includes: statistically analyzing species for which no functional trait data was found or for which functional trait data was still missing. If the relative abundance of a species in the sample plot is below a threshold of 10%, the species is removed; otherwise, the hierarchical Bayesian method BHPMF is used for interpolation, with the formula as follows: Where E is the expected value, n is the row index, m is the column index, and L is the number of levels in the hierarchical structure. For the original object and property matrix, For row-side potential vectors, Let λ be the column-side latent vector. u and λ v For interlayer canonical strength, Let n be the potential vector of the parent node of node n in the (l-1)th layer; σ is the standard deviation of the filled value; As an indicator variable, mark whether there is an observation at position (n,m,l), when X (l) When (n,m) are non-missing values, take Otherwise, it is 0; for the bottom-up approach, l+1 is replaced with l-1, and the parent node p(n) is replaced with the child node c(n).

10. The method according to claim 9, characterized in that, Step 7 includes: extracting the relative height, relative basal diameter, and relative density indices of the species; calculating the species importance index (IV); using species importance as a weight; combining it with the matched functional trait data; calculating the community weighted trait value (CWM); and performing a normality test (Shapiro.test) on the community weighted trait value CWM; and performing a log-log transformation on the community weighted trait value CWM that does not meet the normal distribution. The specific calculation formula is as follows: IV = (X1 + X2 + X3) / 3 Among them, T i X1 represents the individual trait value of the i-th species in each quadrat; X2 represents the relative height, X3 represents the relative basal diameter, and X4 represents the relative density. Finally, the relative abundance calculated in step 1 and the total number of species in each plot are substituted into the Shannon index formula to obtain the Shannon index for each plot: Where S is the number of species.

Citation Information

Patent Citations

  • Water ecological pressure detection method and device based on machine learning

    CN119692853A

  • Aquatic organism diversity rapid evaluation system and method based on GIS and eDNA technologies

    CN120375927A