Biological group zoning algorithm based on R software
By using a biome localization algorithm based on R software, the problem of time-consuming pollen data processing in existing technologies has been solved, and efficient paleovegetation reconstruction and result output have been achieved.
Patent Information
- Application Number
- CN202511774848.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-03-13
AI Technical Summary
Existing pollen biomes delineation software, such as Biomise and PPPBase, is not open source, resulting in time-consuming and inconvenient data processing, making it difficult to efficiently perform batch data calculations.
A biome regionization algorithm based on R software was adopted. By collecting and organizing pollen data, plant functional types and biome region matrices were established. R language was used to perform weight reduction calculation and matrix filtering to output ancient vegetation types.
This invention enables efficient biomesarization of pollen data using the free and open-source software R, simplifying the data processing workflow and improving computational efficiency and the intuitiveness of the output results.
Smart Images

Figure CN121659100A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of plant pollen ecology and vegetation ecology, and specifically relates to a biome regionization algorithm based on R software. Background Technology
[0002] Quantitative reconstruction of paleovegetation distribution patterns not only allows for the study of the evolutionary patterns of paleoenvironments themselves, but also provides crucial scientific data for exploring the interrelationships between paleovegetation, paleoclimate, and ancient human activities. Plant pollen, with its abundant production, wide distribution, ease of preservation, and distinctive morphological characteristics, has become an extremely important proxy indicator for paleoenvironmental research. Utilizing pollen for biomesatization is an important method for paleovegetation reconstruction.
[0003] Currently, pollen biomesatization is performed using specialized software such as Biomise or PPPBase. These software programs are not open-source and require contacting the developers to obtain them; they are time-consuming and inconvenient for batch data processing. Summary of the Invention
[0004] The purpose of this invention is to provide a biome type algorithm based on R software, which uses R software to quantitatively convert pollen data into biome types, thereby restoring ancient vegetation types.
[0005] To further achieve the above objectives, the present invention adopts the following technical solution: A biome regionalization algorithm based on R software includes the following steps: Step 1: Collect and organize pollen datasets, compiling pollen data including sampling point latitude and longitude, altitude, different pollen groups and their percentages into a CSV file; Step 2: Merge the pollen groups from Step 1 into the corresponding plant functional types. If a pollen group belongs to the plant functional type, assign it a value of 1; otherwise, assign it a value of 0. Establish a plant functional type to pollen group matrix according to this rule. Step 3: Define biomes based on the key plant functional type combinations in Step 2. If a plant functional type belongs to a biome, assign it a value of 1; otherwise, assign it a value of 0. Establish a biome region to plant functional type matrix according to this rule. Step 4: Import the pollen groups and their percentage data corresponding to the sampling points into R software for weighted calculation; Step 5: Select the pollen group pair plant functional type matrix from the plant functional type pair pollen group matrix of the sampling points from the plant functional type pair pollen group matrix of Step 2; Step 6: Determine whether the pollen group at the sampling point is present or absent in the pollen group to plant functional type matrix, with a corresponding value of 1 or 0. If the value is 1, assign the value calculated in Step 4 to the plant functional type corresponding to that pollen group; if the value is 0, assign the corresponding plant functional type to 0. Sum the values of all pollen groups corresponding to each plant functional type at the sampling point to obtain the plant functional type score for each sampling point. Step 7: Determine whether each plant functional type at the sampling point exists or not in the matrix of plant functional types in the biome region, with a corresponding value of 1 or 0. If the value is 1, assign the score of the plant functional type to the biome region corresponding to that plant functional type; if the value is 0, assign the corresponding biome region a value of 0. Sum the scores of all plant functional types in each biome region at the sampling point to obtain the biome region score. Step 8: Select the biome type with the highest similarity score for each sampling point and output the calculation results.
[0006] Optionally, in step 4, the weight reduction calculation refers to subtracting the threshold from the percentage of pollen groups and then taking the square root.
[0007] Optionally, in step 6, the determination is achieved using the ifelse function in the R language. If the value of a certain plant functional type corresponding to a certain pollen group in the matrix is 1, then it is determined that the pollen group belongs to that plant functional type.
[0008] Optionally, the summation in steps 6 and 7 reflects the summation over all pollen groups in the following biomesarization algorithm formula: ; In the formula, A ik It is the similarity between pollen sample k and biota region i, S j This is the summation of all pollen groups j, d ij P represents the presence or absence (1 or 0) of biota region i and pollen group j in the biota region and pollen group matrix. jk It is the percentage content of pollen, Q j It is the threshold for the percentage of pollen.
[0009] The key steps of this invention include: 1. Data collection and preparation, which includes collecting pollen data, standardizing the format, and preparing two matrix data. This step is the foundation and prerequisite for all subsequent calculations. 2. Screening of pollen data from sampling points, which filters and extracts pollen data from a smaller number of sampling points from the large overall database, determining the specific target for the next calculation. 3. Biome regionalization calculation. This is a key step in calculating the biome regionalization of pollen groups at sampling points according to a formula. The biome regionalization result of the pollen sample is determined by two summations and screening for the first maximum value. 4. Output of results. The calculation results are output as pollen groups and their corresponding biome regionalization types, intuitively and clearly reflecting the result from pollen data to vegetation reconstruction.
[0010] Compared with the prior art, the present invention has the following beneficial effects: This invention's algorithm is based on the computational principle of biomesification and is written in R language to process pollen data into biomes. R software is free and open-source, readily available, and only requires inputting standard-formatted tabular data to complete batch data calculations and output results. The calculation process is simple, efficient, and convenient. Attached Figure Description
[0011] Figure 1 This is a flowchart of the algorithm of the present invention. Detailed Implementation
[0012] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0013] A biome regionalization algorithm based on R software includes the following steps: Step 1: Collect and organize the pollen dataset. Organize the pollen data, including the latitude and longitude of the sampling points, altitude, different pollen groups, and their percentages, into a CSV file. Collect pollen samples, including pollen-containing topsoil and sediment samples. Extract pollen in the laboratory, identify pollen groups under a microscope, and calculate the percentage of each pollen group to obtain pollen data. Pollen data can also be derived from existing pollen databases and numericalized pollen atlases.
[0014] Step 2: Assign pollen groups from Step 1 to their corresponding plant functional types. If a pollen group belongs to a plant functional type, assign it a value of 1; otherwise, assign it a value of 0. Establish a plant functional type-pollen group matrix according to this rule. Plant functional types play a crucial role in pollen biomes regionalization, serving as a bridge between pollen groups and biomes. Plant functional types are defined by four plant traits: life form, leaf type, leaf stage, and bioclimatic tolerance to cold and drought. Each plant functional type is defined by the intersection of these four traits and named in the order of bioclimatic tolerance, leaf stage, leaf type, and life form. For example, the pollen group *Cephalotaxus japonicus* (Japanese cedar) can be mapped to the plant functional type *Cephalotaxus japonicus*. Following this method, all pollen groups can be mapped to their corresponding plant functional types, thus establishing a plant functional type-pollen group matrix.
[0015] Step 3: Define biomes based on the key plant functional type combinations in Step 2. If a plant functional type belongs to a biome, assign it a value of 1; otherwise, assign it a value of 0. Establish a biome region-plant functional type matrix according to this rule.
[0016] Biomes can be defined based on combinations of characteristic plant functional types. For example, the combination of six characteristic plant functional types—subtropical evergreen conifers, temperate evergreen conifers, subtropical evergreen sclerophyllous broadleaf trees, subtropical evergreen soft-leaved broadleaf trees, subtropical evergreen soft-leaved broadleaf shrubs, and subtropical evergreen sclerophyllous broadleaf shrubs—can define a biome as subtropical evergreen broadleaf forest. By mapping all plant functional types to their corresponding biomes in this way, a biome-plant functional type matrix is established. Currently, the commonly used biome system in China includes 19 biome types, which basically reflect the vegetation changes along the north-south latitudinal temperature gradient and the east-west meridional moisture gradient.
[0017] The 19 biome types include: cold-temperate deciduous forest, cold-temperate evergreen coniferous forest, cold-temperate evergreen coniferous forest and mixed forest, cold-temperate evergreen coniferous forest, cold-temperate mixed forest, temperate deciduous broad-leaved forest, warm-temperate (subtropical) evergreen broad-leaved forest and mixed forest (mixed forest with evergreen components as the main component), warm-temperate (subtropical) pure evergreen broad-leaved forest, tropical semi-evergreen broad-leaved forest, tropical evergreen broad-leaved forest, tropical deciduous broad-leaved forest and sparse forest, temperate arid shrubland, temperate grassland, temperate desert, cushion grassland, grass and miscellaneous grass grassland, creeping dwarf shrub grassland, upright dwarf shrub grassland, and tall and dwarf shrub grassland.
[0018] Currently, China has 76 plant functional types. When defining biomes by combining these plant functional types, the biome type is defined by the combination of characteristic (or key) plant functional types. The characteristic or key plant functional types should reflect the vegetation characteristics of my country and also select dominant plant functional types. Currently, 38 key plant functional types are mainly used to define biomes in China. The 38 plant functional types include ar.cd.mb.eds, ar.cd.mb.pds, ar.enpds, ar.fb,bo.cd.mb.lhs, bo.cd.mb.t,bo.e.mb.lhs, bo.ent,di.sl.lhs,dt.sl.lhs,eu-dt.fb,eu.ent,g, ha, lsuc, rc.fb, s, sl.t, te.cd.mb.lhs, te-dt.fb,te-fa.cd.mb.t,te-fi.cd.mb.t, te-ft.cd.mb.t, tr.e.mb.lhs,tr.e.mb.t,tr.e.sb.t, tr-m.dd.mb.lhs, tr-m.dd.mb.t,tr-x.dd.mb.lhs,tr-x.dd.mb.t, wt.cd.mb.lhs,wt.cd.mb.t, wt.dnt,wt.e.mb.lhs,wt.e.mb.t, wt.ent, wt.e.sb.lhs, wt.e.sb.t. For example, wt.e.sb.t is an abbreviation for Warm-temperate evergreen sclerophyll broad-leaved tree, representing the functional type of subtropical evergreen sclerophyll broad-leaved tree.
[0019] Step 4: Import the pollen groups and their percentage data corresponding to the sampling points into R software for deweighting calculation. Deweighting calculation involves subtracting a threshold (usually set to 0.5) from the percentage of pollen groups and then taking the square root. Because some pollen groups, such as pine pollen, are produced in large quantities and are easily dispersed, they are overexpressed in pollen samples. Deweighting reduces the weight of these overexpressed pollen groups by taking the square root, making the calculation results closer to reality.
[0020] Step 5: Filter the pollen group-plant functional type matrix of the sampling points from the plant functional type-pollen group matrix obtained in Step 2. This step is performed using the `subset` function in R software. Specifically, the pollen groups of the sampling points are a subset of the pollen group data established in Step 1. The `subset` function extracts this subset from the whole set and saves it as a data frame.
[0021] Step 6: Determine whether the pollen group at the sampling point is present or absent in the pollen group to plant functional type matrix, with a corresponding value of 1 or 0. If the value is 1, assign the value calculated in Step 4 to the plant functional type corresponding to that pollen group; if the value is 0, assign the corresponding plant functional type to 0. Sum the values of all pollen groups corresponding to each plant functional type at the sampling point to obtain the plant functional type score for each sampling point.
[0022] The determination is achieved using the ifelse function in the R language. If the value of a certain plant functional type corresponding to a certain pollen group in the matrix is 1, then the pollen group is determined to belong to that plant functional type.
[0023] Step 7: Determine whether each plant functional type at the sampling point exists or not in the matrix of plant functional types in the biome region, with a corresponding value of 1 or 0. If the value is 1, assign the score of the plant functional type to the biome region corresponding to that plant functional type; if the value is 0, assign the corresponding biome region a value of 0. Sum the scores of all plant functional types in each biome region at the sampling point to obtain the biome region score.
[0024] Step 8: Select the biome type with the highest similarity score for each sampling point and output the calculation results.
[0025] Example 1 See Figure 1 The biomesatization process of topsoil pollen data at six altitudes in Lushan using the algorithm of this invention is as follows: Step 1: Collect and organize pollen groups and their percentage data nationwide, and compile them into a CSV file; Step 2: Construct the pollen group-plant functional type matrix from Step 1 and compile it into a CSV file; Step 3: Construct a biome region matrix of plant functional types from Step 2 and compile it into a CSV file; Step 4: Compile the pollen groups and their percentage data of the topsoil at 6 elevations in Lushan into a CSV file. Subtract the threshold (usually set to 0.5) from the percentage of pollen groups in the file and then take the square root (i.e., weight reduction processing). Step 5: Import the CSV file from Step 2 into R software, and use the subset function to filter out the matrix CSV file of pollen groups to plant functional types from Step 4; the filtering is completed using the subset function in R software.
[0026] Specifically, the pollen groups at the Lushan sampling point are a subset of the nationwide pollen group data established in step 1. The subset function extracts this subset from the whole set and saves it as a data frame.
[0027] Step 6: Use the ifelse function to determine whether the pollen groups of the sampling points in Step 5 are present or absent (1 or 0) in the pollen group to plant functional type matrix, and use the colSums function to sum the values of all pollen groups corresponding to each plant functional type of the sampling points. Step 7: Import the CSV file from Step 3 into R software. Use the ifelse function to determine whether the pollen groups from the sampling points in Step 6 are present or absent (1 or 0) in the matrix of plant functional types in the biome. Use the rowSums function to sum the scores of all plant functional types in each biome. The judgment is based on the value 0 or 1 in the matrix. If the value is 1, it is judged as present; if the value is 0, it is judged as absent.
[0028] The summation in steps 6 and 7 reflects the summation over all pollen groups in the following biome regionalization algorithm formula: In the formula, A ik This represents the similarity between pollen sample k and biome region i. This similarity is determined by the biome region type corresponding to the pollen sample, calculated using the biome regionization formula. The final calculated biome region type is considered to be similar to that of the pollen sample. S j This is the summation of all pollen groups j, d ij P represents the presence or absence (1 or 0) of biota region i and pollen group j in the biota region and pollen group matrix. jk It is the percentage content of pollen, Q j This is the threshold for the percentage of pollen. This threshold is set to reduce the weight of certain overexpressed pollen groups. To facilitate comparison of results, 0.5% is uniformly set as the standard threshold.
[0029] Step 8: Use the max.col function to find the first maximum value in each biome region from Step 7, and merge it with the file from Step 4. Output a CSV file, which is the final biome region calculation result.
[0030] Based on pollen data from the topsoil at an altitude of 400 meters in Lushan, the biome type calculated is subtropical evergreen broad-leaved forest.
[0031] The above description is merely a specific embodiment of the present invention, and the scope of protection of the present invention is not limited thereto. Any transformations or substitutions that can be conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A biome regionalization algorithm based on R software, characterized in that, Includes the following steps: Step 1: Collect and organize pollen datasets, compiling pollen data containing information on sampling point latitude and longitude, altitude, different pollen groups and their percentages into a CSV file; Step 2: Establish a matrix of plant functional types and pollen groups; Step 3: Establish a biome region-to-plant functional type matrix; Step 4: Import the pollen groups and their percentage data corresponding to the sampling points into R software for weighted calculation; Step 5: Select the pollen group pair plant functional type matrix from the plant functional type pair pollen group matrix of the sampling points from the plant functional type pair pollen group matrix of Step 2; Step 6: Determine whether the pollen group at the sampling point is present or absent in the pollen group to plant functional type matrix. If present, assign the value calculated in Step 4 to the corresponding plant functional type. If absent, assign the corresponding plant functional type a value of 0. Sum the values of all pollen groups corresponding to each plant functional type at the sampling point to obtain the plant functional type score for each sampling point. Step 7: Determine whether each plant functional type at the sampling point exists or not in the matrix of plant functional types in the biome region. If it exists, assign the score of the plant functional type to the biome region corresponding to that plant functional type. If it does not exist, assign the value of the corresponding biome region to 0. Sum the scores of all plant functional types in each biome region at the sampling point to obtain the biome region score. Step 8: Select the biome type with the highest similarity score for each sampling point and output the calculation results.
2. The biome regionalization algorithm based on R software according to claim 1, characterized in that, In step 2, the rule for establishing the plant functional type to pollen group matrix is as follows: the pollen groups in step 1 are merged into the corresponding plant functional types. If the pollen group belongs to the plant functional type, it is assigned a value of 1; if it does not belong to the plant functional type, it is assigned a value of 0.
3. The biome regionalization algorithm based on R software according to claim 1, characterized in that, In step 3, the rule for establishing the biome region to plant functional type matrix is as follows: define the biome region based on the key plant functional type combination in step 2. If the plant functional type belongs to the biome region, assign a value of 1; otherwise, assign a value of 0.
4. The biome regionalization algorithm based on R software according to claim 1, characterized in that, In step 4, the weight reduction calculation refers to subtracting the threshold from the percentage of pollen groups and then taking the square root.
5. The biome regionalization algorithm based on R software according to claim 1, characterized in that, In step 5, the filtering is performed using the subset function in R software.
6. The biome regionalization algorithm based on R software according to claim 1, characterized in that, In step 6, the ifelse function in the R language is used to make the judgment.
7. The biome regionalization algorithm based on R software according to claim 1, characterized in that, In step 6, the colSums function is used to sum all pollen groups for each plant functional type.
8. The biome regionalization algorithm based on R software according to claim 1, characterized in that, The summation in steps 6 and 7 reflects the summation of all pollen groups in the following biomesarization algorithm formula: ; In the formula, A ik It is the similarity between pollen sample k and biota region i, S j This is the summation of all pollen groups j, d ij P represents the presence or absence of biota region i and pollen group j in the biota region and pollen group matrix. jk It is the percentage content of pollen, Q j It is the threshold for the percentage of pollen.
9. The biome regionalization algorithm based on R software according to claim 1, characterized in that, In steps 6 and 7, there may be a corresponding value of 1 or 0; In step 6, if the value is 1, the value calculated in step 4 is assigned to the plant functional type corresponding to the pollen group; if the value is 0, the corresponding plant functional type is assigned the value 0. In step 7, if the value is 1, the score of the plant functional type is assigned to the biome region corresponding to that plant functional type; if the value is 0, the corresponding biome region is assigned a value of 0.