A method for identifying castanea plant varieties based on digital analysis of leaf morphology

By establishing a database of genus *Castanopsis* varieties and analyzing leaves using image recognition and geometric morphology measurement methods, the problem of confusion in the identification of *Castanopsis* varieties has been solved, enabling accurate and rapid variety identification and improving the utilization rate of *Castanopsis* plant resources.

CN115060720BActive Publication Date: 2025-11-11BEIJING FORESTRY UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210441184.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-25
Publication Date
2025-11-11
Estimated Expiration
2042-04-25

AI Technical Summary

Technical Problem

The lack of accurate and convenient methods for identifying varieties of the genus *Castanopsis* in existing technologies leads to variety confusion and affects their rational utilization and protection.

Method used

By establishing a database of chestnut species, using image recognition software to perform quantitative and graphical analysis of leaf morphology, selecting homology identification points, performing geometric morphological measurements, excluding outliers and asymmetric components, and performing data stratification, digital classification is achieved.

Benefits of technology

This method enables accurate and rapid identification of chestnut species, solves the problem of species confusion, and improves the utilization rate of chestnut resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115060720B_ABST
    Figure CN115060720B_ABST
Patent Text Reader

Abstract

This invention relates to the field of plant variety identification technology, specifically to a method for establishing a database for identifying *Castanea* varieties and a method for identifying *Castanea* varieties using the obtained database. This invention addresses the problems of variety confusion and inaccurate application in production caused by the high similarity of *Castanea* varieties. It involves quantitative and graphical analysis of the leaf morphology of *Castanea* varieties, extracting the main differentiating sites on the leaves of different *Castanea* varieties using geometric morphology measurement methods, establishing a variety identification database, achieving digital classification, and enabling accurate and rapid identification of *Castanea* varieties.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of plant variety identification technology, specifically to a method for establishing a database for identifying varieties of the genus *Castanopsis*, a database for identifying varieties of the genus *Castanopsis*, and a method for identifying varieties of the genus *Castanopsis* using the database. Background Technology

[0002] Chestnut trees belong to the genus *Castanea* of the family Fagaceae. There are approximately 10 species of *Castanea* worldwide, naturally distributed in 4 Asia, 4 Americas, 1 Europe, and 1 Africa. In China, there are three native species of *Castanea*: *Castanea mollissima* Bl., *Castanea henryi(skan) Rehd. et Wils*, and *Castanea seguinii* Dode. Chestnuts have been harvested for consumption as early as 6000 years ago and are known as a "staple crop," an important traditional woody grain tree species in China. Chestnuts are highly adaptable and widely distributed in my country, spanning cold temperate, temperate, warm temperate, subtropical, and marginal tropical zones. They are found in 22 provinces (autonomous regions and municipalities) except for Qinghai, Xinjiang, Inner Mongolia, and Ningxia. In terms of production area distribution, production is mainly concentrated in Hubei, Shandong, Hebei, Henan, and Anhui provinces, with six major cultivar groups: North China, Yangtze River mid-lower reaches, Northwest, Southwest, Southeast, and Northeast. The Japanese chestnut (Castanopsis fargesii) is named for its cone-shaped nut. It grows rapidly, is a deciduous tree, and has a sweet taste. In recent years, with the continuous increase in planting area and yield, the production of Japanese chestnut has received increasing attention. Japanese chestnut and *Castanopsis fargesii* are mainly distributed in the vast subtropical hilly and mountainous areas south of the Qinling Mountains (Zhang Yuhe, Liu Liu, Liang Weijian, et al. 2005. *Chinese Fruit Trees: Chestnut and Hazelnut Volume* [M]. Beijing: China Forestry Publishing House).

[0003] The fruits of the *Castanea* genus contain abundant starch, as well as protein, fat, B vitamins, and other nutrients, making their nutritional value comparable to rice. They also possess the advantages of wheat and surpass corn or rice, boasting the characteristic of "one generation planting, multiple generations enjoying," earning them the nickname "money tree." They integrate ecological, economic, social, carbon sequestration, and cultural functions. Currently, the development and utilization of *Castanea* genus plant resources lacks precise allocation and a complete identification system. Confusion between *Castanea* species (such as chestnut or Japanese chestnut) leads to underdevelopment and low utilization rates. Therefore, accurate identification of *Castanea* species in practical production is crucial for their rational utilization, variety protection, and introduction and cultivation. Exploring accurate, convenient, and practical methods for identifying *Castanea* species is particularly important. Summary of the Invention

[0004] This invention aims to at least partially solve one of the technical problems in related technologies. Therefore, one objective of this invention is to provide a method for establishing a database for identifying *Castanopsis* species, and a method for identifying *Castanopsis* species using the obtained database. This method involves quantitative and graphical analysis of the leaf morphology of *Castanopsis* species, comparing leaf shape differences among different *Castanopsis* species, achieving digital classification, and realizing accurate and rapid species identification.

[0005] Therefore, the first aspect of the present invention provides a method for establishing a database for the identification of *Castanopsis* species. According to an embodiment of the present invention, the method includes:

[0006] (1) Select multiple varieties of chestnut plants in the fruiting period from different regions and collect leaves at the physiological maturity stage.

[0007] (2) The back of the leaf is scanned, wherein the leaves from each of the chestnut species are scanned using uniform scanning parameters and angles;

[0008] (3) Use image recognition software to select identification points for all the leaves and obtain the coordinate data of each identification point of each leaf in order to establish a database of first leaf outline identification points for different chestnut plant varieties.

[0009] (4) Preprocess the first blade profile identification point database to remove outliers and asymmetric components in order to obtain the second blade profile identification point database.

[0010] (5) The second leaf outline identification point database is stratified to obtain a database for the identification of chestnut plant varieties.

[0011] This invention addresses the problem of varietal confusion and inaccurate application in production caused by the high similarity of *Castanea* species. It extracts the main morphological differences in leaves from different *Castanea* species using geometric morphology measurement, establishes a varietal identification database, and enables rapid identification of *Castanea* species. According to one embodiment of this invention, 80 *Castanea* species from 11 provinces (municipalities) across China were selected, and all leaf coordinate data were obtained to construct a database. Then, geometric morphology measurement software was used to perform quantitative and graphical analysis of leaf shape data among varieties, identify the species to be tested, and analyze the main morphological differences in leaves among varieties.

[0012] According to an embodiment of the present invention, the chestnut plant varieties in the fruiting period mentioned in step (1) include at least 80 varieties.

[0013] According to an embodiment of the present invention, at least 10 plants of each of the Chestnut species are collected.

[0014] According to an embodiment of the present invention, leaves of each chestnut plant are collected from at least four different locations, with an angle difference of 60°-90° between adjacent locations.

[0015] According to an embodiment of the present invention, the scanning parameters are: resolution 300-600 dpi, brightness 0-30 L.

[0016] Resolution and brightness affect the clarity and light transmittance of scanned images. Too high a resolution or too much brightness will cause the leaves to be translucent, appear greenish, and have unclear veins. Too low a resolution will result in dark gray leaves with indistinct leaf edges. Resolution also affects the size of the formed image; using consistent parameters allows for comparison of leaf size, which is beneficial for morphological comparison at the same level.

[0017] According to an embodiment of the present invention, the identification point includes the homologous point of the outer edge contour of the leaf of the chestnut plant variety in the fruiting stage described in step (1).

[0018] Homologous points refer to points common to the leaves of all selected *Cassia* species. Homologous points along the outer margin of the leaf are selected as identification points because they more accurately characterize the size and overall morphology of the leaf, are representative points in morphological research, and are more conducive to identification in practical life. According to the embodiment of this invention, at least 14 identification points are selected for each leaf. These 14 identification points are primary identification points, namely:

[0019] The identification points are as follows: the first serration pointing upwards from the petiole; the serration pointing at the widest part of the leaf; the serration pointing upwards from the widest part of the leaf; the serration pointing downwards from the widest part of the leaf; the serration pointing downwards from the leaf apex; the serration pointing downwards from the leaf apex; and the serration pointing downwards from the leaf apex. Each of these primary identification points is paired and located on both sides of the leaf vein. These 14 primary identification points play an important role in variety classification, and they can further improve the accuracy of identifying *Castanea* varieties using the established database.

[0020] According to an embodiment of the present invention, for each leaf, the identification point further includes three secondary identification points: the leaf tip identification point, the intersection of the leaf veins at the widest serration, and the junction of the petiole and the leaf blade.

[0021] According to an embodiment of the present invention, for each blade, the identification point further includes 7 supplementary identification points, namely:

[0022] The indentation at the upper part of the serration at the widest point of the leaf, the identification point of the first serration downward from the tip of the leaf, the intersection of the main vein immediately adjacent to the center point, the intersection of the veins of the first serration upward from the petiole, and the starting point of the petiole.

[0023] The indentation at the upper part of the serration at the widest point of the leaf and the identification point of the first serration downward from the leaf tip are paired and distributed on both sides of the leaf vein. The intersection of the main vein immediately adjacent to the center point, the intersection of the veins of the first serration upward from the petiole, and the starting point of the petiole are individual identification points. According to an embodiment of the present invention, the preprocessing of the first leaf contour identification point database further includes:

[0024] A generalized Protodyakonov analysis is performed on the first blade profile identification point database to maximize the concentration of coordinate points of all blades.

[0025] Generalized Protodyakonov analysis is performed to concentrate the coordinates of all leaves together to the greatest extent possible. This eliminates interference from non-shape factors such as placement and orientation, and also separates leaf shape from size.

[0026] According to an embodiment of the present invention, in step (4), outliers are eliminated by performing a Fourier transform on the coordinate data to form a new dataset for analysis. This process excludes leaves with significant morphological variations due to environmental factors, making the analysis results more accurate.

[0027] According to an embodiment of the present invention, the second leaf outline identification point database is stratified according to at least one of the following: origin region, variety, and individual plant.

[0028] According to an embodiment of the present invention, in step (5), when the second leaf outline identification point database is stratified according to variety, average leaf shape data is created. Creating average leaf shape data facilitates subsequent comparisons between varieties and regions.

[0029] A second aspect of the present invention provides a database for the identification of species of the genus *Castanea*. According to an embodiment of the present invention, the database is obtained by the establishment method described in the first aspect.

[0030] A third aspect of this invention provides a method for identifying varieties of the genus *Castanopsis*. According to an embodiment of the invention, the identification method includes:

[0031] 1) Collect leaves at physiological maturity from the species of the genus *Castanopsis* to be identified;

[0032] 2) Scan the back of all leaves from the chestnut species to be identified, using uniform scanning parameters and angles;

[0033] 3) Using image recognition software, homology identification points are selected for all scanned leaves of the chestnut species to be identified, and the homology identification points are compared with the database for chestnut species identification described in the second aspect. Through canonical variable analysis, the Protodyakonov distance matrix and variety scatter plot are obtained.

[0034] 4) Based on the Protodyakonov distance matrix and the position of the scatter plot of the varieties, determine the species of the *Castanopsis* genus to be identified.

[0035] The method for identifying *Castanopsis* species provided by this invention is based on quantitative and graphical analysis of leaf morphology, and is suitable for many *Castanopsis* species that are highly similar and cannot be distinguished in production. It enables digital classification, achieving accurate and rapid species identification.

[0036] According to an embodiment of the present invention, in step 1), the species of the *Castanopsis* to be identified originates from the source of *Castanopsis* species in a database used for the identification of *Castanopsis* species.

[0037] According to an embodiment of the present invention, at least 10 plants of the *Castanopsis* species to be identified are collected, and leaves of each plant are collected from at least 4 different locations, with an angle difference of 60°-90° between adjacent locations.

[0038] According to an embodiment of the present invention, in step 2), the scanning parameters are: resolution 300-600 dpi, brightness 0-30L.

[0039] According to an embodiment of the present invention, in step 3), the homology identification point includes the homology point of the outer edge contour of the leaf of the chestnut plant variety to be identified.

[0040] According to an embodiment of the present invention, at least 14 homology identification points are collected for each blade.

[0041] According to an embodiment of the present invention, for each blade, the homology identification points include 14 primary identification points, namely:

[0042] The identification points are: the first serration point upward from the petiole, the serration point at the widest part of the leaf, the serration point upward from the widest part of the leaf, the serration point downward from the widest part of the leaf, the second serration point downward from the leaf tip, the third serration point downward from the leaf tip, and the fourth serration point downward from the leaf tip. The primary identification points at each location are distributed in pairs on both sides of the vein of the leaf.

[0043] According to an embodiment of the present invention, the homology identification point further includes three secondary identification points, namely, the leaf tip identification point, the intersection of the leaf veins at the widest serration, and the junction of the petiole and the leaf blade.

[0044] According to an embodiment of the present invention, the homology identification points further include 7 supplementary identification points, namely:

[0045] The indentation at the upper part of the serration at the widest point of the leaf, the identification point of the first serration downward from the tip of the leaf, the intersection of the main vein immediately adjacent to the center point, the intersection of the veins of the first serration upward from the petiole, and the starting point of the petiole.

[0046] Among them, the indentation point at the upper part of the serration at the widest point of the leaf and the identification point of the first serration downward from the tip of the leaf are distributed in pairs on both sides of the leaf vein. The main vein intersection point adjacent to the center point, the vein intersection point of the first serration upward from the petiole, and the petiole starting point are individual identification points.

[0047] The fourth aspect of the present invention provides the use of the database for identifying varieties of the genus *Castanopsis* obtained by the establishment method described in the first aspect or the database for identifying varieties of the genus *Castanopsis* described in the second aspect in the identification of varieties of the genus *Castanopsis*.

[0048] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0049] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:

[0050] Figure 1 This invention shows the location of the identification point on a chestnut leaf according to an embodiment of the invention. The chestnut leaf is derived from the chestnut variety "Yanshan Short Branch" and originates from Qianxi County, Hebei Province.

[0051] Figure 2 The flowchart of a chestnut variety identification method according to an embodiment of the present invention is shown, wherein ①- collected leaves; ②- scanner; ③- scanned leaves; ④- ImageJ image recognition software; ⑤- identification point location; ⑥- MorphoJ geometric morphology measurement and analysis software; ⑦- symmetric and asymmetric components; ⑧- canonical variable analysis;

[0052] Figure 3 The symmetric component CVA analysis based on variety level is shown. Detailed Implementation

[0053] The present invention will now be described with reference to specific embodiments. It should be noted that these embodiments are merely descriptive and do not limit the present invention in any way.

[0054] Unless otherwise specified, all reagents used in the experiments of the examples are commercially available.

[0055] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this disclosure, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0056] Traditionally, the most common method for plant identification has been manual identification, which mainly utilizes traditional morphological measurement methods and human sensory observation to obtain characteristics such as the shape, color, and odor of plant branches, leaves, flowers, fruits, and bark, and then compares them between species. Leaves, as a salient observational structure of plants, have advantages such as long observation periods and easy acquisition, playing a crucial role in plant identification. With the development of science and technology, DNA molecular marker technology, due to its objectivity and accuracy, is applied to the identification of plant authenticity and purity, but its practical applicability in production is relatively low. Plant phenotypic traits, as intuitive results, are more convenient and practical for production. However, traditional morphological measurement methods suffer from drawbacks such as being time-consuming, labor-intensive, inefficient, highly subjective, and prone to large measurement errors. In the development of morphology, Geometric Morphometry (GMM) emerged, a method for multivariate statistical analysis of Cartesian coordinate data. By quantifying and graphically representing the morphology of different plants, it analyzes the morphological differences and relationships between similar species. This method is low-cost and easy to implement in practical production applications, and it can effectively distinguish differences caused by geographical origin, thus providing a guarantee for plant identification and classification.

[0057] my country boasts a rich variety of *Castanea* species. Previous research on the phenotypic traits of *Castanea* has focused on analyzing genetic diversity, lacking a systematic approach to identification. However, by employing geometric morphometry to precisely classify chestnut leaves and construct a database of leaf shape coordinates for *Castanea* species, a more efficient and intuitive method can be used to reflect the differences in leaf morphology between species and varieties. This provides a precise and convenient method for phenotypic identification of *Castanea* species, ensuring their introduction and utilization across regions and supporting processing and production, thus possessing significant practical value.

[0058] To this end, the inventors have developed a method for identifying varieties of the genus *Castanopsis*, based on quantitative and graphical analysis of leaf morphology. This method is suitable for many *Castanopsis* varieties (such as chestnut and Japanese chestnut) that are highly similar and difficult to distinguish in production. It enables digital classification, achieving accurate and rapid variety identification.

[0059] First, the inventors developed a method for establishing a database for the identification of *Castanea* species. According to a specific embodiment of the present invention, the method includes:

[0060] (1) Select multiple varieties of chestnut plants in the fruiting period from different regions and collect leaves at the physiological maturity stage.

[0061] (2) The back of the leaf is scanned, wherein the leaves from each of the chestnut species are scanned using uniform scanning parameters and angles;

[0062] (3) Use image recognition software to select identification points for all the leaves and obtain the coordinate data of each identification point of each leaf in order to establish a database of first leaf outline identification points for different chestnut plant varieties.

[0063] (4) Preprocess the first blade profile identification point database to remove outliers and asymmetric components in order to obtain the second blade profile identification point database.

[0064] (5) The second leaf outline identification point database is stratified to obtain a database for the identification of chestnut plant varieties.

[0065] According to a specific embodiment of the present invention, the different regions can be, for example, different provinces or cities, and there are at least 10 different regions. The more regions there are, the more varieties can be established in the database, which is more conducive to matching with the varieties to be tested during identification.

[0066] According to a specific embodiment of the present invention, for each region, a variety of chestnut species can be selected, and the chestnut species included in the database include at least 80, or even 90, 100 or more chestnut species.

[0067] According to a specific embodiment of the present invention, for each region, multiple chestnut varieties can be selected, and the chestnut varieties included in the database can be at least 80, or even 90, 100, or more. These chestnut varieties are all approved and confirmed varieties. The more chestnut varieties in the database, the more beneficial it is for variety identification.

[0068] According to a specific embodiment of the present invention, at least 10 plants of each of the chestnut species are collected. If the number of collected plants is too small, the data obtained cannot accurately characterize the leaf characteristics of the chestnut species.

[0069] According to a specific embodiment of the present invention, leaves of each chestnut plant are collected from at least four different locations, with an angle difference of 60°-90° between adjacent locations. For example, leaves can be collected from the east, west, south, and north directions respectively.

[0070] According to one specific embodiment of the present invention, the scanning parameters are: resolution 300-600 dpi, brightness 0-30L. According to another embodiment of the present invention, the scanning parameters may be, for example, a resolution of 600 dpi and a brightness of 30L. The scanning angle is not limited; for example, it can be vertical or horizontal, but it is necessary to ensure that the direction is uniform when all blades are scanned.

[0071] According to a specific embodiment of the present invention, the identification points include the homologous points of the outer edge contours of the leaves of all chestnut plant varieties in the fruiting stage described in step (1).

[0072] Homologous points are points shared by leaves of all selected chestnut species. They have relative positional consistency and repeatability. Only by selecting homologous points can leaf overprints of the same species be made, as well as comparisons between different species.

[0073] The basic principles for selecting leaf identification points are: all leaves must be homologous (shared by all leaves); they must be sufficiently representative; they must be able to reflect the morphological and structural information of the research sample; their relative positions must be consistent; they must be repeatable, and they must be able to be marked intuitively and accurately during the repetition process.

[0074] According to a specific embodiment of the present invention, for each blade, at least 14 identification points are selected. These 14 identification points are primary identification points, namely:

[0075] The identification points are: the first serration point upward from the petiole, the serration point at the widest part of the leaf, the serration point upward from the widest part of the leaf, the serration point downward from the widest part of the leaf, the second serration point downward from the leaf tip, the third serration point downward from the leaf tip, and the fourth serration point downward from the leaf tip. The primary identification points at each location are distributed in pairs on both sides of the vein of the leaf.

[0076] In this invention, "upward" refers to the direction from the petiole towards the tip of the blade, and "downward" refers to the direction from the tip of the blade towards the petiole.

[0077] According to a specific embodiment of the present invention, for each leaf, in addition to the 14 primary identification points mentioned above, the selected identification points further include 3 secondary identification points. The 3 secondary identification points are the leaf tip identification point, the intersection of the leaf veins at the widest serration, and the junction of the petiole and the leaf blade.

[0078] According to a specific embodiment of the present invention, for each blade, in addition to the aforementioned 14 primary identification points and 3 secondary identification points, the selected identification points further include 7 supplementary identification points, namely:

[0079] The indentation at the upper part of the serration at the widest point of the leaf, the identification point of the first serration downward from the tip of the leaf, the intersection of the main vein immediately adjacent to the center point, the intersection of the veins of the first serration upward from the petiole, and the starting point of the petiole.

[0080] Among them, the indentation point at the upper part of the serration at the widest point of the leaf and the identification point of the first serration downward from the tip of the leaf are distributed in pairs on both sides of the leaf vein. The main vein intersection point adjacent to the center point, the vein intersection point of the first serration upward from the petiole, and the petiole starting point are individual identification points.

[0081] The basic principles for selecting specific identification points in this invention are as follows:

[0082] ① Describe the overall shape and size of the leaves of plants in the genus *Castanopsis* (such as chestnut);

[0083] ②The homologous sites shared by all leaves of the experimental chestnut species showed correlation and reproducibility among different samples;

[0084] ③ The identification points on both sides are in the same relative position, for example, the first serration from the base and the widest serration are both selected (which is more conducive to separating symmetrical and asymmetrical components. Symmetrical components refer to shape traits affected by heredity. Asymmetrical components refer to the random deviation of the bilateral symmetry of the leaf caused by environmental factors, i.e., leaf shape variation).

[0085] ④ Based on the morphological characteristics of chestnut leaves, the upper part of the leaves has denser serrations and there are significant differences in the upper part of the leaves among different varieties. Therefore, more identification points are selected from the upper part than from the lower part.

[0086] According to a preferred embodiment of the present invention, at least 24 identification points are collected for each leaf, including the aforementioned 14 primary identification points, 3 secondary identification points, and 7 supplementary identification points. Of course, the number of identification points can also be 30, 40, or more. However, collecting 24 identification points for each leaf yields a database sufficient for preparing to identify *Castanopsis* species, ensuring a high accuracy rate in identification.

[0087] According to a specific embodiment of the present invention, the image recognition software can be ImageJ image recognition software or other image recognition software known in the art, as long as it can identify the coordinate position and obtain accurate identification points.

[0088] According to a specific embodiment of the present invention, the preprocessing of the first blade contour identification point database further includes:

[0089] A generalized Protodyakonov analysis is performed on the first blade contour identification point database to maximize the concentration of coordinate points of all the blades together. This eliminates interference from non-shape factors such as placement and orientation, and also separates leaf shape from size.

[0090] According to a specific embodiment of the present invention, outliers are eliminated by performing a Fourier transform on the coordinate data, resulting in a new dataset for analysis. This process excludes leaves with significant morphological variations due to environmental factors, making the analysis results more accurate.

[0091] According to a specific embodiment of the present invention, asymmetric components are excluded from the coordinate data, while symmetric components are retained. Symmetrical components refer to shape traits influenced by heredity. Asymmetric components refer to leaf characteristics influenced by environmental factors.

[0092] According to a specific embodiment of the present invention, the second leaf outline identification point database is stratified according to at least one of the following: origin region, variety, and individual plant.

[0093] According to a specific embodiment of the present invention, in step (5), when the second leaf outline identification point database is stratified according to variety, average leaf shape data is created.

[0094] According to a specific embodiment of the present invention, the present invention provides a method for identifying varieties of the genus *Castanopsis*, the method comprising:

[0095] 1) Collect leaves at physiological maturity from the species of the genus *Castanopsis* to be identified;

[0096] 2) Scan the back of all leaves from the chestnut species to be identified, using uniform scanning parameters and angles;

[0097] 3) Using image processing software, homology identification points were selected for all scanned leaves of the chestnut species to be identified, and the homology identification points were compared with a database used for chestnut species identification. Through canonical variable analysis, the Protodyakonov distance matrix and variety scatter plot were obtained.

[0098] 4) Based on the Protodyakonov distance matrix and the scatter plot of the varieties, determine the species of the *Castanopsis* genus to be identified.

[0099] According to a specific embodiment of the present invention, in step 1), the species of the *Castanopsis* to be identified originates from the source of *Castanopsis* species in a database used for the identification of *Castanopsis* species.

[0100] According to a specific embodiment of the present invention, in step 1), at least 10 plants of the *Castanopsis* species to be identified are collected, and leaves of each plant are collected from at least 4 different locations, with an angle difference of 60°-90° between adjacent locations. For example, leaves of the *Castanopsis* species to be identified can be collected from the east, west, south, and north directions respectively.

[0101] According to one specific embodiment of the present invention, in step 2), the scanning parameters are: resolution 300-600 dpi, brightness 0-30L. According to another embodiment of the present invention, the scanning parameters may be, for example, a resolution of 600 dpi and a brightness of 30L. The scanning angle is not limited; for example, it can be vertical or horizontal, but it is necessary to ensure that the direction is uniform when all blades are scanned.

[0102] According to a specific embodiment of the present invention, in step 3), the homology identification point includes the homology point of the outer edge contour of the leaf of the chestnut plant variety to be identified.

[0103] According to a specific embodiment of the present invention, for each blade, at least 14 primary identification points of homology identification are collected, namely:

[0104] The identification points are: the first serration point upward from the petiole, the serration point at the widest part of the leaf, the serration point upward from the widest part of the leaf, the serration point downward from the widest part of the leaf, the second serration point downward from the leaf tip, the third serration point downward from the leaf tip, and the fourth serration point downward from the leaf tip. The primary identification points at each location are distributed in pairs on both sides of the vein of the leaf.

[0105] According to a specific embodiment of the present invention, for each leaf, in addition to the 14 primary identification points mentioned above, the selected homology identification points further include 3 secondary identification points. The 3 secondary identification points are the leaf tip identification point, the leaf vein intersection point of the widest serration, and the junction point of the petiole and the leaf blade.

[0106] According to a specific embodiment of the present invention, for each blade, in addition to the aforementioned 14 primary identification points and 3 secondary identification points, the selected homology identification points further include 7 supplementary identification points, namely:

[0107] The indentation at the upper part of the serration at the widest point of the leaf, the identification point of the first serration downward from the tip of the leaf, the intersection of the main vein immediately adjacent to the center point, the intersection of the veins of the first serration upward from the petiole, and the starting point of the petiole.

[0108] Among them, the indentation point at the upper part of the serration at the widest point of the leaf and the identification point of the first serration downward from the tip of the leaf are distributed in pairs on both sides of the leaf vein. The main vein intersection point adjacent to the center point, the vein intersection point of the first serration upward from the petiole, and the petiole starting point are individual identification points.

[0109] According to a specific embodiment of the present invention, the method for identifying varieties of the genus *Castanopsis* is as follows: Figure 2 As shown, the process includes collecting leaf samples → scanning the leaf samples with a scanner → obtaining the identification point locations using ImageJ image recognition software → performing canonical variable analysis using MorphoJ geometric morphology measurement and analysis software (asymmetric components need to be removed).

[0110] According to one embodiment of the present invention, a method for identifying varieties of *Castanopsis* includes temporarily preserving standard-collected leaves in a specimen clip, scanning the back of the leaves with a scanner to generate an A4-sized image, selecting 24 identification points using Image J to obtain coordinate data, performing canonical variable analysis using Morpho J 1.07a, and determining the variety based on the scatter plot position and Protodyakonov distance matrix.

[0111] According to a specific embodiment of the present invention, the requirements for leaves of the *Castanopsis* species to be identified are as follows: the species from the source region that exist in the database are selected, and the number of leaves reaches the biomass sample (≥30); the tree is required to be 8 years old or above, with good growth, and complete leaves without disease or pests are selected from the four directions of the outer and middle canopy of the tree.

[0112] According to a specific embodiment of the present invention, in the method for identifying varieties of the genus *Castanopsis*, the back of all selected leaves of the variety to be identified are scanned, with uniform parameters and angles, a resolution of 300–600 dpi, A4 vertical orientation, and a brightness of 0–30 L, ensuring that the veins on the back of the leaves are clear.

[0113] According to a specific embodiment of the present invention, ImageJ is used to select homology identification points from all scanned leaves of the variety to be identified, and canonical variable analysis is performed on the coordinate data of all varieties in the database. This yields a Protodyakonov distance matrix and a scatter plot of the varieties. Based on the distance and the position of the scatter plot, the closest variety is identified as the variety to be determined.

[0114] Cross-validation using discriminant analysis was used to prove the success rate of the discrimination. A grid variation diagram generated through canonical variable analysis was used to determine the changing trends of leaf morphology among varieties.

[0115] According to a specific embodiment of the present invention, the present invention provides a method for chestnut variety identification based on digital analysis of leaf morphology, the steps of which are as follows:

[0116] 1) Select well-developed and robust chestnut varieties to be identified, and collect leaves at the physiological maturity stage of the leaves. Collect 2 complete leaves from each variety and 2 leaves from each of the four directions of east, south, west and north on the outer periphery of the canopy layer of each tree. Collect 8 leaves per tree and 80 leaves per variety.

[0117] 2) Scan the back of all selected leaves of the variety to be identified, and standardize the parameters and angles;

[0118] 3) Using ImageJ, 24 homology identification points were selected from all scanned leaves of the variety to be identified. These points were compared with a chestnut variety coordinate point database. Through canonical variable analysis, a Protodyakonov distance matrix and a variety scatter plot were obtained. Based on the distance and position in the scatter plot, the closest variety was selected as the identified variety. The success rate of the discrimination was demonstrated using cross-validation of discriminant analysis.

[0119] According to a specific embodiment of the present invention, the specific locations of the 24 homology identification points are as follows: Figure 1 And as shown in Table 1 below:

[0120] Table 1. Description of Leaf Identification Point Locations

[0121]

[0122] Table 1 lists the names of the identification points and... Figure 1 The identification points in the table correspond one-to-one. For example, identification point IM1-2 in Table 1 refers to... Figure 1 The positions marked with the numbers "1" and "2" on the middle leaf blade. Among them, IM1-2, 3-4, 5-6, 7-8, 9-10, 11-12, and 13-14 are primary identification points; IM19, 21, and 23 are secondary identification points; and IM15, 16, 17, 18, 20, 22, and 24 are supplementary identification points. IM1-IM21 (21 identification points in total) represent the complete chestnut leaf outline; IM2, 22, 23, and 24 (4 identification points in total) describe the position and length of the main vein; IM23 represents the center position of the vein. Selecting the center point allows for better Protodyakonov overprint analysis and measurement of the centroid size. Leaf size is obtained by measuring the centroid of the outline structure formed by the identification points (the centroid is measured by the distance from the identification point to the center of the outline). IM23 and IM24 (two identification points in total) describe the spacing between the intersections of the central veins. Accurately displaying the center position of the veins can reflect the density of the primary veins to a certain extent.

[0123] The order of selecting identification points should be based on a combination of importance and convenience. First, select 14 primary identification points, then select other identification points on the leaf edges, and finally select identification points on the main veins in a top-to-bottom order. In actual operation, the order can be changed, but the selection order of all leaf points must be consistent.

[0124] According to a specific embodiment of the present invention, the method for identifying chestnut varieties comprises the following steps: scanning all leaves of the variety to be identified → selecting 24 homologous identification points → comparing with a chestnut variety coordinate point database → performing multivariate statistical data analysis, primarily using canonical variable analysis to obtain a Protodyakonov distance matrix, a variety scatter plot, and a grid variation map. Based on the distance and scatter plot position, the closest variety is identified as the determined variety. The grid variation map is used to analyze the leaf morphology change trends and main differences among varieties. Cross-validation using discriminant analysis is used to prove the success rate of the identification.

[0125] *Castanopsis* species have opposite, oblong or lanceolate leaves with serrated margins and numerous pinnate parallel lateral veins. *Chestnut chinensis* has simple, alternate leaves, 6–20 cm long and 4–10 cm wide; leaf blade ovate-elliptic, obovate-elliptic, or broadly lanceolate-elliptic; apex acuminate or acute, base cuneate or nearly cordate, leaf margin serrate, serrate or shallowly serrate; petiole about 1–2 cm long. *Chestnut arborescens* has simple, alternate leaves, 14–19 cm long and 4–5 cm wide; leaf blade oblong-ovate to ovate-lanceolate; apex caudate-acuminate, base cuneate to nearly rounded, leaf margin serrate; petiole about 1.5–2 cm long. (Zhang Yuhe, Liu Liu, Liang Weijian, et al. 2005. *Chinese Fruit Trees*, Chestnut and Hazelnut Volume [M]. Beijing: China Forestry Publishing House. pp. 22-25). The *Castanopsis* species in this invention include, but are not limited to, *Chestnut chinensis* and *Chestnut arborescens*. These types of chestnut plants have similar leaf characteristics. Although only the chestnut database and chestnut variety identification method are shown in the examples, the chestnut plant variety identification database and chestnut plant variety identification method of the present invention are also applicable to other chestnut plant varieties (e.g., Castanopsis fargesii).

[0126] The present disclosure will be explained below with reference to embodiments. Those skilled in the art will understand that the following embodiments are for illustrative purposes only and should not be construed as limiting the scope of the disclosure. Where specific techniques or conditions are not specified in the embodiments, they are performed in accordance with the techniques or conditions described in the literature in the art or according to the product instructions. Reagents or instruments whose manufacturers are not specified are all conventional products that can be obtained commercially.

[0127] Example 1

[0128] 1. Collection of different chestnut leaf samples

[0129] Eighty chestnut varieties that had entered their peak fruiting period were selected from chestnut-growing areas in 11 provinces (municipalities) across China. All varieties were selected from locally developed and robust trees. Leaves were collected at the physiological maturity stage, from 10 trees of each variety. Two complete leaves were collected from each tree at four locations on the east, south, west, and north sides of the outer canopy, totaling eight leaves per tree and eighty leaves per variety, for a total of 6400 leaves from the 80 varieties. These leaves were flattened using a specimen clip and then scanned and analyzed.

[0130] 2. Data Acquisition Methods

[0131] Wipe the collected leaves clean, removing any folded or damaged leaves. Scan the back of the leaves using an EPSON Scan scanner. Use uniform scaling parameters: 600 dpi resolution, A4 portrait orientation, 30 dpi brightness, and save the image as a .jpg file.

[0132] ImageJ software (Dai Zhicong, Du Daolin, Si Chuncan, et al. 2009, A method for accurately measuring leaf morphological and quantitative characteristics using a scanner and ImageJ software, Guangxi Plants, 29(3), 342-347) was used to select identification points for leaves of all varieties, including all homologous points of the outer edge contour of chestnut leaves (see Appendix for details). Figure 2 (The identification point locations are shown in section ⑤). 24 (x,y) coordinate data were obtained for each leaf and saved in .txt format to establish a database of leaf outline identification points for different chestnut qualities.

[0133] 3. Data Processing

[0134] 1) Data was processed using Microsoft Excel 2021. Individual ID naming convention: ID = 1 (region / group) - 01 (variety number) - 01 (individual number), imported into a .txt file.

[0135] 2) Import the coordinate point data into Morpho J 1.07 software: File → Create New Project → Select .txt format file → TPS file type → Object Symmetry → Create Dataset → Rename the imported data file

[0136] 3) Using Morpho J 1.07 software, perform generalized procrustes analysis (GPA) to maximize the concentration of coordinate points of all leaves together. This eliminates interference from non-shape factors such as placement and orientation, and also separates leaf shape from size. Specific steps: Preliminaries → New Procrustes Fit → Check if the identification point association is correct → Accept

[0137] 4) By performing a Fourier transform on the coordinate data, outliers are eliminated, creating a new dataset for analysis. This excludes leaves with significant morphological variations due to environmental factors, making the analysis more accurate. Specific steps: Preliminaries → Find Outliers

[0138] 5) This invention separates leaf shape data into symmetrical and asymmetrical components during data import. Symmetrical components refer to shape traits influenced by heredity. Asymmetrical components refer to leaf shape variations caused by environmental and other factors, resulting in random deviations in bilateral symmetry. This study's division of leaves into symmetrical and asymmetrical components allows for better analysis of leaf morphological changes and identification of key morphological differences.

[0139] Subsequent analysis of the results proved that the leaf morphology of different chestnut varieties can only be distinguished by the symmetrical components, while the asymmetrical components are random and have no specific rules, and therefore have no significance for identification.

[0140] 6) Data Stratification: The newly formed dataset is stratified by source region, variety, and individual, and an average leaf shape is created at the variety level to facilitate subsequent comparisons between varieties and regions. Specific steps: Preliminaries → Extract new classifier from ID strings → Name for new classifier: region, variety, individual; Preliminaries → Average observations by

[0141] 7) Multivariate statistical analysis includes Principal Component Analysis (PCA) to extract principal differences and key identification points. Specific steps: Preliminaries → Generate Covariance Matrices → Select all types to be used (Symmetric and Asymmetry components) → Accept → Variation → Principal Component Analysis

[0142] In the symmetrical component, the cumulative contribution rate of PC1 and PC2 was 80.6%; in the asymmetrical component, the cumulative contribution rates of PC1 and PC2 were 72.0% (see Table 2), which illustrates the variation in leaf morphology among varieties. Extensive research on leaf materials has demonstrated that, as... Figure 1The 24 identification points shown in Table 1 have varying degrees of importance in the identification process. Based on the comprehensive scores and contribution rates of the identification points obtained from principal component analysis (see Table 3), important identification sites were screened out. The reliability of the results was verified by combining the allometric growth analysis of leaves of different chestnut varieties (see Table 4). Together, they proved that 14 identification points, IM1-2, 3-4, 5-6, 7-8, 9-10, 11-12, and 13-14 (mainly representing the morphology and relative position of the leaf base, the widest part of the leaf, and the leaf tip), play an important role in variety classification, with a cumulative contribution rate of 70.14%.

[0143] Based on the results in Tables 3 and 4, in the process of chestnut variety identification, 14 identification points (IM1-2, 3-4, 5-6, 7-8, 9-10, 11-12, and 13-14) were used as primary identification points; 3 identification points (IM19, 21, and 23) were used as secondary identification points. The addition of IM19 and 23 made the leaf morphology more complete and facilitated the measurement of leaf length. IM21 is the center point of the leaf and is a key point for measuring leaf size; 7 identification points (IM15, 16, 17, 18, 20, 22, and 24) were used as supplementary identification points, with a cumulative contribution rate of 20.54%, which also played an important role in the overall identification.

[0144] Table 2. Top 5 principal components of 80 chestnut varieties based on symmetric and asymmetric components.

[0145]

[0146] Table 3. Comprehensive scores and contribution rates of identification points for 80 chestnut varieties based on principal component analysis of symmetric components.

[0147]

[0148] 8) Partial Least Squares (2B-PLS) was used for allometric growth analysis (AGA) to explore the symbiotic relationship between leaf morphology (symmetric and asymmetric components) and leaf size, i.e., allometric growth, to demonstrate the significant role of the extracted identification points. This validated the main identification points extracted by principal component analysis, and the results were consistent with those of principal component analysis (see Table 4). Specific steps: Covariation → Partial Least Squares → Two Separate → Blocks → Block 1: Log Centroid Size; Block 2: Symmetric component → Perform permutation test 10000

[0149] Table 4. Comprehensive scores of identification points for 80 chestnut varieties based on symmetric component allometric growth analysis.

[0150]

[0151] 9) Canonical Variance Analysis (CVA) compares the average leaf morphology differences among 10 regions and 80 chestnut varieties to achieve variety identification; Discriminant Analysis (DA), based on cross-validation and discriminant functions, is used to distinguish between the two and verify the reliability of the identification results. This invention is used for inter-regional discrimination. Specific steps: Comparison → Canonical Variance → Analysis → Date type: symmetric / asymmetry component Classifier variable(s) to use for grouping: region / variety; Comparison → Discriminant function analysis → Data type: symmetric component Classifier(s) to be used as grouping certification: region / variety (Include all pairs of groups, Permutation test: 10000)

[0152] A total of 45 discriminant analysis data sets were obtained from 10 regions. The results (see Table 5) show that Hubei and Anhui had the lowest discrimination rates, at 87.5% vs. 92.9%, while the discrimination rates of the remaining regions were all between 95% and 100%. This demonstrates that chestnut varieties can be basically completely distinguished between regions. A total of 3160 discriminant analysis data sets were obtained from 80 varieties (see Table 6). Among them, only 21 groups had discrimination rates below 100%. The varieties with the lowest discrimination rates were 'Yanlong' and 'Yanming' (93.33% vs. 84.62%), 'Libo Early Chestnut' and 'Libo Middle Chestnut' (86.67% vs. 92.30%), and 'Dongfeng' and 'Jinfeng' (96.67% vs. 84.62%). This demonstrates that 99.34% of the varieties in this analysis achieved a 100% discrimination accuracy rate, with only a few varieties having lower discrimination rates, all of which were above 80%. Cluster analysis (CA) was performed on the Mahalanobis distance matrix obtained from CVA using Origin 2021. A mesh variation diagram based on thin plate spline (TPS) was plotted to visualize the morphological changes of the blade.

[0153] Table 5. Discriminant analysis of leaf shape among 10 cultivation regions

[0154]

[0155] Table 6. Discriminant analysis of interleaf shape in 80 varieties

[0156] (Except for those shown in the table, the discrimination rate among all other varieties reached 100%)

[0157]

[0158] Example 2: Verification of Chestnut Variety Identification Method

[0159] 1. Morphological determination

[0160] Obtain leaves of the variety to be identified using the above steps, with variety numbers 1, 2, and 3 → Data acquisition → Data processing.

[0161] 1) Requirements for collecting samples from the blades to be tested

[0162] Select varieties from the source regions that exist in the database, with a leaf count reaching the biomass sample size (≥30); the trees should be 8 years old or older, with good growth, and select complete leaves from the outer and central canopy layers in the four cardinal directions (east, west, south, and north) free from pests and diseases.

[0163] 2) Data Acquisition

[0164] Scan the back of all selected leaves of the variety to be identified, using uniform parameters and angles, a resolution of 300-600 dpi, A4 vertical orientation, and a brightness of 0-30L, ensuring that the veins on the back of the leaves are clearly visible.

[0165] 3) Data processing

[0166] ImageJ was used to select homology identification points from all scanned leaves of the variety to be identified, and canonical variable analysis was performed on the coordinate data of all varieties in the database. This yielded a Protodyakonov distance matrix and a scatter plot of the varieties. Based on the distance and position in the scatter plot, the closest variety was selected as the identification variety. Cross-validation using discriminant analysis was used to verify the success rate of the discrimination. A grid variation diagram generated through canonical variable analysis was used to determine the changing trends of leaf morphology among varieties.

[0167] Canonical variable analysis and discriminant analysis were performed on the data. The canonical variable scatter plot shows (see...) Figure 2Three varieties to be identified almost completely overlapped with 35'Dabanhong', 36'Qianxizaohong', and 34'Yanshanzaofeng' in the database. CVA analysis showed that variety 1 had the smallest Mahalanobis distance (the smaller the distance, the greater the morphological similarity) with 'Dabanhong', variety 2 with 'Qianxizaohong', and variety 3 with 'Yanshanzaofeng', and the P-values ​​were also extremely insignificant, proving a very high degree of similarity (see Table 7). Therefore, the unknown varieties can be preliminarily identified as 'Dabanhong', 'Qianxizaohong', and 'Yanshanzaofeng'.

[0168] Table 7. Mahalanobis distance and discriminant analysis for leaf characteristics of different varieties (based on 10,000 replicates)

[0169]

[0170] 2. Molecular marker determination

[0171] 1) DNA was extracted from the leaves of the variety to be identified using the CTAB method;

[0172] 2) Select primers (see Table 8) to amplify clear and different bands among chestnut varieties for variety identification.

[0173] Table 8 SSR Primer Information

[0174]

[0175] 3) The PCR reaction system consisted of 20 μL containing 100-200 ng of genomic DNA, 0.8 pmol of a forward primer with the 5' end of the universal M13 sequence, 3.2 pmol each of a reverse primer and fluorescently labeled universal M13 primers (M13F(-47): 5′-CGCCAGGGTTTTCCCAGTCACGAC-3′ (SEQ ID NO:13), M13R Primer: 5′-CACACAGGAAACAGCTATGAC-3′ (SEQ ID NO:14)), and 10 μL of 2×Taq PCR Mix. The PCR amplification program was as follows: 94℃ pre-denaturation for 5 min; 94℃ denaturation for 30 s, 56℃ annealing for 30 s, 72℃ extension for 45 s, 30 cycles; 72℃ extension for 7 min followed by storage at 4℃.

[0176] 4) The PCR products were detected by capillary electrophoresis, and the size of the amplified fragments was read using Gene-Marker v 4.0 software. The results (see Table 9) showed that, based on the previously established fingerprint profiles, the varieties to be identified were 'Dabanhong', 'Qianxizaohong', and 'Yanshanzaofeng'.

[0177] Table 9. SSR fingerprints of three chestnut varieties

[0178]

[0179] The results showed that the identification results of the two methods were consistent. The varieties selected for the verification experiment were from the same varietal group. This invention found that the similarity between varieties from the same region was higher than that between varieties from different regions, thus proving that the accuracy of geometric morphology discrimination is extremely high. Compared with molecular identification, using geometric morphology methods for variety identification significantly shortens the identification time and saves costs, making it more suitable for application in production practice.

[0180] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," "some implementations," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0181] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention. SEQUENCE LISTING <110> Beijing Forestry University <120> A Method for Identifying Chrysanthemum Varieties Based on Digital Analysis of Leaf Morphology <130> BI3220441 <160> 14 <170> PatentIn version 3.3 <210> 1 <211> twenty three <212> DNA <213> Artificial <220> <223> ICMA005 forward primer <400> 1 aaataaaacc cctcatcaac aca 23 <210> 2 <211> twenty three <212> DNA <213> Artificial <220> <223> ICMA005 reverse primer <400> 2 gaactcaaaa cctcaaaacc tca 23 <210> 3 <211> 20 <212> DNA <213> Artificial <220> <223> ICMA018 forward primer <400> 3 acaacgatcc cagaccaaag 20 <210> 4 <211> 20 <212> DNA <213> Artificial <220> <223> ICMA018 reverse primer <400> 4 ctaggcgatc ggagagagac 20 <210> 5 <211> twenty three <212> DNA <213> Artificial <220> <223> CmCTR4 forward primer <400> 5 cataggttca aaccataccc gtg 23 <210> 6 <211> twenty four <212> DNA <213> Artificial <220> <223> CmCTR4 reverse primer <400> 6 ctcatctttg tagggtataa tacc 24 <210> 7 <211> 20 <212> DNA <213> Artificial <220> <223> CmCTR24 forward primer <400> 7 ctgcaagaca agaattacac 20 <210> 8 <211> 18 <212> DNA <213> Artificial <220> <223> CmCTR24 reverse primer <400> 8 gaataacctg cagaaggc 18 <210> 9 <211> 20 <212> DNA <213> Artificial <220> <223> CsCAT5 forward primer <400> 9 cattttctca ttgtggctgc 20 <210> 10 <211> 20 <212> DNA <213> Artificial <220> <223> CsCAT5 reverse primer <400> 10 cacttgcaca tccaattagg 20 <210> 11 <211> 20 <212> DNA <213> Artificial <220> <223> CsCAT8 forward primer <400> 11 ctgcaagaca agaattacac 20 <210> 12 <211> 18 <212> DNA <213> Artificial <220> <223> CsCAT8 anisotropy <400> 12 gaatacctg cagaaggc <210> 13 <211> 24 <212> DNA <213> Artificial <220> <223> M13F(‐47) <400> 13 cgccagggtt ttcccagtca cgac <210> 14 <211> 21 <212> DNA <213> Artificial <220> <223> M13R Primer <400> 14 cacacagga acagctatga c

Claims

1. A method for establishing a database for identifying varieties of the genus *Castanea*, characterized in that, include: (1) Select multiple varieties of chestnut plants in the fruiting period from different regions and collect leaves at physiological maturity. (2) The back of the leaf is scanned, wherein the leaves from each of the chestnut species are scanned using uniform scanning parameters and angles; (3) Use image recognition software to select identification points for all the leaves and obtain the coordinate data of each identification point of each leaf in order to establish a database of the first leaf outline identification points of different chestnut plant varieties; (4) Preprocess the first blade profile identification point database to remove outliers and asymmetric components in order to obtain the second blade profile identification point database. (5) The second leaf outline identification point database is stratified to obtain a database for the identification of *Castanopsis* species. The chestnut cultivars in the peak fruiting period in step (1) include at least 80 varieties. At least 10 plants of each of the aforementioned *Castanopsis* species were collected; Leaves were collected from each *Castanopsis* species at least four different locations, with an angle difference of 60°–90° between adjacent locations. The identification points include the homologous points of the outer edge contour of the leaves of chestnut plant varieties in the fruiting stage in step (1). For each leaf, the identification points include 14 primary identification points, 3 secondary identification points, and 7 supplementary identification points. The 14 primary identification points are as follows: The identification points are: the first serration point upward from the petiole, the serration point at the widest part of the leaf blade, the first serration point upward from the widest part of the leaf blade, the first serration point downward from the widest part of the leaf blade, the second serration point downward from the leaf tip, the third serration point downward from the leaf tip, and the fourth serration point downward from the leaf tip. The primary identification points at each location are paired and distributed on both sides of the vein of the leaf. The three secondary identification points are the leaf tip identification point, the intersection of the leaf veins at the widest serration, and the junction of the petiole and the leaf blade. The seven supplementary identification points are as follows: The indentation at the upper part of the serration at the widest point of the leaf, the identification point of the first serration downward from the tip of the leaf, the intersection of the main vein immediately adjacent to the center point, the intersection of the veins of the first serration upward from the petiole, and the starting point of the petiole. Among them, the indentation point at the upper part of the serration at the widest point of the leaf and the identification point of the first serration downward from the tip of the leaf are distributed in pairs on both sides of the leaf vein. The main vein intersection point adjacent to the center point, the vein intersection point of the first serration upward from the petiole, and the petiole starting point are individual identification points.

2. The method for establishing according to claim 1, characterized in that, The scanning parameters are: resolution 300~600dpi, brightness 0~30L.

3. The method for establishing according to claim 1, characterized in that, Preprocessing of the first blade profile identification point database further includes: A generalized Protodyakonov analysis is performed on the first blade profile identification point database to maximize the concentration of coordinate points of all blades. The second leaf outline identification point database is stratified according to at least one of the following: origin region, variety, and individual plant; In step (5), when the second leaf outline identification point database is stratified according to variety, average leaf shape data is created.

4. A database for identifying varieties of the genus *Castanopsis*, characterized in that, The database is obtained by the establishment method according to any one of claims 1-3.

5. A method for identifying varieties of the genus *Castanopsis*, characterized in that, include: 1) Collect leaves at physiological maturity from the species of the genus *Castanopsis* to be identified; 2) The abaxial surfaces of all leaves from the chestnut species to be identified were scanned, using uniform scanning parameters and angles. 3) Using image recognition software, homology identification points are selected for all scanned leaves of the chestnut plant varieties to be identified, and the homology identification points are compared with the database for chestnut plant variety identification as described in claim 4. Through canonical variable analysis, the Protodyakonov distance matrix and variety scatter plot are obtained. 4) Based on the Protodyakonov distance matrix and the scatter plot positions of the varieties, determine the species of the *Castanopsis* genus to be identified. In step 1), the *Castanopsis* species to be identified originates from the place of origin of *Castanopsis* species in a database used for the identification of *Castanopsis* species. At least 10 *Castanopsis* species to be identified were collected, with leaves collected from each plant at least from 4 different locations, with an angle difference of 60°-90° between adjacent locations. In step 3), the homology identification points include homology points on the outer edge contour of the leaves of the chestnut plant species to be identified. For each leaf, the homology identification points include 14 primary identification points, 3 secondary identification points, and 7 supplementary identification points. The 14 primary identification points are as follows: The identification points are: the first serration point upward from the petiole, the serration point at the widest part of the leaf blade, the first serration point upward from the widest part of the leaf blade, the first serration point downward from the widest part of the leaf blade, the second serration point downward from the leaf tip, the third serration point downward from the leaf tip, and the fourth serration point downward from the leaf tip. The primary identification points at each location are paired and distributed on both sides of the vein of the leaf. The three secondary identification points are as follows: Identification point at the tip of the leaf, the intersection of the veins at the widest serration, and the junction of the petiole and the leaf blade; The seven supplementary identification points are as follows: The indentation at the upper part of the serration at the widest point of the leaf, the identification point of the first serration downward from the tip of the leaf, the intersection of the main vein immediately adjacent to the center point, the intersection of the veins of the first serration upward from the petiole, and the starting point of the petiole. Among them, the indentation point at the upper part of the serration at the widest point of the leaf and the identification point of the first serration downward from the tip of the leaf are distributed in pairs on both sides of the leaf vein. The main vein intersection point adjacent to the center point, the vein intersection point of the first serration upward from the petiole, and the petiole starting point are individual identification points.

6. The identification method according to claim 5, characterized in that, In step 2), the scanning parameters are: resolution 300~600dpi, brightness 0~30L.

7. The use of the database for identifying varieties of the genus *Castanopsis* as described in claim 4 in the identification of varieties of the genus *Castanopsis*.