Method and System for Generating Flora and Species Identification Keys Based on Deep Learning

Automatically generate and update the plant archaeology and species search tables through deep learning methods, solving the problem of time-consuming and laborious and human errors in the compilation of plant archaeology and improving the efficiency of plant species recognition.

CN120196791BActive Publication Date: 2025-07-29KUNMING INST OF BOTANY CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510686492.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-07-29
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

The writing and updating of existing plant journals relies on a lot of manual work, which is time-consuming and labor-intensive, and is prone to human errors, making it difficult to improve the efficiency of plant species identification.

Method used

Deep learning methods are adopted to collect and label plant information, build sample feature sets, train deep learning models, generate plant archetype contents and construct species search tables, and realize automated identification and updates.

Benefits of technology

It realizes automatic generation and real-time update of plant archaeology and species retrieval tables, improves plant species recognition efficiency and reduces human errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196791B_ABST
    Figure CN120196791B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for generating a flora and a species identification key based on deep learning, which relates to the technical field of atlas construction. Sample data is collected, and feature extraction is performed on the sample data to obtain a corresponding sample feature set. The deep learning model is trained according to the obtained sample feature set, and flora content is generated according to the training result, and a plant information database is constructed. A species identification key is constructed to identify the plants to be identified through the species identification key. The flora content and the species identification key are automatically generated, and the model can continuously learn new plant data to update the flora content and the identification key in real time. With the discovery of new species and the update of sample data, the model will automatically adapt to new features, update the identification key and expand the content of the flora. This function ensures that the flora and the identification key can follow the progress of scientific research, enabling botanists to significantly improve the efficiency of plant species identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of atlas construction, and specifically to a method and system for generating a flora and species checklist based on deep learning. Background Art

[0002] As an important tool for botanical research, the flora has long been used to record and describe the morphological characteristics, distribution, ecological habits, and other relevant information of various plants. The flora not only provides systematic data for plant taxonomy but also provides a large amount of basic research data for botanists, environmentalists, ecologists, etc. However, the current compilation and update of the flora still rely on a large amount of manual work, including the collection, classification, morphological description of plant specimens, and the compilation of plant checklists. This process is not only time-consuming and laborious but also easily affected by human factors, resulting in errors and inconsistent results. Therefore, how to improve the efficiency of flora compilation and species identification and reduce human errors is a problem that needs to be solved. For this reason, a method and system for generating a flora and species checklist based on deep learning are provided. Summary of the Invention

[0003] The purpose of the present invention is to provide a method and system for generating a flora and species checklist based on deep learning.

[0004] The purpose of the present invention can be achieved through the following technical solutions: A method for generating a flora and species checklist based on deep learning, comprising the following steps:

[0005] Step S1: Collect sample data, extract features from the sample data, and obtain the corresponding sample feature set;

[0006] Step S2: Train a deep learning model according to the obtained sample feature set, generate flora content according to the training results, and build a plant information database;

[0007] Step S3: Build a species checklist and identify the plants to be identified through the species checklist.

[0008] Further, the process of collecting sample data, extracting features from the sample data, and obtaining the corresponding sample feature set includes:

[0009] Collect plant information of different plants based on existing public materials; wherein, the existing public materials include public databases, scientific research institutions, floras, and other literatures;

[0010] The plant information includes plant image information and plant text information, and the plant text information includes morphological feature description information, ecological habit information, and growth distribution information;

[0011] Summarize the plant information of each plant separately to obtain a set of plant information corresponding to each plant;

[0012] Annotate the plant image information and plant text information in each obtained set of plant information, and summarize the annotated parts to obtain a corresponding set of sample features.

[0013] Furthermore, the process of training a deep learning model based on the obtained set of sample features includes:

[0014] Construct an image sample training set and a text sample training set respectively according to the obtained set of sample features;

[0015] Obtain several batches of data pairs according to the image samples in the image sample training set and the text samples in the text sample training set;

[0016] Input the obtained data pairs into the deep learning model batch by batch;

[0017] The deep learning model extracts the image feature vectors and text feature vectors corresponding to the image samples and text samples in the data pair respectively, as well as the predicted correlation coefficients between the image feature vectors and text feature vectors, and then obtains the corresponding contrast loss function;

[0018] Iteratively train the deep learning model until the contrast loss function of the deep learning is lower than the preset value or the number of iterations reaches the preset number, thus completing the training of the deep learning model;

[0019] Construct a validation set and a test set, input the constructed validation set and test set into the trained deep learning model, evaluate the performance of the deep learning model according to the output of the deep learning model, and after passing the evaluation, output the corresponding recognition features and generate the corresponding flora content.

[0020] Furthermore, the specific process of generating the flora content is as follows:

[0021] Generate corresponding plant labels with the names of each plant, associate the corresponding recognition features with the plant labels, generate the flora content corresponding to the plant labels according to the plant labels and the associated recognition features, and output the flora content according to the natural language text;

[0022] Associate the output flora content with the corresponding plant image information and generate the corresponding plant information entries;

[0023] Summarize the plant information entries of all plants to complete the construction of the plant information database.

[0024] Furthermore, the process of constructing a species key includes:

[0025] For each plant information entry, obtain all the characteristics corresponding to the plant information entry, and create corresponding retrieval nodes respectively according to each of the obtained characteristics;

[0026] Classify each retrieval node into a judgment node and a selection node according to the type of the characteristic;

[0027] Connect each retrieval node to obtain a corresponding retrieval path;

[0028] Summarize the retrieval paths corresponding to all plant information entries, match the retrieval nodes on each retrieval path, merge the retrieval nodes of the same type, and denote the merged retrieval nodes as set nodes;

[0029] Summarize the plants that meet each set node to obtain a corresponding plant species set;

[0030] Obtain the number of retrieval nodes merged by each set node, set the retrieval priority of the set node according to the number of merged retrieval nodes, and juxtapose the set nodes with the same priority, thereby completing the construction of the species retrieval table.

[0031] Further, the process of identifying a plant to be identified through the species retrieval table includes:

[0032] The user retrieves the plant to be identified according to the constructed species retrieval table, traverses each set node in the species retrieval table in turn, and determines the retrieval path of the plant to be identified;

[0033] Mark the plant species sets corresponding to all set nodes on the retrieval path, and obtain the intersection of the plants within all set nodes;

[0034] Then the plants within the intersection are the plants to be identified, and retrieve the plant flora content corresponding to the plant, thereby completing the user's retrieval of the plant to be identified.

[0035] Further, a plant flora and species retrieval table generation system includes a management center, and the management center is communicatively connected to a data input module, a deep learning module, and a user retrieval module;

[0036] The data input module is used to input the plant information of different plants in existing public materials, and use the input plant information as sample data;

[0037] The deep learning module is used to perform model training according to the sample data, generate plant flora content according to the training result, and construct a plant information database;

[0038] The user retrieval module is used to construct a corresponding species retrieval table according to the plant information library, and the user retrieves the plants to be identified and the corresponding flora content according to the constructed species retrieval table.

[0039] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0040] The present invention can not only automatically generate flora content and species retrieval tables, but also continuously learn new plant data through the model to update the flora content and retrieval tables in real time. With the discovery of new species and the update of sample data, the model will automatically adapt to new features, update the retrieval table and expand the content of the flora. This function ensures that the flora and retrieval table can follow the progress of scientific research, enabling botanists to significantly improve the efficiency of plant species identification. Description of the Drawings

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0042] Figure 1 It is the schematic diagram of the present invention. Detailed Embodiments

[0043] As Figure 1 shown, the method for generating a flora and a species retrieval table based on deep learning includes the following steps:

[0044] Step S1: Collect sample data, extract features from the sample data, and obtain the corresponding sample feature set;

[0045] Step S2: Train the deep learning model according to the obtained sample feature set, generate flora content according to the training results, and construct a plant information library;

[0046] Step S3: Construct a species retrieval table and identify the plants to be identified through the species retrieval table.

[0047] It should be further noted that in the specific implementation process, the process of collecting sample data, extracting features from the sample data, and obtaining the corresponding sample feature set includes:

[0048] Based on existing public materials, collect plant information of different plants; among them, the existing public materials include public databases, scientific research institutions, floras, and other documents;

[0049] The plant information includes plant image information and plant text information, and the plant text information includes morphological feature description information, ecological habit information, and growth distribution information;

[0050] The plant image information includes pictures of each plant from different angles and at different growth stages, such as pictures of leaves, flowers, and fruits;

[0051] The morphological feature description information includes morphological descriptions of each plant at different growth stages, such as leaf shape, flower color, shape, fruit type, etc.

[0052] The ecological habit information includes the growth environment of the plant, such as climate, soil, altitude;

[0053] The growth distribution information includes global or regional distribution;

[0054] Summarize the plant information of each plant respectively to obtain a plant information set corresponding to each plant;

[0055] Label the plant image information and plant text information in each obtained plant information set, and summarize the labeled parts to obtain a corresponding sample feature set.

[0056] It should be further noted that in the specific implementation process, the process of training the deep learning model according to the obtained sample feature set includes:

[0057] Construct an image sample training set and a text sample training set respectively according to the obtained sample feature set;

[0058] Obtain several batches of data pairs according to the image samples in the image sample training set and the text samples in the text sample training set;

[0059] Input the obtained data pairs into the deep learning model batch by batch;

[0060] The deep learning model extracts the image feature vectors and text feature vectors corresponding to the image samples and text samples in the data pair respectively, as well as the predicted correlation coefficients between the image feature vectors and the text feature vectors;

[0061] Denote the obtained image feature vectors as i, i = 1, 2,..., n;

[0062] Denote the obtained text feature vectors as j, j = 1, 2,..., n;

[0063] Then the corresponding contrast loss function is obtained as:

[0064] ;

[0065] Among them, is a boundary value, and > 0, is the predicted correlation coefficient between the image feature vector labeled i and the text feature vector labeled j;

[0066] Iteratively train the deep learning model until the contrast loss function of the deep learning is lower than the preset value or the number of iterations reaches the preset number, thereby completing the training of the deep learning model;

[0067] Construct a validation set and a test set, input the constructed validation set and test set into the trained deep learning model, evaluate the performance of the deep learning model according to the output of the deep learning model, and after passing the evaluation, output the corresponding recognition features and generate the corresponding flora content.

[0068] It should be further noted that in the specific implementation process, the specific process of generating the flora content is as follows:

[0069] Generate corresponding plant labels for the names of each plant, associate the corresponding recognition features with the plant labels, generate flora content corresponding to the plant labels according to the plant labels and the associated recognition features, and output the flora content according to natural language text;

[0070] Illustrate with examples:

[0071] Set the plant label as A, and the associated recognition features are respectively:

[0072] Morphological features: elliptical leaves, white flowers, five-petaled flowers, and the fruit is a round berry;

[0073] Ecological features: moist soil environment, warm climate;

[0074] Distribution features: southeastern Asia, Southeast Asian region, forest edge;

[0075] Then the corresponding flora content generated is: "This plant has long elliptical leaves, white five-petaled flowers, and the fruit is a round berry. It usually grows at the forest edge in a warm climate, adapts to a moist soil environment, and is distributed in southeastern Asia, especially common in southern China and the Southeast Asian region;

[0076] Associate the output flora content with the corresponding plant image information and generate corresponding plant information entries;

[0077] Summarize the plant information entries of all plants to complete the construction of the plant information database.

[0078] It should be further noted that in the specific implementation process, the process of constructing a species key and identifying the plants to be identified through the species key includes:

[0079] For each plant information entry, obtain all the characteristics corresponding to the plant information entry, and establish corresponding retrieval nodes according to each of the obtained characteristics;

[0080] Classify each retrieval node into a judgment node and a selection node according to the type of the characteristic; it should be further noted that in the specific implementation process, the judgment node is specifically a "yes" or "no" judgment, such as whether the plant has fruits; the selection node is specifically an "either-or selection", such as the number of petals of the plant being "1", "2", "3",..., and select the selection node that conforms to the number of petals of the plant;

[0081] Connect each retrieval node to obtain a corresponding retrieval path;

[0082] Summarize the retrieval paths corresponding to all plant information entries, match the retrieval nodes on each retrieval path, merge the retrieval nodes of the same type, and record the merged retrieval nodes as set nodes, where the set nodes include at least one retrieval node;

[0083] Summarize the plants that meet each set node to obtain a corresponding plant species set;

[0084] Obtain the number of retrieval nodes merged by each set node, set the retrieval priority of the set node according to the number of merged retrieval nodes, where the more the number of merged retrieval nodes, the higher the retrieval priority of the set node, and juxtapose the set nodes with the same priority, thereby completing the construction of the species retrieval table;

[0085] The user retrieves the plant to be identified according to the constructed species retrieval table, traverses each set node in the species retrieval table in turn, and determines the retrieval path of the plant to be identified;

[0086] Mark the plant species sets corresponding to all set nodes on the retrieval path, and obtain the intersection of the plants in all set nodes;

[0087] Then the plants in the intersection are the plants to be identified, and retrieve the plant flora content corresponding to the plant, thereby completing the user's retrieval of the plant to be identified.

[0088] In another embodiment of the present invention, a plant flora and species retrieval table generation system is also disclosed, including a management center, and the management center is communicatively connected to a data input module, a deep learning module, and a user retrieval module;

[0089] The data input module is used to input the plant information of different plants in the existing public materials, and use the input plant information as sample data;

[0090] The deep learning module is used to perform model training based on sample data, generate flora content according to the training results, and construct a plant information database;

[0091] The user retrieval module is used to construct a corresponding species retrieval table according to the plant information database, and the user retrieves the plant to be identified and the corresponding flora content according to the constructed species retrieval table.

[0092] The above are only preferred embodiments of the present invention, and do not impose any form of limitation on the present invention. Although the present invention has been disclosed above with preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to equivalent embodiments by using the above-disclosed technical content within the scope of the technical solution of the present invention. However, as long as it does not depart from the content of the technical solution of the present invention, any modification or equivalent replacement made to the above embodiments based on the technical essence of the present invention still falls within the scope of the technical solution of the present invention.

Claims

1. A method for generating a flora and species checklist based on deep learning, characterized in that, It includes the following steps: Step S1: Collect sample data, extract features from the sample data, and obtain the corresponding sample feature set; Step S2: Train the deep learning model according to the obtained sample feature set, generate flora content based on the training results, and construct a plant information database; Step S3: Construct a species key, and identify the plants to be identified through the species key; Construct a validation set and a test set, input the constructed validation set and test set into the trained deep learning model, evaluate the performance of the deep learning model according to the output of the deep learning model, and after passing the evaluation, output the corresponding identification features and generate the corresponding flora content; The process of generating flora content is as follows: Generate corresponding plant labels with the names of each plant, associate the corresponding identification features with the plant labels, generate flora content corresponding to the plant labels according to the plant labels and the associated identification features, and output the flora content in natural language text; Associate the output flora content with the corresponding plant image information, and generate corresponding plant information entries; Summarize the plant information entries of all plants to complete the construction of the plant information database; The process of constructing a species key includes: According to each plant information entry, obtain all the features corresponding to the plant information entry, and establish corresponding retrieval nodes respectively according to the obtained features; Classify each retrieval node into a judgment node and a selection node according to the type of the feature; Connect each retrieval node to obtain the corresponding retrieval path; Summarize the retrieval paths corresponding to all plant information entries, match the retrieval nodes on each retrieval path, merge the retrieval nodes of the same type, and record the merged retrieval nodes as set nodes; Summarize the plants that meet each set node to obtain the corresponding plant species set; Obtain the number of retrieval nodes merged by each set node, set the retrieval priority of the set node according to the number of merged retrieval nodes, and juxtapose the set nodes with the same priority to complete the construction of the species key.

2. The method for generating a flora and a species checklist based on deep learning according to claim 1, wherein The process of collecting sample data, extracting features from the sample data, and obtaining the corresponding sample feature set includes: Collect plant information of different plants based on existing public materials; among them, the existing public materials include public databases, scientific research institutions, floras, and other literatures; The plant information includes plant image information and plant text information, and the plant text information includes morphological feature description information, ecological habit information, and growth distribution information; Summarize the plant information of each plant respectively to obtain a plant information set corresponding to each plant; Annotate the plant image information and plant text information in each obtained plant information set, and summarize the annotated parts to obtain the corresponding sample feature set.

3. The method for generating a flora and species checklist based on deep learning according to claim 2, wherein The process of training the deep learning model according to the obtained sample feature set includes: According to the obtained sample feature set, construct an image sample training set and a text sample training set respectively; Obtain a number of batches of data pairs based on the image samples in the image sample training set and the text samples in the text sample training set; Input the obtained data pairs into the deep learning model batch by batch; The deep learning model respectively extracts the image feature vector and the text feature vector corresponding to the image sample and the text sample in the data pair, as well as the predicted correlation coefficient between the image feature vector and the text feature vector, and then obtains the corresponding contrast loss function; Perform iterative training on the deep learning model until the contrast loss function of the deep learning is lower than the preset value or the number of iterations reaches the preset number, thereby completing the training of the deep learning model.

4. The method for generating a flora and a species checklist based on deep learning according to claim 3, wherein The process of identifying the plant to be identified through the species checklist includes: The user retrieves the plant to be identified according to the constructed species checklist, traverses each set node in the species checklist in turn, and determines the retrieval path of the plant to be identified; Mark the plant species sets corresponding to all the set nodes on the retrieval path, and obtain the intersection of the plants in all the set nodes; Then the plant in the intersection is the plant to be identified, and the content of the flora corresponding to the plant is retrieved, thereby completing the user's retrieval of the plant to be identified.

5. A flora and species checklist generation system applied to the method for generating a flora and species checklist based on deep learning according to any one of claims 1 to 4, including a management center, characterized in that, The management center is communicatively connected to a data input module, a deep learning module, and a user retrieval module; The data input module is used to input the plant information of different plants in the existing public materials, and use the input plant information as sample data; The deep learning module is used to perform model training according to the sample data, generate the content of the flora according to the training results, and construct a plant information database; The user retrieval module is used to construct a corresponding species checklist according to the plant information database, and the user retrieves the plant to be identified and the corresponding content of the flora according to the constructed species checklist.

Citation Information

Patent Citations

  • Desert vegetation recovery species configuration method based on community construction mechanism

    CN116433447A

  • Unified multi-agent system for abnormality detection and isolation

    US20220327204A1