A method for identifying a plant specimen

By extracting and fusing multidimensional data information from plant specimen images, and utilizing multimodal features and classification models, the problems of low efficiency and unstable accuracy in traditional methods are solved, achieving efficient and accurate plant specimen identification.

CN120747077BActive Publication Date: 2025-11-25KUNMING INST OF BOTANY CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511225807.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-11-25
Estimated Expiration
2045-08-29

AI Technical Summary

Technical Problem

Traditional plant specimen identification methods rely on human experience, which is inefficient and has unstable accuracy. Single-modal deep learning technology is easily affected by missing information or errors, and its recognition accuracy is difficult to meet the requirements.

Method used

Multidimensional data information is extracted from plant specimen images, feature recognition and encoding are performed, multimodal features of images and text are fused, and a classification model is used to make a comprehensive decision to obtain the recognition result.

Benefits of technology

It significantly improves the automation and accuracy of plant specimen identification, and is suitable for plant taxonomy research, ecological protection, and biodiversity monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747077B_ABST
    Figure CN120747077B_ABST
Patent Text Reader

Abstract

The application discloses a plant specimen identification method, and relates to the technical field of plant identification, which comprises the following steps: extracting multi-dimensional data information in a plant specimen image, performing feature recognition on the extracted multi-dimensional data information, and obtaining corresponding multi-dimensional data features; performing feature coding based on the obtained multi-dimensional data features, obtaining a corresponding feature matrix, performing feature fusion on the obtained feature matrix, and obtaining corresponding multi-modal features; performing feature classification on the obtained multi-modal features by using a classification model, comprehensively deciding a classification result, and obtaining an identification result of the plant specimen; the application effectively fuses data of two different modes of images and texts, overcomes the problem of limited single-mode identification information, and can significantly improve the automation level and accuracy of plant specimen identification.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of plant identification, and specifically relates to a plant specimen identification method. BACKGROUND

[0002] Traditional plant specimen identification methods usually rely on the manual experience and intuitive observation of experts, and have problems of low efficiency, unstable accuracy and high cost. With a substantial increase in the number of specimen collections, manual work alone cannot meet the actual demand. Although single-modal deep learning technology (such as pure image recognition or pure text analysis) has been applied in the field of plant classification, single-modal is susceptible to information loss or errors, and the recognition accuracy is difficult to meet the actual demand. Therefore, the present application provides a plant specimen identification method. SUMMARY

[0003] The present application aims to provide a plant specimen identification method.

[0004] The purpose of the present application can be achieved by the following technical solution: a plant specimen identification method, comprising:

[0005] extracting multi-dimensional data information in a plant specimen image, and performing feature recognition on the extracted multi-dimensional data information to obtain corresponding multi-dimensional data features;

[0006] performing feature coding based on the obtained multi-dimensional data features to obtain a corresponding feature matrix, and performing feature fusion on the obtained feature matrix to obtain a corresponding multi-modal feature;

[0007] performing feature classification on the obtained multi-modal feature using a classification model, and making a comprehensive decision on the classification result to obtain the identification result of the plant specimen.

[0008] Further, the process of extracting multi-dimensional data information in the plant specimen image comprises:

[0009] performing image element labeling in the obtained plant specimen image to obtain a plurality of image element regions;

[0010] constructing a plane coordinate system in the plant specimen image, and mapping the plant specimen image labeled by the image elements into the plane coordinate system;

[0011] generating corresponding local images according to the obtained image element regions, and classifying the local images according to the types of the image elements to obtain corresponding text local images and plant local images;

[0012] obtaining the coordinate ranges of each text local image and plant local image in the plane coordinate system;

[0013] The text partial image and the plant partial image with the coordinate range existing overlap are marked as a parent image, and the image part existing overlap is extracted as a child image, the child image is associated with the corresponding parent image, and the obtained child image is recorded as an interference image;

[0014] The obtained text partial image, plant partial image and interference image are converted to obtain corresponding gray-scale images;

[0015] The obtained text partial image and plant partial image are binarized to obtain corresponding binary images;

[0016] The corresponding text region and plant region are calibrated according to the obtained binary images;

[0017] Each interference image is divided to obtain a corresponding multi-value image;

[0018] The text part and the plant part in the multi-value image are calibrated respectively;

[0019] The parent image associated with the interference image is obtained, and the text region and the plant region in the parent image are compared with the calibrated text part and the plant part in the multi-value image respectively to determine whether there is deviation in the calibration result, and the interference image and the corresponding parent image with deviation are corrected to obtain the final calibration result;

[0020] Text feature extraction and plant feature extraction are performed according to the calibrated binary image to obtain corresponding text data features and plant data features, and the obtained text data features and plant data features are summarized to obtain multi-dimensional data features in the corresponding plant specimen image.

[0021] Further, the process of correcting the deviation of the interference image and the corresponding parent image with deviation includes:

[0022] A two-dimensional plane coordinate system is constructed, and the interference image and the corresponding parent image with deviation are mapped into the two-dimensional plane coordinate system;

[0023] The stamping tool is used to sample in the plant partial image, and then the corresponding text part in the interference image is filled according to the text part in the text partial image;

[0024] The filled interference image and the plant partial image are fused to obtain the final plant partial image;

[0025] Similarly, the stamping tool is used to sample in the text partial image, and then the corresponding plant part in the interference image is filled according to the plant part in the plant partial image;

[0026] Fuse the filled interference image with the text local image to obtain a final text local image.

[0027] Further, the feature encoding process based on the obtained multi-dimensional data features comprises:

[0028] Divide the area corresponding to the obtained plant data features to obtain a plurality of sub-areas;

[0029] Generate a corresponding feature vector for each sub-area by the completed DINOv2, and obtain the correlation between each feature vector by self-attention mechanism on the generated feature vector;

[0030] Integrate the same feature vectors to obtain a corresponding global feature vector, and integrate all sub-areas corresponding to the obtained global feature vector and associate the corresponding morphological features.

[0031] Further, the process of obtaining the feature matrix comprises:

[0032] Input the obtained multi-dimensional data features into the completed CLIP model, and obtain the similarity of the text data features and the plant data features by the CLIP model;

[0033] According to the obtained similarity, determine the correlation between the text data features and the plant data features, and collect the text data features and the plant data features with correlation to obtain a corresponding correlation feature matrix.

[0034] Further, the process of obtaining the corresponding multi-modal features by feature fusion on the obtained feature matrix comprises:

[0035] Obtain the corresponding positive sample data set and negative sample data set from the obtained each correlation feature matrix;

[0036] Fuse the feature of each correlation feature matrix belonging to the same positive sample data set to obtain the corresponding multi-modal feature, and correspond the obtained multi-modal feature to the corresponding plant specimen image.

[0037] Further, the process of obtaining the recognition result of the plant specimen by classifying the obtained multi-modal features by the classification model and comprehensively deciding the classification result comprises:

[0038] Input the obtained multi-modal features into the completed Concat model, MLP model and Conv model respectively;

[0039] Output the classification prediction result of the multi-modal features by the Concat model, MLP model and Conv model respectively;

[0040] The classification prediction results output by the Concat model, the MLP model and the Conv model are respectively summarized to obtain corresponding classification prediction result sets;

[0041] According to the obtained various classification prediction result sets, corresponding comprehensive decisions are set, and according to the set comprehensive decisions, the various classification prediction results in the various classification prediction result sets are finally determined to obtain corresponding final classification results.

[0042] Further, the comprehensive decision is that the prediction results of each associated feature matrix in the multi-modal feature in the Concat model, the MLP model and the Conv model are respectively marked;

[0043] According to the marked prediction results, a weighted voting method is used for voting, and a final prediction result is determined according to the voting result.

[0044] Compared with the prior art, the beneficial effects of the present application are:

[0045] The present application effectively fuses image and text data of two different modalities, overcomes the problem of limited single modal recognition information, can significantly improve the automation level and accuracy of plant specimen identification, is widely applicable to the fields of plant classification research, ecological protection, biodiversity monitoring, etc., has good practical application value and broad application prospect. BRIEF DESCRIPTION OF DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments or the prior art, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments described in the present application, and other drawings can also be obtained by those skilled in the art based on these drawings.

[0047] Figure 1 The schematic diagram of the present application. DETAILED DESCRIPTION

[0048] As shown in Figure 1 A plant specimen identification method, comprising:

[0049] Extracting multi-dimensional data information in the plant specimen image, and performing feature recognition on the extracted multi-dimensional data information to obtain corresponding multi-dimensional data features;

[0050] Based on the obtained multi-dimensional data features, feature coding is performed to obtain corresponding feature matrices, and feature fusion is performed on the obtained feature matrices to obtain corresponding multi-modal features;

[0051] The obtained multi-modal features are classified by using a classification model, and a comprehensive decision is made on the classification results to obtain the identification result of the plant specimen.

[0052] It needs to be further explained that in the specific implementation process, the process of extracting multi-dimensional data information in the plant specimen image includes:

[0053] Image element marking is performed in the obtained plant specimen image to obtain a plurality of image element regions, wherein the image element marking is based on marking the text region and the plant region in the plant specimen image;

[0054] A plane coordinate system is constructed in the plant specimen image, and the plant specimen image subjected to image element marking is mapped into the plane coordinate system;

[0055] According to the obtained image element regions, corresponding local images are generated, and the local images are classified according to the types of image elements to obtain corresponding text local images and plant local images;

[0056] The coordinate ranges of the text local images and the plant local images in the plane coordinate system are obtained;

[0057] The text local images and the plant local images with overlapping coordinate ranges are marked as parent images, and the overlapping image parts are extracted as child images. The child images are associated with the corresponding parent images, and the obtained child images are recorded as interference images;

[0058] The obtained text local images, plant local images and interference images are converted to obtain corresponding gray-scale images;

[0059] The obtained text local images and plant local images are subjected to binarization processing to obtain corresponding binary images;

[0060] According to the obtained binary images, the corresponding text regions and plant regions are respectively calibrated;

[0061] Two pixel intervals are set for each interference image, and the two pixel intervals correspond to the text part and the plant part respectively;

[0062] The obtained interference images are divided according to the set pixel intervals to obtain corresponding multi-value images;

[0063] The text part and the plant part in the multi-value images are respectively calibrated;

[0064] The parent image associated with the interference image is acquired, and the text region and plant region in the parent image are compared with the calibrated text part and plant part in the multi-value image respectively, if the calibrated parts are completely overlapped, it means that the corresponding binary image calibration result is accurate, if the calibrated parts are not completely overlapped, it means that the calibration result has deviation, and the interference image with deviation and the corresponding parent image are corrected for deviation to obtain the final calibration result;

[0065] According to the calibrated binary image, text feature extraction and plant feature extraction are performed to obtain corresponding text data features and plant data features, and the obtained text data features and plant data features are summarized to obtain multi-dimensional data features in the corresponding plant specimen image.

[0066] It should be further explained that, in the specific implementation process, the process of correcting the interference image with deviation and the corresponding parent image for deviation includes:

[0067] A two-dimensional plane coordinate system is constructed, and the interference image with deviation and the corresponding parent image are mapped into the two-dimensional plane coordinate system;

[0068] The stamping tool is used to sample in the plant local image, and then the corresponding text part in the interference image is filled according to the text part in the text local image;

[0069] The filled interference image and the plant local image are fused to obtain the final plant local image;

[0070] Similarly, the stamping tool is used to sample in the text local image, and then the corresponding plant part in the interference image is filled according to the plant part in the plant local image;

[0071] The filled interference image and the text local image are fused to obtain the final text local image.

[0072] It should be further explained that, in the specific implementation process, the process of encoding features based on the obtained multi-dimensional data features to obtain a corresponding feature matrix includes:

[0073] The region corresponding to the obtained plant data features is divided to obtain a plurality of sub-regions;

[0074] Each sub-region generates a corresponding feature vector through the completed DINOv2, and the correlation between the generated feature vectors is obtained through a self-attention mechanism, for example, the positional relationship between the leaves and stems of the plant;

[0075] The same feature vectors are integrated to obtain a corresponding global feature vector, and all sub-regions corresponding to the obtained global feature vector are integrated and associated with corresponding morphological features, for example, if there are several feature vectors that are all leaf blades, then the region obtained by integrating all sub-regions corresponding to the feature vectors of the leaf blades is the morphological feature of the leaf blades.

[0076] The obtained multi-dimensional data features are input into the trained CLIP model, and the similarity between the text data features and the plant data features is obtained by the CLIP model, denoted as S (x, y).

[0077]

[0078] S (x, y) represents the similarity between the i th text data feature and the j th plant data feature, x i represents the i th text data feature, y j represents the j th plant data feature.

[0079] According to the obtained similarity, the relevance between the text data features and the plant data features is determined, and the text data features and the plant data features with relevance are summarized to obtain a corresponding association feature matrix.

[0080] It should be further explained that, in the specific implementation process, the process of feature fusion of the obtained feature matrix to obtain the corresponding multi-modal feature includes:

[0081] The obtained each association feature matrix obtains a corresponding positive sample data set and a negative sample data set; it should be further explained that the positive sample data set represents that when the text data features and the plant data features in two association feature matrices belong to the same plant specimen image, then the two association feature matrices belong to the same positive sample data set, and vice versa.

[0082] The feature fusion of each association feature matrix in the same positive sample data set is performed to obtain a corresponding multi-modal feature, and the obtained multi-modal feature is corresponding to the corresponding plant specimen image.

[0083] It should be further explained that, in the specific implementation process, the process of feature classification of the obtained multi-modal feature by using a classification model and comprehensive decision-making of the classification result to obtain the recognition result of the plant specimen includes:

[0084] The obtained multi-modal features are respectively input into the trained Concat model, MLP model and Conv model.

[0085] ​​​The classification prediction results output by the Concat model, the MLP model and the Conv model are respectively summarized to obtain a corresponding classification prediction result set;

[0086] The classification prediction results output by the Concat model, the MLP model and the Conv model are respectively summarized to obtain a corresponding classification prediction result set;

[0087] According to the obtained classification prediction result set, a corresponding comprehensive decision is set, and according to the set comprehensive decision, each classification prediction result in each classification prediction result set is finally determined to obtain a corresponding final classification result;

[0088] The comprehensive decision is specifically marking the prediction results of each correlation feature matrix in the multi-modal feature in the Concat model, the MLP model and the Conv model;

[0089] According to the marked prediction results, a weighted voting method is used for voting, and the final prediction result is determined according to the voting result; it should be noted that the weight of the weighted voting method depends on the performance of each model, the better the performance, the higher the weight, the probability corresponding to the prediction result of each model is multiplied by the corresponding weight, the prediction results corresponding to all correlation feature matrices are summed, and the probability with the largest total sum is selected as the final prediction category.

[0090] The above is only a preferred embodiment of the present application, and does not limit the present application in any form. Although the present application has been disclosed as above, it is not intended to limit the present application. Any person skilled in the art can make some changes or modifications to the above disclosed technical content without departing from the scope of the present application, and any modification or equivalent replacement of the above embodiments according to the technical essence of the present application still belongs to the scope of the present application.

Claims

1. A method of identifying a plant specimen, characterized by, The method comprises the following steps: extracting multi-dimensional data information in the plant specimen image, and performing feature recognition on the extracted multi-dimensional data information to obtain corresponding multi-dimensional data features; based on the obtained multi-dimensional data features, performing feature coding to obtain a corresponding feature matrix, and performing feature fusion on the obtained feature matrix to obtain a corresponding multi-modal feature; using a classification model to classify the obtained multi-modal feature, and making a comprehensive decision on the classification result to obtain the recognition result of the plant specimen; the process of extracting multi-dimensional data information in the plant specimen image comprises: performing image element labeling in the obtained plant specimen image to obtain a plurality of image element regions; constructing a plane coordinate system in the plant specimen image, and mapping the plant specimen image labeled by the image element into the plane coordinate system; generating corresponding local images according to the obtained image element regions, and classifying the local images according to the types of image elements to obtain corresponding text local images and plant local images; obtaining the coordinate ranges of each text local image and plant local image in the plane coordinate system; labeling the text local images and plant local images with overlapping coordinate ranges as parent images, extracting the overlapping image parts as sub-images, associating the sub-images with the corresponding parent images, and recording the obtained sub-images as interference images; transforming the obtained text local images, plant local images and interference images to obtain corresponding gray-scale images; performing binaryzation processing on the obtained text local images and plant local images to obtain corresponding binaryzation images; labeling the corresponding text regions and plant regions according to the obtained binaryzation images; dividing each interference image to obtain a corresponding multi-value image; labeling the text part and plant part in the multi-value image respectively; obtaining the parent image associated with the interference image, and comparing the text region and plant region in the parent image with the labeled text part and plant part in the multi-value image respectively to determine whether there is deviation in the labeling result, and correcting the deviation of the interference image and the corresponding parent image to obtain the final labeling result; performing text feature extraction and plant feature extraction according to the correctly labeled binaryzation images to obtain corresponding text data features and plant data features, and summarizing the obtained text data features and plant data features to obtain the multi-dimensional data features in the corresponding plant specimen image.

2. The method of claim 1, wherein the plant specimen is identified by the steps of: The process of correcting the deviation of the interference image and the corresponding parent image comprises: constructing a two-dimensional plane coordinate system, and mapping the interference image with deviation and the corresponding parent image into the two-dimensional plane coordinate system; using a stamping tool to sample in the plant local image, and then filling the corresponding text part in the interference image according to the text part in the text local image; fusing the filled interference image and the plant local image to obtain the final plant local image; Similarly, using a stamping tool to sample in the text local image, and then filling the corresponding plant part in the interference image according to the plant part in the plant local image; The filled interference image is fused with the text local image to obtain a final text local image.

3. The method of claim 2, wherein the plant specimen is identified by the steps of: The process of feature coding based on the obtained multi-dimensional data features includes: The area corresponding to the obtained plant data features is divided to obtain a plurality of sub-areas; Each sub-area generates a corresponding feature vector through the trained DINOv2, and the correlation between the generated feature vectors is obtained through a self-attention mechanism. The same feature vectors are integrated to obtain a corresponding global feature vector, and all sub-areas corresponding to the obtained global feature vector are integrated and associated with the corresponding morphological features.

4. The method of claim 3, wherein the plant specimen is identified by the steps of: The process of obtaining the feature matrix includes: The obtained multi-dimensional data features are input into the trained CLIP model to obtain the similarity of the text data features and the plant data features through the CLIP model; According to the obtained similarity, the correlation between the text data features and the plant data features is determined, and the text data features and the plant data features with correlation are summarized to obtain a corresponding correlation feature matrix.

5. The method of claim 4, wherein the plant specimen is identified by the steps of: The process of obtaining the corresponding multi-modal features by feature fusion of the obtained feature matrix includes: Obtain the corresponding positive sample data set and negative sample data set from the obtained correlation feature matrix; The feature fusion of the correlation feature matrix in the same positive sample data set is obtained to obtain the corresponding multi-modal features, and the obtained multi-modal features are corresponding to the plant specimen image.

6. The method of claim 5, wherein the plant specimen is identified by the steps of: The process of feature classification of the obtained multi-modal features by using the classification model and the comprehensive decision of the classification result to obtain the recognition result of the plant specimen includes: The obtained multi-modal features are input into the trained Concat model, MLP model and Conv model respectively; The classification prediction results of the multi-modal features are output by the Concat model, MLP model and Conv model respectively; The classification prediction results output by the Concat model, MLP model and Conv model are summarized respectively to obtain a corresponding classification prediction result set; According to the obtained classification prediction result set, a corresponding comprehensive decision is set, and according to the set comprehensive decision, the classification prediction results in each classification prediction result set are finally determined to obtain a corresponding final classification result.

7. The method of claim 6, wherein the plant specimen is identified by the steps of: The comprehensive decision is to mark the prediction results of each correlation feature matrix in the multi-modal features in the Concat model, MLP model and Conv model respectively; According to the marked prediction results, the weighted voting method is used for voting, and the final prediction result is determined according to the voting result.

Citation Information

Patent Citations

  • Plant specimen classification system and method fusing multi-scale direction texture features

    CN119478842A

  • Crop intelligent identification system based on CLIP multi-mode method

    CN119580049A