Plant identification method and device and electronic equipment
By performing multimodal feature analysis on plant images, more accurate and comprehensive plant recognition results are generated, and the problems of low plant recognition accuracy and insufficient information in the prior art are solved.
Patent Information
- Application Number
- CN202510240073.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-17
AI Technical Summary
The existing plant classification methods have low plant identification accuracy and cannot provide comprehensive general plant information.
By analyzing the multimodal feature information of plants in plant images, including color, texture, attributes, geographical distribution, functional value and morphological description information for different growth periods, more accurate and comprehensive plant identification results are generated.
It improves the accuracy of plant recognition and provides more comprehensive and detailed plant information, solving the problems of inaccurate identification results and insufficient information in the prior art.
Smart Images

Figure CN120164104A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of artificial intelligence, and particularly relates to a plant recognition method, device and electronic device. Background Art
[0002] Plant classification is a technology that requires precise recognition of plants in images, and it has practical significance in aspects such as plant classification and recognition, protection and utilization of plant resources, exploration of the genetic relationships between plants, elucidation of the evolutionary laws of plants, and agricultural applications.
[0003] Currently, the main method of plant classification is to manually extract the color, texture features, etc. of plants, and then combine shallow machine learning algorithms, such as support vector machine algorithms, to classify plants. This method has a relatively low classification accuracy for plants. Summary of the Invention
[0004] The purpose of the embodiments of this application is to provide a plant recognition method, device and electronic device, which can improve the accuracy of plant recognition and provide comprehensive plant recognition information.
[0005] In a first aspect, the embodiments of this application provide a plant recognition method, which includes:
[0006] Recognize the plants in the plant image to obtain plant recognition result information; wherein, the plant recognition result information is generated according to the plant multi-modal feature information of the plants;
[0007] Display the plant recognition result information; wherein, the plant recognition result information includes at least one of the following: the attribute information of the plant, the geographical distribution information of the plant, the functional value information of the plant, the morphological description information and images of different growth periods of the plant.
[0008] In a second aspect, the embodiments of this application provide a plant recognition device, which includes:
[0009] A recognition module, configured to recognize the plants in the plant image to obtain plant recognition result information; wherein, the plant recognition result information is generated according to the plant multi-modal feature information of the plants;
[0010] A display module, configured to display the plant recognition result information; wherein, the plant recognition result information includes at least one of the following: the attribute information of the plant, the geographical distribution information of the plant, the functional value information of the plant, the morphological description information and images of different growth periods of the plant.
[0011] In a third aspect, an embodiment of the present application provides a readable storage medium, on which a program or instructions are stored, and when the program or instructions are executed by a processor, the steps of the method described in the first aspect are implemented.
[0012] In a fourth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run a program or instructions to implement the method described in the first aspect.
[0013] In a fifth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the method described in the first aspect.
[0014] In the embodiment of the present application, when identifying a plant, multimodal feature information of the plant in the plant image is referred to. That is, in the embodiment of the present application, when identifying a plant, in addition to referring to the color and texture features of the plant in the plant image, other feature information of the plant is also referred to, thereby improving the accuracy of plant identification. In addition, after identifying the plant, the obtained plant identification result information includes at least one of the plant's attribute information, geographical distribution information, functional value information, morphological description information of different growth periods of the plant, and the image, that is, the obtained plant identification result information is more comprehensive. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is a schematic flowchart of a plant identification method provided by some embodiments of the present application;
[0016] Figure 2 is a schematic diagram showing the display of plant identification result information of Lonicera japonica provided by some embodiments of the present application;
[0017] Figure 3 is a schematic diagram showing the display of plant identification result information of Lilium brownii provided by some embodiments of the present application;
[0018] Figure 4 is a schematic diagram of an artificial intelligence assistant interface provided by some embodiments of the present application;
[0019] Figure 5 is a schematic diagram of an artificial intelligence assistant interface provided by some embodiments of the present application;
[0020] Figure 6 is a schematic diagram of an artificial intelligence assistant interface provided by some embodiments of the present application;
[0021] Figure 7 is a schematic diagram of a knowledge graph of Lonicera japonica provided by some embodiments of the present application;
[0022] Figure 8 It is a schematic diagram showing the morphological description information of Lonicera provided by some embodiments of the present application;
[0023] Figure 9 It is a schematic diagram showing the attribute information of Lonicera provided by some embodiments of the present application;
[0024] Figure 10 It is a schematic diagram showing the functional value information of Lonicera provided by some embodiments of the present application;
[0025] Figure 11 It is a schematic diagram showing the geographical distribution information of Lonicera provided by some embodiments of the present application;
[0026] Figure 12 It is a schematic diagram of the images of other plants associated with the plant in the plant image provided by some embodiments of the present application;
[0027] Figure 13 It is a schematic diagram of the structure of a plant recognition device shown by some embodiments of the present application;
[0028] Figure 14 It is a schematic diagram of the structure of an electronic device shown by some embodiments of the present application;
[0029] Figure 15 It is a schematic diagram of the hardware structure of an electronic device shown by some embodiments of the present application. Detailed implementation manners
[0030] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, rather than all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.
[0031] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such used data may be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same type, and the number of objects is not limited. For example, the first object may be one or N. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / ", generally represents an "or" relationship between the associated objects before and after.
[0032] The terms used in the implementation manner part of the present application are only used to explain the specific embodiments of the present application, rather than to limit the present application.
[0033] The terms related to the embodiments of the present invention are explained below.
[0034] Multimodal information: Information of multiple modalities, and multiple modalities include but are not limited to: multiple modalities such as images, texts, videos, audios, etc.
[0035] Artificial Intelligence (AI): A new technical science that studies, develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence.
[0036] Large Language Model (LLM), abbreviated as large model, is an advanced artificial intelligence algorithm trained on a large amount of data.
[0037] Graph Neural Networks (GNN) model: A deep learning model specifically used to process graph data. Compared with traditional deep learning models, GNN can better capture the relationships and topological structures between nodes in a graph. It obtains the representation of nodes by iteratively updating the feature vectors of nodes.
[0038] Knowledge graph: A structured semantic knowledge base used to describe concepts in the physical world and their interrelationships in symbolic form. Its basic unit of composition is the "entity-relationship-entity" triple, as well as entity and its related attribute-value pairs. Entities are connected to each other through relationships, forming a network-like knowledge structure.
[0039] Confidence level: Also known as reliability, or confidence level, confidence coefficient, which refers to the probability that the population parameter value falls within a certain area of the sample statistic value.
[0040] The technical solutions of the embodiments of the present application can be applied to identifying plants in plant images, and then obtaining plant attribute information, plant geographical distribution information, morphological description information of different growth periods of plants, and the scene of the image. For example, when a user is playing in a park and takes a photo of honeysuckle casually, there is only honeysuckle in the photo, but the user doesn't know what kind of plant it is and wants to know the name of the plant and other information about it, such as the family and genus to which the plant belongs, as well as its geographical distribution information, morphological description information of different growth periods of the plant, and the image, etc. Another example is that when a user sees a photo on the web page, there are lilies and roses in the photo, but the user doesn't know lilies and only knows there are roses in the photo. The user wants to know the name of the lily and other information about it, such as the family and genus to which the lily belongs, as well as its geographical distribution information, morphological description information of different growth periods of the lily, and the image, etc.
[0041] The following will combine the accompanying drawings and, through specific embodiments and their application scenarios, elaborate in detail on the plant recognition method provided by the embodiments of the present application.
[0042] Figure 1 FIG. 1 is a schematic flowchart of a plant recognition method provided by an embodiment of the present application. The execution subject of this plant recognition method can be an electronic device, which can be, but is not limited to, a personal computer (PC), a smart phone, a tablet computer, or a personal digital assistant (PDA), etc.
[0043] As Figure 1 shown, the plant recognition method provided by the embodiment of the present application may include step 110 - step 120.
[0044] Step 110: Recognize the plants in the plant image to obtain plant recognition result information.
[0045] Among them, the plant image can be an image containing plants. For example, it can be an image containing only plants, or it can also be an image with other objects in addition to plants. For example, when a user is playing in the park and takes a photo of himself with a flower casually, the other objects contained in the specific plant image are not limited in addition to plants.
[0046] It should be noted that the types of plants contained in the above plant images are not limited. For example, they can be flowers, trees, or other types of plants.
[0047] The plant recognition result information can be the recognition result information of the plants obtained after recognizing the plants in the plant image.
[0048] The above plant recognition result information may include at least one of the following: the attribute information of the plant, the geographical distribution information of the plant, the functional value information of the plant, the morphological description information and images of different growth periods of the plant.
[0049] The above attribute information of the plant can be the family information, genus information, and species information of the plant. For example, for the plant Lonicera japonica, its family information is: Caprifoliaceae, its genus information is: Lonicera, and its species information is: Lonicera japonica. Another example is for the plant Lilium brownii var. viridulum, its family information is Liliaceae, its genus information is: Lilium, and its species information is: Lilium brownii var. viridulum.
[0050] The morphological description information of a plant at different growth stages can be information used to describe the external morphological characteristics of the plant at different growth stages. For example, for Lonicera japonica at the mature stage, its morphological description information can be: the branches are grayish-brown, with hard rough hairs or bristles, the hair color is brownish-red, and the leaves are disc-shaped. For Lilium brownii var. viridulum at the mature stage, its morphological description information can be that the petals are light white, with a light brown line in the middle, the leaves are alternate, without petioles, and the shape is lanceolate or ovate-lanceolate.
[0051] The above plant recognition result information is generated based on the plant's multi-modal feature information. Here, the plant's multi-modal feature information can be the plant's multi-modal feature information, such as but not limited to including: the image of the plant, the morphological description information of the plant, the geographical distribution information of the plant, and the functional value information of the plant.
[0052] In one example, a user is playing in the park and casually takes a photo of Lonicera japonica. There is only Lonicera japonica in this photo, but the user doesn't know what kind of plant it is. The user wants to know the name of the plant, as well as other information about the plant, such as the family and genus to which the plant belongs, the geographical distribution information of the plant, the morphological description information of the plant at different growth stages, and the image, etc. Then, the Lonicera japonica in the photo of Lonicera japonica can be recognized to obtain the attribute information of Lonicera japonica: Lonicera japonica, belonging to the genus Lonicera, family Caprifoliaceae, the geographical distribution information of Lonicera japonica: widely produced in East Asia, the functional value information of Lonicera japonica: cold in nature, sweet in taste, slightly sour in taste, as well as the morphological description information and image of Lonicera japonica at different growth stages.
[0053] In another example, a user sees a photo on the web page. There are Lilium brownii var. viridulum and Rosa rugosa in this photo, but the user doesn't know Lilium brownii var. viridulum and only knows that there is Rosa rugosa in the photo. The user wants to know the name of Lilium brownii var. viridulum, as well as other information about the flower, such as the family and genus to which the Lilium brownii var. viridulum belongs, the geographical distribution information of the Lilium brownii var. viridulum, the morphological description information of the Lilium brownii var. viridulum at different growth stages, and the image, etc. Then, the Lilium brownii var. viridulum in the photo can be recognized to obtain the attribute information of Lilium brownii var. viridulum: Lilium brownii var. viridulum, belonging to the family Liliaceae, genus Lilium, the geographical distribution information of Lilium brownii var. viridulum: mainly distributed in the temperate regions of the Northern Hemisphere such as eastern Asia, Europe, and North America, the functional value information of Lilium brownii var. viridulum: the bulbs of Lilium brownii var. viridulum are rich in starch, edible, and also used medicinally, as well as the morphological description information and image of Lilium brownii var. viridulum at different growth stages.
[0054] Step 120: Display the plant recognition result information.
[0055] Continuing to refer to the first example above, after recognizing the photo of Lonicera japonica, the plant recognition result information of Lonicera japonica is obtained, and the plant recognition result information of Lonicera japonica is displayed, that is, as Figure 2As shown, the property information of Lonicera in the display image 20: Lonicera, belonging to the genus Lonicera, Caprifoliaceae, geographical distribution information: widely produced in East Asia, functional value: cold in nature, sweet in taste, slightly sour, morphological description information and images at different growth stages (not shown in Figure 2 ).
[0056] Continuing to refer to the second example above, after identifying Lilium, the plant identification result information of Lilium can be obtained, and the plant identification result information of Lilium is displayed, that is, as Figure 3 shown, the property information of Lilium in the display image 30: Lilium, belonging to the family Liliaceae, genus Lilium, geographical distribution information: mainly distributed in the temperate regions of the Northern Hemisphere such as eastern Asia, Europe, and North America, functional value: the bulbs of Lilium are rich in starch, edible, and also used medicinally, morphological description information and images at different growth stages (not shown in Figure 3 ).
[0057] It should be noted that only part of the plant identification result information of Lonicera is shown in Figure 2 , and similarly, only part of the plant identification result information of Lilium is shown in Figure 3 . If you want to view more specific information, you can perform corresponding operations to view more plant identification result information. As Figure 2 shown, if you want to view more functional value information of Lonicera, you can click on the "Functional Value" control 21 in Figure 2 to view more functional value information of Lonicera. If you want to view more geographical distribution information of Lonicera, you can click on the "Geographical Distribution" control 22 in Figure 2 to view more geographical distribution information of Lonicera.
[0058] In some embodiments of the present application, in order to meet the time requirements of users for plant identification, before step 110, the above-mentioned method may further include:
[0059] Receiving a plant image uploaded by the user;
[0060] Receiving plant identification instruction information input by the user on the artificial intelligence assistant interface;
[0061] Step 110 may specifically include:
[0062] When receiving the plant identification instruction information input by the user on the artificial intelligence assistant interface, identifying the plant in the plant image to obtain plant identification result information.
[0063] Among them, the artificial intelligence assistant interface may be an operating system management interface provided based on AI technology. The operating system here may specifically be a Linux system.
[0064] The plant recognition indication information can be information used to indicate the recognition of the plant in the plant image, such as "Please recognize the plant in this plant image".
[0065] In some embodiments of the present application, the user can upload a plant image containing the plant to be recognized in the AI assistant interface, and then input the plant recognition indication information for recognizing the plant in the uploaded plant image in the AI assistant interface. When the electronic device receives the plant recognition indication information input by the user in the artificial intelligence assistant interface, it can recognize the plant in the plant image based on the plant recognition indication information to obtain the plant recognition result information.
[0066] It should be noted that the plant image uploaded by the user can be taken by the user himself and then uploaded to the AI assistant interface, or downloaded by the user from the web page and then uploaded to the AI assistant interface, or shared by other users to the user and then the user uploads it to the AI assistant interface, which is not limited in the embodiments of the present application.
[0067] Continuing to refer to the first example above, after the user takes a photo of honeysuckle, as Figure 4 shown, click on the control 42 in the AI assistant interface 41, an image list 43 can be displayed. The image list 43 includes photo 431, photo 432, and photo 433. Among them, photo 431 is the photo of honeysuckle taken by the user, photo 432 is the user's personal photo, and photo 433 is a photo containing lilies and roses downloaded by the user from the web page. If the user clicks on photo 431, then as Figure 5 shown, photo 431 can be uploaded in the AI assistant interface 41. Then the user inputs the plant recognition indication information "Please recognize the plant in this picture" in the input box 424 of the AI assistant interface 41, and the honeysuckle in photo 431 can be recognized to obtain the recognition result information of honeysuckle as Figure 2 shown.
[0068] Continuing to refer to the second example above, after the user obtains the photo of lilies and roses, as Figure 4 shown, click on the control 42 in the AI assistant interface 41, an image list 43 can be displayed. If the user clicks on photo 433 in the image list 43, then as Figure 6 shown, photo 433 can be uploaded in the AI assistant interface 41. Then the user inputs the plant recognition indication information "Please recognize the plant on the right side in this picture" in the input box 424 of the AI assistant interface 41, and the lily on the right side in photo 433 can be recognized to obtain the recognition result information of lily as Figure 3 shown.
[0069] In an embodiment of the present application, when it is determined that plant recognition instruction information for recognizing the plants in the uploaded plant image is received from the user on the artificial intelligence assistant interface, the plants in the plant image are then recognized. In this way, the plants in the plant image can be recognized according to the timing of plant recognition required by the user, meeting the time requirements of the user for plant recognition.
[0070] It should be noted that there may be multiple plants in the plant image, and the user can select the plants they want to recognize according to their needs for recognition.
[0071] In some embodiments of the present application, in order to improve the flexibility of plant recognition, the recognition of the plants in the plant image specifically may include:
[0072] When the plant image includes at least two plants and a selection input of one of the plants in the plant image is received from the user, the plant selected by the selection input is recognized.
[0073] In some embodiments of the present application, when the plant image includes at least two plants, if the user only wants to recognize one of them, the user can make a selection input for the plant they want to recognize in the plant image. For example, by long pressing the plant they want to recognize, the plant the user wants to recognize can be extracted. In this way, when recognizing the plants in the plant image, only the plant selected by the user's selection input can be recognized.
[0074] In some embodiments of the present application, when extracting the plant the user wants to recognize based on the user's selection input, the bounding box of the plant selected by the user's selection input can be extracted based on a target detection algorithm, and the target detection algorithm can be, for example, the YOLO algorithm.
[0075] Continuing to refer to the second example above, refer to Figure 6 , photo 433 includes two plants: plant 4331 and plant 4332. The user knows that plant 4331 is a rose. If the user wants to recognize plant 4332, the user can long press plant 4332, and then plant 4332 can be extracted. Specifically, plant 4332 is framed by a dotted line box 61. In this way, when subsequently recognizing the plants in photo 433, only the plant 4332 framed by the dotted line box 61 in photo 433 can be recognized to obtain the recognition result information of plant 4332.
[0076] In an embodiment of the present application, when the plant image includes at least two plants, if a selection input of one of the plants in the plant image is received from the user, only the plant selected by the selection input can be recognized. In this way, the plants the user wants to recognize can be recognized according to the user's needs, improving the flexibility of plant recognition.
[0077] In some embodiments of the present application, in order to facilitate the user's awareness of the plant recognition result information, step 120 may specifically include:
[0078] On the artificial intelligence assistant interface, display the plant recognition result information.
[0079] In some embodiments of the present application, when the user uploads a plant image in the AI assistant interface and inputs plant recognition instruction information in the AI assistant interface, the plant recognition result information can be displayed in the AI assistant interface when it is displayed.
[0080] In the embodiments of the present application, by displaying the plant recognition result information in the artificial intelligence assistant interface, it is possible to facilitate the user's awareness of the plant recognition result information.
[0081] In some embodiments of the present application, in order to improve the acquisition efficiency of the multi-modal feature information of the plant in the plant image, step 110 may specifically include:
[0082] Input the plant image into the feature extractor in the large language model to identify the plant in the plant image and output the morphological description information of the plant;
[0083] Input the plant image and the morphological description information into the multi-modal feature extractor in the large language model to output the plant multi-modal feature information;
[0084] Obtain the plant recognition result information according to the plant multi-modal feature information.
[0085] Among them, the feature extractor can be a device for extracting the external morphological features of the plant in the plant image.
[0086] The multi-modal feature extractor can be a device for extracting multi-modal feature information.
[0087] The multi-modal feature information can be information of multiple modalities of the plant, such as the image of the plant, the geographical information of the plant, and the functional value information, etc.
[0088] In some embodiments of the present application, by inputting the plant image into the feature extractor in the large language model, the morphological description information of the plant can be output, and then by inputting the plant image and the morphological description information into the multi-modal feature extractor in the large language model, the plant multi-modal feature information can be output.
[0089] It should be noted that the electronic device may be integrated with a large language model to identify the plant in the plant image based on the large language model and obtain the plant recognition result information.
[0090] Continue to refer to the above Figure 5, after uploading photo 431 to the AI assistant interface 41, based on the feature extractor in the large language model integrated in the electronic device, the plants in the plant image can be identified, and the morphological description information of the plants can be output: the branches are grayish-brown, with hard rough hairs or bristles, the hair color is reddish-brown, and the leaves are disc-shaped. Then, photo 431 and the morphological description information "the branches are grayish-brown, with hard rough hairs or bristles, the hair color is reddish-brown, and the leaves are disc-shaped" are input into the multi-modal feature extractor in the large language model, and the multi-modal feature information of the plants is output: Lonicera: belonging to the genus Lonicera, Caprifoliaceae, widely produced in East Asia, it is cold in nature, sweet in taste, and slightly sour. Then, according to the multi-modal feature information of the plants, the plant recognition result information can be obtained: the attribute information of Lonicera: Lonicera, belonging to the genus Lonicera, Caprifoliaceae, the geographical distribution information of Lonicera: widely produced in East Asia, the functional value information of Lonicera: cold in nature, sweet in taste, slightly sour, and the morphological description information and images of Lonicera at different growth stages.
[0091] In the embodiment of the present application, the multi-modal feature information is obtained by identifying the plants in the plant image through the large language model, which improves the acquisition efficiency of the multi-modal feature information of the plants in the plant image.
[0092] In some embodiments of the present application, in order to improve the accuracy of the plant recognition result information, before obtaining the plant recognition result information according to the multi-modal feature information of the plants, the methods involved above may further include:
[0093] Input the plant image and the morphological description information into the multi-modal feature extractor, and output the confidence level of the multi-modal feature information of the plants;
[0094] The obtaining of the plant recognition result information according to the multi-modal feature information of the plants may specifically include:
[0095] In the case where the confidence level is greater than or equal to the confidence level threshold, obtain the plant recognition result information according to the multi-modal feature information of the plants.
[0096] Among them, the confidence level threshold may be a pre-set threshold of the confidence level of the multi-modal feature information of the plants, and the value range of the confidence level threshold may be 70%-100%. For example, the confidence level threshold may be set to 80% or 85%. The specific value of the confidence level threshold can be set according to user needs and is not limited in the embodiment of the present application.
[0097] In some embodiments of the present application, after inputting the plant image and the morphological description information into the multi-modal feature extractor in the large language model, the confidence level of the multi-modal feature information of the plants will be output while outputting the multi-modal feature information of the plants. In the case where the confidence level is greater than or equal to the confidence level threshold, then obtain the plant recognition result information according to the multi-modal feature information of the plants.
[0098] In the case where the confidence level of the output plant multi-modal feature information is less than the confidence level threshold, it may be that the clarity of the plant image uploaded by the user is not good enough. In this case, the user can re-upload an image of the plant to be recognized, that is, the electronic device can re-receive the plant image uploaded by the user, and re-recognize the plant in the re-uploaded plant image to obtain the plant multi-modal feature information and confidence level of the plant, until the obtained confidence level is greater than or equal to the confidence level threshold, and the user stops re-uploading the plant image.
[0099] It should be noted that the plant in the re-uploaded plant image belongs to the same plant category as the plant in the initially uploaded plant image, but only the image information of the re-uploaded plant image is different from the image information of the initially uploaded plant image. The image information here can be information representing the image quality, such as the contrast, brightness, and shooting angle of the image.
[0100] Continue to refer to Figure 5 , taking the confidence level threshold as 80% as an example, after uploading photo 431 to the AI assistant interface 41, if the confidence level of the plant multi-modal feature information of Lonicera in photo 431 obtained is lower than 80%, then re-upload a photo containing Lonicera, and re-recognize Lonicera in the photo to obtain the plant multi-modal feature information of Lonicera and the confidence level of the plant multi-modal feature information of Lonicera. If the confidence level is higher than 80%, then according to the plant multi-modal feature information, the recognition result information of Lonicera can be obtained. If the confidence level is lower than 80%, then re-upload a photo containing Lonicera until the confidence level of the plant multi-modal feature information of Lonicera obtained is higher than 80%.
[0101] In some embodiments of the present application, in the case where the confidence level of the plant multi-modal feature information output by the multi-modal feature extractor is greater than or equal to the confidence level threshold, when obtaining the plant recognition result information according to the plant multi-modal feature information, the plant multi-modal feature information with a confidence level lower than the confidence level threshold extracted in all history is weighted and fused to obtain the final plant multi-modal feature information, and then the plant recognition result information is obtained according to the final plant multi-modal feature information.
[0102] In the embodiments of the present application, in the case where the confidence level of the plant multi-modal feature information output by the multi-modal feature extractor is greater than or equal to the confidence level threshold, and then the plant recognition result information is obtained according to the plant multi-modal feature information, thereby improving the accuracy of the plant recognition result information.
[0103] In some embodiments of the present application, for the determination efficiency of the plant recognition result information, the obtaining of the plant recognition result information according to the plant multi-modal feature information may specifically include:
[0104] Compare the plant multi-modal feature information with the reference multi-modal feature information of each reference plant image in the plant image library, and output a set of plant candidate recognition results;
[0105] Construct a plant knowledge graph based on the plant multi-modal feature information and the alternative multi-modal feature information;
[0106] Generate a node matrix and an adjacency matrix for each graph node of the plant knowledge graph based on the plant multi-modal feature information and the alternative multi-modal feature information;
[0107] Input the node matrix and the adjacency matrix of each graph node into a graph neural network model, and output plant recognition result information.
[0108] Among them, the plant image library can be a pre-set library containing multiple reference plant images, and the multi-modal feature information corresponding to each reference plant image is also included in the plant image library.
[0109] The reference plant image can be a plant image used to compare with the plant image uploaded by the user.
[0110] The reference multi-modal feature information can be the multi-modal feature information of the reference plant image.
[0111] The set of plant candidate recognition results can be the recognition result set obtained after comparing the plant multi-modal feature information with the reference multi-modal feature information of each reference plant image in the plant image library.
[0112] The above-mentioned set of plant candidate recognition results may include alternative multi-modal feature information and the reference plant image corresponding to the alternative multi-modal feature information. The alternative multi-modal feature information can be part of the reference multi-modal feature information selected from the reference multi-modal feature information of multiple reference plant images in the plant image library. Specifically, the alternative multi-modal feature information may include the reference multi-modal feature information in each item of reference multi-modal feature information whose similarity with the plant multi-modal feature information is greater than the similarity threshold, that is, the alternative multi-modal feature information is the reference multi-modal feature information in each reference multi-modal feature information whose similarity with the plant multi-modal feature information is greater than the similarity threshold.
[0113] The above-mentioned similarity threshold can be a pre-set threshold for the similarity between the reference multi-modal feature information and the plant multi-modal feature information. For example, the value range of the similarity threshold can be 60%-100%. For example, the similarity threshold can be set to 70%, and the similarity threshold can also be set to 80%. The specific value of the similarity threshold can be set according to user needs and is not limited in the embodiments of the present application.
[0114] The plant knowledge graph can be a knowledge graph constructed with plant multi-modal feature information and alternative multi-modal feature information as nodes, that is, each graph node of the plant knowledge graph includes plant multi-modal feature information and alternative multi-modal feature information.
[0115] For any graph node, its node matrix can be a matrix used to represent the graph node.
[0116] For any graph node, its adjacency matrix can be a matrix used to describe the relationship between the graph node and other graph nodes.
[0117] In some embodiments of the present application, the plant multi-modal feature information can be compared with the reference multi-modal feature information of each reference plant image in the plant image library to obtain a set of plant candidate recognition results. Specifically, it can be retrieved in the plant image library based on the plant multi-modal feature information using the K-Nearest Neighbor (KNN) classification algorithm to obtain the set of candidate recognition results.
[0118] Then, based on the plant multi-modal feature information and the alternative multi-modal feature information in the set of plant candidate recognition results, a plant knowledge graph is constructed. By calculating the node matrix and adjacency matrix of each graph node of the plant knowledge graph, the target recognition result information of the plant can be obtained.
[0119] It should be noted that the determination method of the reference multi-modal feature information of each reference plant image in the plant image library is the same as the determination method of the above plant multi-modal feature information, except that the confidence of the multi-modal feature information is not judged.
[0120] It should be noted that each reference plant image in the plant image library, as well as the reference multi-modal feature information of each reference plant image, can be displayed in the form of a knowledge graph, such as Figure 7 shown, Figure 7 is a schematic diagram of the knowledge graph of Lonicera japonica.
[0121] If the user wants to view Figure 7 the detailed information of each modal feature information in Figure 7 , the corresponding control of each modal feature information in Figure 7 can be clicked, and the detailed information of the modal feature information will be displayed. As Figure 8 shown, when the user clicks on "Morphological description information" 71, the morphological description information of Lonicera japonica "The branches are grayish-brown, hard, rough hairs or bristles, the hair color is brownish-red, and the leaves are disk-shaped" can be displayed as Figure 7 shown. When the user clicks on "Attribute information" 72 in Figure 9 , the attribute information of Lonicera japonica "Lonicera japonica, belonging to the genus Lonicera, Caprifoliaceae" can be displayed as Figure 7 shown. When the user clicks on "Functional value" 73 inFigure 10 The displayed functional value information of honeysuckle is "cold in nature, sweet in taste, and slightly sour". When the user clicks Figure 7 on "Geographical Distribution" 74 in Figure 11 it, the geographical distribution information of honeysuckle, "widely produced in East Asia", can be displayed as shown. When the user clicks Figure 7 on "Picture Gallery" 75 in
[0122] In the embodiments of the present application, by selecting reference plant images with high similarity to the plant multimodal feature information from the plant image library, and the reference multimodal feature information of the reference plant images, a plant candidate recognition result set is obtained. Then, based on the plant multimodal feature information and the reference multimodal feature information in the plant candidate recognition result set, the plant recognition result information is determined. In this way, it is not necessary to calculate the reference multimodal feature information of all reference plant images in the plant image library, improving the determination efficiency of the plant recognition result information.
[0123] In some embodiments of the present application, in order to improve the recognition accuracy of the plants in the plant images, the plant multimodal feature information is compared with the reference multimodal feature information of each reference plant image in the plant image library, and a plant candidate recognition result set is output, which may specifically include:
[0124] Determine the similarity between the plant multimodal feature information and the reference multimodal feature information of each reference plant image in the plant image library;
[0125] Determine all the reference multimodal feature information with similarity greater than the similarity threshold as the alternative multimodal feature information;
[0126] Determine the alternative multimodal feature information and the reference plant images corresponding to the alternative multimodal feature information as the plant candidate recognition result set.
[0127] Among them, the similarity threshold can be a pre-set threshold for the similarity between the plant multimodal feature information and the reference multimodal feature information of each reference plant image in the plant image library. The similarity threshold can take a value of 80%. The specific value of the similarity threshold can be set by the user according to needs and is not limited in the embodiments of the present application.
[0128] In some embodiments of the present application, by determining the similarity between the plant multimodal feature information and the reference multimodal feature information of each reference plant image in the plant image library, the reference multimodal feature information with similarity greater than the similarity threshold is selected as the alternative multimodal feature information, and then the alternative multimodal feature information and the reference plant images corresponding to the alternative multimodal feature information are jointly determined as the plant candidate recognition result set.
[0129] Continuing to refer to the first example above, taking the similarity threshold as 80% as an example, the reference plant images in the plant image library include Image 1 of the mature stage of Lonicera japonica, Image 2 of tulips, and Image 3 of cypress. The multimodal feature information of Lonicera japonica, "Lonicera japonica: belonging to the genus Lonicera, family Caprifoliaceae, widely produced in East Asia, with a cold nature, sweet taste, and slightly sour taste", is compared with the multimodal feature information of Image 1, the multimodal feature information of Image 2, and the multimodal feature information of Image 3 respectively. If the similarity between the multimodal feature information of Image 1 and the multimodal feature information of Lonicera japonica is 90%, the similarity between the multimodal feature information of Image 2 and the multimodal feature information of Lonicera japonica is 40%, and the similarity between the multimodal feature information of Image 3 and the multimodal feature information of Lonicera japonica is 48%, then the multimodal feature information of Image 1 can be used as the alternative multimodal feature information, and the multimodal feature information of Image 1, as well as Image 1, are used as the plant candidate recognition result set.
[0130] In some embodiments of the present application, when selecting the plant candidate recognition result set, by comparing the similarity between the plant multimodal feature information and the reference multimodal feature information of each reference plant image in the plant image library, the reference multimodal feature information with a similarity higher than the similarity threshold is selected as the alternative multimodal feature information, and the alternative multimodal feature information and the reference plant image corresponding to the alternative multimodal feature information are used as the plant candidate recognition result set. In this way, it can be ensured that the plants in the reference plant images of the selected plant candidate recognition result set are highly similar to the plants in the plant image, improving the recognition accuracy of the plants in the plant image.
[0131] In some embodiments of the present application, in order to reduce the processing amount of multimodal feature information and improve the determination efficiency of plant recognition result information, generating the node matrix and adjacency matrix of each graph node of the plant knowledge graph based on the plant multimodal feature information and the alternative multimodal feature information specifically may include:
[0132] Normalize the plant multimodal feature information and the alternative multimodal feature information respectively to obtain the node matrix of each graph node of the plant knowledge graph;
[0133] According to the node matrix of each graph node of the plant knowledge graph and the transposed matrix of the node matrix of each graph node, obtain the adjacency matrix of each graph node.
[0134] In some embodiments of the present application, the plant multimodal feature information and the alternative multimodal feature information can be normalized respectively. Specifically, the plant multimodal feature information and the alternative multimodal feature information can be normalized by L2 norm respectively, and the node matrix of each graph node of the plant knowledge graph can be obtained.
[0135] It should be noted that for each graph node, the node matrix corresponding to the graph node can be a matrix with M rows and N columns. The data in each row is used to represent one modal feature information in the multi-modal feature information corresponding to the graph node, and the data in each column of each row is used to represent an information element of the modal feature information corresponding to the row.
[0136] It should be noted that the number of columns in each row is the same. If the information represented by a certain row is insufficient, the value "0" is used instead.
[0137] In an example, the multi-modal feature information corresponding to a certain graph node is "Lonicera, belonging to the genus Lonicera, family Caprifoliaceae, widely produced in East Asia, cold in nature, sweet in taste, slightly sour in taste". It has a total of three levels of feature information, namely: attribute information "Lonicera, belonging to the genus Lonicera, family Caprifoliaceae", geographical distribution information "widely produced in East Asia", and functional value information "cold in nature, sweet in taste, slightly sour in taste". Then the values in the node matrix of this graph node are three rows. Among them, the first row is the data representing the attribute information, the second row is the data representing the geographical distribution information, and the third row is the data representing the functional value information. The first row has 3 columns of data. Among them, the first column is the data representing "Lonicera species", the second column is the data representing "genus Lonicera", and the third column is the data representing "family Caprifoliaceae". The second row also has 3 columns of data. Among them, the first column is the data representing "widely produced in East Asia", and the second and third columns are both the value "0". The third row also has 3 columns of data. Among them, the first column is the data representing "cold in nature", the second column is the data representing "sweet in taste", and the third column is the data representing "slightly sour in taste".
[0138] After obtaining the node matrix of each graph node, the adjacency matrix of each graph node can be obtained according to the node matrix of each graph node of the plant knowledge graph and the transpose matrix of the node matrix of each graph node. Specifically, for any graph node, the adjacency matrix of each graph node can be obtained according to the following formula (1):
[0139] A = X Τ · X (1)
[0140] In the above formula (1), A is the adjacency matrix of a certain graph node, X is the node matrix of a certain graph node, and X Τ is the transpose of X.
[0141] In an embodiment of the present application, by performing normalization processing on the plant multimodal feature information and the alternative multimodal feature information respectively, a node matrix of each graph node of the plant knowledge graph can be obtained. Then, according to the node matrix of each graph node of the plant knowledge graph and the transposed matrix of the node matrix of each graph node, an adjacency matrix of each graph node can be obtained. In this way, each graph node can be quantified, and there is no need to use the multimodal feature information represented by the graph node for subsequent processing, reducing the processing amount of the multimodal feature information and improving the determination efficiency of the plant recognition result information.
[0142] In some embodiments of the present application, the above graph neural network model may include a graph convolution module, a fully connected layer, and an activation function layer. In order to accurately obtain the plant recognition result information, inputting the node matrix and the adjacency matrix of each graph node into the graph neural network model and outputting the plant recognition result information may specifically include:
[0143] Inputting the node matrix and the adjacency matrix of each graph node into the graph convolution module to obtain the first hidden state parameter of the plant;
[0144] Performing global average pooling processing on the first hidden state parameter to obtain the second hidden state parameter;
[0145] Inputting the second hidden state parameter into the fully connected layer, and through the fully connected layer, performing feature extraction on the second hidden state parameter to obtain the feature information of the second hidden state parameter;
[0146] Inputting the feature information of the second hidden state parameter into the activation function layer, and through the activation function in the activation function layer, performing nonlinear transformation on the feature information of the second hidden state parameter and outputting the plant recognition result information.
[0147] Among them, the first hidden state parameter may be the hidden state parameter of the plant obtained after inputting the node matrix and the adjacency matrix of each graph node into the graph convolution module.
[0148] The second hidden state parameter may be the hidden state parameter obtained by performing global average pooling processing on the first hidden state parameter.
[0149] In some embodiments of the present application, the node matrix and adjacency matrix of each graph node can be input into a graph convolution module to obtain the first hidden state parameter of the plant. Then, the first hidden state parameter is subjected to global average pooling to obtain the second hidden state parameter. The second hidden state parameter is input into a fully connected layer to extract features of the second hidden state parameter, obtaining the feature information of the second hidden state parameter. Furthermore, the feature information of the second hidden state parameter is input into an activation function layer, and based on the activation function in the activation function layer, a non-linear transformation is performed on the feature information of the second hidden state parameter, and the plant recognition result information can be obtained.
[0150] It should be noted that the activation function in the activation function layer can be a softmax activation function.
[0151] In the embodiments of the present application, by processing the node matrix and adjacency matrix of each graph node through a graph neural network model, the relationship and topological structure between each graph node can be well captured, so that accurate plant recognition result information can be obtained.
[0152] In some embodiments of the present application, the graph convolution module may include M graph convolution layers, where M is a positive integer; the step of inputting the node matrix and adjacency matrix of each graph node into the graph convolution module to obtain the first hidden state parameter of the plant may specifically include:
[0153] Input the adjacency matrix and node matrix of each graph node into the i-th graph convolution layer to obtain the third hidden state parameter, where the initial value of i is 1;
[0154] Input the third hidden state parameter into the (i + 1)-th graph convolution layer, and according to the third hidden state parameter and the network parameters of the (i + 1)-th graph convolution layer, obtain the fourth hidden state parameter;
[0155] Update the third hidden state parameter to the fourth hidden state parameter, and return to execute the step of inputting the third hidden state parameter into the (i + 1)-th graph convolution layer, and according to the third hidden state parameter and the network parameters of the (i + 1)-th graph convolution layer, obtain the fourth hidden state parameter, until i + 1 = M, to obtain the first hidden state parameter of the plant.
[0156] Among them, the third hidden state parameter may be the hidden state parameter obtained after inputting the adjacency matrix and node matrix of each graph node into the i-th graph convolution layer.
[0157] The fourth hidden state parameter may be the hidden state parameter obtained by inputting the third hidden state parameter into the (i + 1)-th graph convolution layer and according to the third hidden state parameter and the network parameters of the (i + 1)-th graph convolution layer.
[0158] Specifically, according to the hidden state parameters output by the i-th layer of graph convolutional layer, the hidden state parameters output by the (i + 1)-th layer of graph convolutional layer can be obtained according to the following formula (2):
[0159] H k+1 = H k + σ(H k ) (2)
[0160] In the above formula (2), H k is the hidden state parameter output by the i-th layer of graph convolutional layer, σ is the network parameter of the (i + 1)-th layer of graph convolutional layer, and H k+1 is the hidden state parameter output by the (i + 1)-th layer of graph convolutional layer.
[0161] In some embodiments of the present application, the adjacency matrix and node matrix of each graph node can be input into the first layer of graph convolutional layer first to obtain the hidden state parameter 1, and then the hidden state parameter 1 is input into the second layer of graph convolutional layer. In order to fully retain the original information of the plant in the plant image, the hidden state parameter 1 is used for skip connection, that is, according to the hidden state parameter 1 and the network parameter of the second layer of graph convolutional layer, the hidden state parameter 2 is obtained, and then the hidden state parameter 2 is input into the third layer of graph convolutional layer to obtain the hidden state parameter 3, and so on, until the last layer of graph convolutional layer outputs the hidden state parameter, and the hidden state parameter output by the last layer of graph convolutional layer is used as the first hidden state parameter.
[0162] It should be noted that the value of the above M can be selected by the user according to needs, such as 16, and it is not limited in the embodiments of the present application.
[0163] In the embodiments of the present application, the adjacency matrix and node matrix of each graph node are processed by M graph convolutional layers in the graph convolutional module. In this way, by iteratively updating the feature vector of the plant image, the plant representation in the plant image can be accurately obtained, and then the plant recognition result information can be accurately obtained.
[0164] In some embodiments of the present application, when the graph neural network model processes the adjacency matrix and node matrix of each graph node, it can also obtain the relevant information of other plants associated with the plant in the plant image. If the user wants to view the relevant information of other plants associated with the plant in the plant image, it can be realized through the relevant controls in the AI assistant interface. As Figure 2 shown, if the user wants to view which plants are associated with Lonicera japonica, or which plants are in the same family as Lonicera japonica, then they can click Figure 2 the "Match Associated Data" control 23 in Figure 12 and then, as Figure 12The image 121 of the red spider lily and the image 122 of gladiolus shown.
[0165] In an embodiment of the present application, by responding to the user's input to the matching associated data control, information about other plants associated with the plant in the plant image can be displayed, so that the user can understand more information about other plants associated with the plant in the plant image.
[0166] For the plant recognition method provided by the embodiments of the present application, the execution subject can be a plant recognition device. In the embodiments of the present application, taking the plant recognition device executing the plant recognition method as an example, the plant recognition device provided by the embodiments of the present application is described.
[0167] Figure 13 It is a schematic structural diagram of a plant recognition device shown according to an exemplary embodiment. As Figure 13 shown, the plant recognition device 1300 may include:
[0168] The recognition module 1310 is used to recognize the plant in the plant image to obtain plant recognition result information; wherein, the plant recognition result information is generated according to the plant's multi-modal feature information of the plant.
[0169] The display module 1320 is used to display the plant recognition result information; wherein, the plant recognition result information includes at least one of the following: the attribute information of the plant, the geographical distribution information of the plant, the functional value information of the plant, the morphological description information and images of different growth periods of the plant.
[0170] In the embodiments of the present application, compared with the prior art in which plants are recognized based on the extracted color and texture features of plants, the solution of the embodiments of the present application refers to the multi-modal feature information of the plant in the plant image when recognizing plants, that is, in the embodiments of the present application, when recognizing plants, in addition to referring to the color and texture features of the plant in the plant image, other feature information of the plant is also referred to, thus improving the accuracy of plant recognition. In addition, after the plant is recognized, the obtained plant recognition result information includes at least one of the attribute information of the plant, the geographical distribution information of the plant, the functional value information of the plant, the morphological description information of different growth periods of the plant and the image, that is, the obtained plant recognition result information is more comprehensive, solving the limitations of the existing solutions in plant recognition.
[0171] In some embodiments of the present application, the above-mentioned device may further include:
[0172] The receiving module is used to receive the plant image uploaded by the user before recognizing the plant in the plant image to obtain the plant recognition result information; receive the plant recognition instruction information input by the user on the artificial intelligence AI assistant interface.
[0173] The recognition module 1310 is specifically configured to:
[0174] When receiving the plant recognition instruction information input by the user in the AI assistant interface, recognize the plant in the plant image to obtain the plant recognition result information.
[0175] In some embodiments of the present application, the recognition module 1310 is specifically configured to:
[0176] When the plant image includes at least two plants and a selection input of one plant in the plant image is received by the user, recognize the plant selected by the selection input.
[0177] In some embodiments of the present application, the display module 1320 is specifically configured to:
[0178] On the artificial intelligence AI assistant interface, display the plant recognition result information.
[0179] In some embodiments of the present application, the recognition module 1310 is specifically configured to:
[0180] Input the plant image into the feature extractor in the large language model, recognize the plant in the plant image, and output the morphological description information of the plant;
[0181] Input the plant image and the morphological description information into the multi-modal feature extractor in the large language model, and output the plant multi-modal feature information;
[0182] Obtain the plant recognition result information according to the plant multi-modal feature information.
[0183] In some embodiments of the present application, the recognition module 1310 is specifically configured to:
[0184] Compare the plant multi-modal feature information with the reference multi-modal feature information of each reference plant image in the plant library to obtain a plant candidate recognition result set. The plant candidate recognition result set includes alternative multi-modal feature information and the reference plant images corresponding to the alternative multi-modal feature information. The alternative multi-modal feature information includes the reference multi-modal feature information in each reference multi-modal feature information whose similarity with the plant multi-modal feature information is greater than the similarity threshold;
[0185] Based on the plant multi-modal feature information and the alternative multi-modal feature information, construct a plant knowledge graph. Each graph node of the plant knowledge graph includes the plant multi-modal feature information and the alternative multi-modal feature information;
[0186] Generate the node matrix and adjacency matrix of each graph node of the plant knowledge graph based on the plant multimodal feature information and the alternative multimodal feature information;
[0187] Input the node matrix and adjacency matrix of each graph node into the graph neural network model, and output the plant recognition result information.
[0188] In some embodiments of the present application, the recognition module 1310 is specifically configured to:
[0189] Determine the similarity between the plant multimodal feature information and the reference multimodal feature information of each reference plant image in the plant image library;
[0190] Determine all the reference multimodal feature information with a similarity greater than the similarity threshold as the alternative multimodal feature information;
[0191] Determine the alternative multimodal feature information and the reference plant image corresponding to the alternative multimodal feature information as the plant candidate recognition result set.
[0192] In some embodiments of the present application, the recognition module 1310 is specifically configured to:
[0193] Perform normalization processing on the plant multimodal feature information and the alternative multimodal feature information respectively to obtain the node matrix of each graph node of the plant knowledge graph;
[0194] According to the node matrix of each graph node of the plant knowledge graph and the transposed matrix of the node matrix of each graph node, obtain the adjacency matrix of each graph node.
[0195] In some embodiments of the present application, the graph neural network model includes a graph convolution module, a fully connected layer, and an activation function layer; the recognition module 1310 is specifically configured to:
[0196] Input the node matrix and the adjacency matrix of each graph node into the graph convolution module to obtain the first hidden state parameter of the plant;
[0197] Perform global average pooling processing on the first hidden state parameter to obtain a second hidden state parameter;
[0198] Input the second hidden state parameter into the fully connected layer, and through the fully connected layer, perform feature extraction on the second hidden state parameter to obtain the feature information of the second hidden state parameter;
[0199] Input the feature information of the second hidden state parameter into the activation function layer, and through the activation function in the activation function layer, perform nonlinear transformation on the feature information of the second hidden state parameter, and output the plant recognition result information.
[0200] In some embodiments of the present application, the recognition module 1310 is further configured to:
[0201] Input the plant image and the morphological description information into the multimodal feature extractor, and output the confidence of the plant multimodal feature information;
[0202] Specifically, the recognition module 1310 is configured to:
[0203] When the confidence is greater than or equal to the confidence threshold, recognize the plant in the plant image to obtain plant recognition result information.
[0204] The plant recognition device in the embodiments of the present application may be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device may be a terminal or other devices other than the terminal. Exemplarily, the electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. It may also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.
[0205] The plant recognition device in the embodiments of the present application may be a device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.
[0206] The plant recognition device provided in the embodiments of the present application can implement Figure 1 each process implemented by the method embodiments. To avoid repetition, it will not be elaborated here.
[0207] Optionally, as Figure 14As shown in the figure, an embodiment of the present application further provides an electronic device 1400, including a processor 1401 and a memory 1402. A program or instruction that can run on the processor 1401 is stored on the memory 1402. When the program or instruction is executed by the processor 1401, it implements each step of the above-mentioned embodiment of the plant recognition method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0208] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.
[0209] Figure 15 It is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application.
[0210] The electronic device 1500 includes, but is not limited to: a radio frequency unit 1501, a network module 1502, an audio output unit 1503, an input unit 1504, a sensor 1505, a display unit 1506, a user input unit 1507, an interface unit 1508, a memory 1509, and a processor 1510 and other components.
[0211] Those skilled in the art can understand that the electronic device 1500 may further include a power supply (such as a battery) for supplying power to each component. The power supply can be logically connected to the processor 1510 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. Figure 15 The structure of the electronic device shown in the figure does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0212] Among them, the processor 1510 is used to identify the plants in the plant image to obtain plant recognition result information; among them, the plant recognition result information is generated according to the plant multi-modal feature information of the plant;
[0213] The display unit 1506 is used to display the plant recognition result information; among them, the plant recognition result information includes at least one of the following: the attribute information of the plant, the geographical distribution information of the plant, the functional value information of the plant, the morphological description information and images of different growth periods of the plant.
[0214] Thus, compared with the prior art that identifies plants based on the extracted color and texture features of plants, the solution of the embodiment of the present application refers to the multi-modal feature information of plants in the plant image when identifying plants. That is, in the embodiment of the present application, when identifying plants, in addition to referring to the color and texture features of plants in the plant image, other feature information of plants is also referred to, thereby improving the accuracy of plant identification. In addition, after identifying plants, the obtained plant identification result information includes at least one of the plant's attribute information, geographical distribution information, functional value information, morphological description information of different growth periods of the plant, and the image, that is, the obtained plant identification result information is more comprehensive, solving the limitations of the existing solutions in plant identification.
[0215] Optionally, the user input unit 1507 is configured to receive a plant image uploaded by a user;
[0216] The user input unit 1507 is further configured to receive plant identification instruction information input by the user in the artificial intelligence AI assistant interface;
[0217] The processor 1510 is further configured to identify the plant in the plant image to obtain plant identification result information when receiving the plant identification instruction information input by the user in the AI assistant interface.
[0218] Thus, when it is determined that the plant identification instruction information for identifying the uploaded plant image input by the user in the artificial intelligence assistant interface is received, the plant in the plant image is then identified. In this way, according to the user's needs, when the user needs to obtain plant identification result information, the plant in the plant image is identified according to the plant identification instruction information input by the user, improving the flexibility of the plant identification time.
[0219] Optionally, the processor 1510 is further configured to identify the plant selected by the selection input when the plant image includes at least two plants and the user's selection input for one of the plants in the plant image is received.
[0220] Thus, when the plant image includes at least two plants, if the user's selection input for one of the plants in the plant image is received, only the plant selected by the selection input can be identified. In this way, the plant that the user wants to identify can be identified according to the user's needs, improving the flexibility of plant identification.
[0221] Optionally, the display unit 1506 is further configured to display the plant identification result information in the artificial intelligence AI assistant interface.
[0222] Thus, by displaying the plant identification result information in the artificial intelligence assistant interface, it is convenient for the user to know the plant identification result information.
[0223] Optionally, the processor 1510 is further configured to input the plant image into a feature extractor in a large language model to identify the plant in the plant image and output morphological description information of the plant; input the plant image and the morphological description information into a multi-modal feature extractor in the large language model to output plant multi-modal feature information; and obtain plant recognition result information according to the plant multi-modal feature information.
[0224] In this way, the large language model is used to identify the plant in the plant image to obtain multi-modal feature information, improving the acquisition efficiency of the multi-modal feature information of the plant in the plant image.
[0225] Optionally, the processor 1510 is further configured to compare the plant multi-modal feature information with the reference multi-modal feature information of each reference plant image in the plant image library to output a plant candidate recognition result set, where the plant candidate recognition result set includes alternative multi-modal feature information and the reference plant image corresponding to the alternative multi-modal feature information, and the alternative multi-modal feature information includes the reference multi-modal feature information in each item of reference multi-modal feature information whose similarity with the plant multi-modal feature information is greater than a similarity threshold; construct a plant knowledge graph based on the plant multi-modal feature information and the alternative multi-modal feature information, where each graph node of the plant knowledge graph includes the plant multi-modal feature information and the alternative multi-modal feature information; generate a node matrix and an adjacency matrix for each graph node of the plant knowledge graph based on the plant multi-modal feature information and the alternative multi-modal feature information; and input the node matrix and the adjacency matrix of each graph node into a graph neural network model to output plant recognition result information.
[0226] In this way, by selecting the reference plant image with a high similarity to the plant multi-modal feature information and the reference multi-modal feature information of the reference plant image from the plant image library, a plant candidate recognition result set is obtained, and then based on the plant multi-modal feature information and the reference multi-modal feature information in the plant candidate recognition result set, the plant recognition result information is determined. In this way, it is not necessary to calculate the reference multi-modal feature information of all reference plant images in the plant image library, improving the determination efficiency of the plant recognition result information.
[0227] Optionally, the processor 1510 is further configured to determine the similarity between the plant multi-modal feature information and the reference multi-modal feature information of each reference plant image in the plant image library; determine all the reference multi-modal feature information with a similarity greater than the similarity threshold as alternative multi-modal feature information; and determine the alternative multi-modal feature information and the reference plant image corresponding to the alternative multi-modal feature information as the plant candidate recognition result set.
[0228] In this way, when selecting the plant candidate recognition result set, by comparing the similarity between the plant multi-modal feature information and the reference multi-modal feature information of each reference plant image in the plant image library, the reference multi-modal feature information with a similarity higher than the similarity threshold is selected as the alternative multi-modal feature information, and the alternative multi-modal feature information and the reference plant image corresponding to the alternative multi-modal feature information are used as the plant candidate recognition result set. In this way, it can be ensured that the plants in the reference plant images of the selected plant candidate recognition result set are highly similar to the plants in the plant image, improving the recognition accuracy of the plants in the plant image.
[0229] Optionally, the processor 1510 is further configured to perform normalization processing on the plant multi-modal feature information and the alternative multi-modal feature information respectively to obtain the node matrix of each graph node of the plant knowledge graph; according to the node matrix of each graph node of the plant knowledge graph and the transposed matrix of the node matrix of each graph node, the adjacency matrix of each graph node is obtained.
[0230] In this way, by performing normalization processing on the plant multi-modal feature information and the alternative multi-modal feature information respectively, the node matrix of each graph node of the plant knowledge graph can be obtained. Then, according to the node matrix of each graph node of the plant knowledge graph and the transposed matrix of the node matrix of each graph node, the adjacency matrix of each graph node can be obtained. In this way, each graph node can be quantified, and there is no need to use the multi-modal feature information represented by the graph node for subsequent processing, reducing the processing amount of the multi-modal feature information and improving the determination efficiency of the plant recognition result information.
[0231] Optionally, the graph neural network model includes a graph convolution module, a fully connected layer, and an activation function layer; the processor 1510 is further configured to input the node matrix and the adjacency matrix of each graph node into the graph convolution module to obtain the first hidden state parameter of the plant; perform global average pooling processing on the first hidden state parameter to obtain a second hidden state parameter; input the second hidden state parameter into the fully connected layer, and through the fully connected layer, perform feature extraction on the second hidden state parameter to obtain the feature information of the second hidden state parameter; input the feature information of the second hidden state parameter into the activation function layer, and through the activation function in the activation function layer, perform non-linear transformation on the feature information of the second hidden state parameter, and output the plant recognition result information.
[0232] In this way, by processing the node matrix and the adjacency matrix of each graph node through the graph neural network model, the relationship and topological structure between each graph node can be well captured, and thus accurate plant recognition result information can be obtained.
[0233] Optionally, the processor 1510 is further configured to input the plant image and the morphological description information into the multimodal feature extractor, and output the confidence of the plant multimodal feature information; when the confidence is greater than or equal to a confidence threshold, obtain plant recognition result information according to the plant multimodal feature information.
[0234] In this way, when the confidence of the plant multimodal feature information output by the multimodal feature extractor is greater than or equal to the confidence threshold, the plant recognition result information is obtained according to the plant multimodal feature information, thereby improving the accuracy of the plant recognition result information.
[0235] It should be understood that, in the embodiment of the present application, the input unit 1504 may include a Graphics Processing Unit (GPU) 15041 and a microphone 15042. The graphics processor 15041 processes the image data of static pictures or videos obtained by an image capture device (such as a color camera) in a video capture mode or an image capture mode. The display unit 1506 may include a display panel 15061, and the display panel 15061 may be configured in the form of a liquid crystal display, an organic light emitting diode, or the like. The user input unit 1507 includes at least one of a touch panel 15071 and other input devices 15072. The touch panel 15071 is also referred to as a touch screen. The touch panel 15071 may include two parts: a touch detection device and a touch controller. The other input devices 15072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated herein.
[0236] The memory 1509 can be used to store software programs and various data. The memory 1509 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area may store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 1509 may include a volatile memory or a non-volatile memory, or the memory 1509 may include both a volatile memory and a non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM), and a direct rambus random access memory (DRRAM). The memory 1509 in the embodiments of the present application includes, but is not limited to, these and any other suitable types of memories.
[0237] The processor 1510 may include one or more processing units; optionally, the processor 1510 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above modem processor may not be integrated into the processor 1510 either.
[0238] The embodiments of the present application further provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above-mentioned embodiment of the plant recognition method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0239] Among them, the processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory ROM, random access memory RAM, magnetic disks, or optical discs, etc.
[0240] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement each process of the above-mentioned plant recognition method embodiment, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0241] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.
[0242] The embodiments of the present application provide a computer program product. The program product is stored in a storage medium. The program product is executed by at least one processor to implement each process of the above-mentioned plant recognition method embodiment, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0243] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article, or device. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed. It may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0244] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0245] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.
Claims
1. A plant identification method, characterized in that: The method comprises: Identify the plants in the plant image to obtain plant identification result information; wherein the plant identification result information is generated based on the plant multimodal feature information of the plant; Display plant identification result information; wherein the plant identification result information includes at least one of the following: attribute information of the plant, geographical distribution information of the plant, functional value information of the plant, morphological description information and images of the plant at different growth periods.
2. The method according to claim 1, characterized in that Before identifying the plants in the plant image and obtaining the plant identification result information, the method further includes: Receiving plant images uploaded by users; Receive plant identification instruction information input by the user in the artificial intelligence AI assistant interface; The step of identifying the plants in the plant image and obtaining plant identification result information includes: Upon receiving plant identification indication information input by the user in the AI assistant interface, the plants in the plant image are identified to obtain plant identification result information.
3. The method according to claim 2, characterized in that The identifying of plants in the plant image includes: When the plant image includes at least two plants and a selection input of a plant in the plant image is received from a user, the plant selected by the selection input is identified.
4. The method according to claim 1, characterized in that: The display of plant identification result information includes: On the artificial intelligence AI assistant interface, the plant identification result information is displayed.
5. The method according to claim 1, characterized in that: The step of identifying the plants in the plant image and obtaining plant identification result information includes: Inputting the plant image into a feature extractor in a large language model, identifying the plant in the plant image, and outputting morphological description information of the plant; Inputting the plant image and the morphological description information into a multimodal feature extractor in the large language model, and outputting plant multimodal feature information; Plant identification result information is obtained based on the plant multimodal feature information.
6. The method according to claim 5, characterized in that The obtaining of plant identification result information according to the plant multimodal feature information includes: Comparing the plant multimodal feature information with the reference multimodal feature information of each reference plant image in the plant library, and outputting a plant candidate recognition result set, wherein the plant candidate recognition result set includes candidate multimodal feature information and reference plant images corresponding to the candidate multimodal feature information, and the candidate multimodal feature information includes reference multimodal feature information in each reference multimodal feature information whose similarity with the plant multimodal feature information is greater than a similarity threshold; Based on the plant multimodal feature information and the candidate multimodal feature information, construct a plant knowledge graph, wherein each graph node of the plant knowledge graph includes the plant multimodal feature information and the candidate multimodal feature information; Based on the plant multimodal feature information and the candidate multimodal feature information, generating a node matrix and an adjacency matrix of each graph node of the plant knowledge graph; The node matrix and adjacency matrix of each graph node are input into the graph neural network model, and the plant identification result information is output.
7. The method according to claim 6, characterized in that The step of comparing the plant multimodal feature information with the reference multimodal feature information of each reference plant image in the plant library and outputting a plant candidate recognition result set includes: Determining the similarity between the plant multimodal feature information and reference multimodal feature information of each reference plant image in the plant gallery; Determine all reference multimodal feature information with similarities greater than a similarity threshold as candidate multimodal feature information; The candidate multimodal feature information and a reference plant image corresponding to the candidate multimodal feature information are determined as a plant candidate recognition result set.
8. The method according to claim 6, characterized in that The step of generating a node matrix and an adjacency matrix of each graph node of the plant knowledge graph based on the plant multimodal feature information and the candidate multimodal feature information includes: Normalizing the plant multimodal feature information and the candidate multimodal feature information respectively to obtain a node matrix of each graph node of the plant knowledge graph; According to the node matrix of each graph node of the plant knowledge graph and the transposed matrix of the node matrix of each graph node, the adjacency matrix of each graph node is obtained.
9. The method according to claim 6, characterized in that The graph neural network model includes a graph convolution module, a fully connected layer and an activation function layer; The node matrix and adjacency matrix of each graph node are input into the graph neural network model, and the plant identification result information is output, including: Inputting the node matrix and the adjacency matrix of each graph node into the graph convolution module to obtain the first hidden state parameter of the plant; Performing a full average pooling process on the first hidden state parameter to obtain a second hidden state parameter; Inputting the second hidden state parameter into the fully connected layer, and performing feature extraction on the second hidden state parameter through the fully connected layer to obtain feature information of the second hidden state parameter; The characteristic information of the second hidden state parameter is input into the activation function layer, and the characteristic information of the second hidden state parameter is nonlinearly transformed through the activation function in the activation function layer, and the plant identification result information is output.
10. The method according to claim 5, characterized in that Before obtaining plant identification result information according to the plant multimodal feature information, the method further includes: Inputting the plant image and the morphological description information into the multimodal feature extractor, and outputting the confidence of the plant multimodal feature information; The obtaining of plant identification result information according to the plant multimodal feature information includes: When the confidence level is greater than or equal to a confidence level threshold, plant identification result information is obtained according to the plant multimodal feature information.
11. A plant identification device, characterized in that: The device comprises: An identification module, used to identify plants in the plant image and obtain plant identification result information; wherein the plant identification result information is generated based on plant multimodal feature information of the plant; The display module is used to display plant identification result information; wherein the plant identification result information includes at least one of the following: attribute information of the plant, geographical distribution information of the plant, functional value information of the plant, morphological description information and images of the plant at different growth periods.
12. The device according to claim 11, characterized in that The device also includes: The receiving module is used to receive the plant image uploaded by the user before identifying the plant in the plant image and obtaining the plant identification result information; and receive the plant identification indication information input by the user in the artificial intelligence AI assistant interface; The identification module is specifically used for: Upon receiving plant identification indication information input by the user in the AI assistant interface, the plants in the plant image are identified to obtain plant identification result information.
13. The device according to claim 12, characterized in that The identification module is specifically used for: When the plant image includes at least two plants and a selection input of a plant in the plant image is received from a user, the plant selected by the selection input is identified.
14. The device according to claim 11, characterized in that The display module is specifically used for: On the artificial intelligence AI assistant interface, the plant identification result information is displayed.
15. The device according to claim 11, characterized in that The identification module is specifically used for: Inputting the plant image into a feature extractor in a large language model, identifying the plant in the plant image, and outputting morphological description information of the plant; Inputting the plant image and the morphological description information into a multimodal feature extractor in the large language model, and outputting plant multimodal feature information; Plant identification result information is obtained based on the plant multimodal feature information.
16. The device according to claim 15, characterized in that The identification module is specifically used for: Comparing the plant multimodal feature information with the reference multimodal feature information of each reference plant image in the plant library to obtain a plant candidate recognition result set, wherein the plant candidate recognition result set includes candidate multimodal feature information and reference plant images corresponding to the candidate multimodal feature information, and the candidate multimodal feature information includes reference multimodal feature information in each reference multimodal feature information whose similarity with the plant multimodal feature information is greater than a similarity threshold; Based on the plant multimodal feature information and the candidate multimodal feature information, construct a plant knowledge graph, wherein each graph node of the plant knowledge graph includes the plant multimodal feature information and the candidate multimodal feature information; Based on the plant multimodal feature information and the candidate multimodal feature information, generating a node matrix and an adjacency matrix of each graph node of the plant knowledge graph; The node matrix and adjacency matrix of each graph node are input into the graph neural network model, and the plant identification result information is output.
17. The device according to claim 16, characterized in that The identification module is specifically used for: Determining the similarity between the plant multimodal feature information and reference multimodal feature information of each reference plant image in the plant gallery; Determine all reference multimodal feature information with similarities greater than a similarity threshold as candidate multimodal feature information; The candidate multimodal feature information and a reference plant image corresponding to the candidate multimodal feature information are determined as a plant candidate recognition result set.
18. The device according to claim 16, characterized in that The identification module is specifically used for: Normalizing the plant multimodal feature information and the candidate multimodal feature information respectively to obtain a node matrix of each graph node of the plant knowledge graph; According to the node matrix of each graph node of the plant knowledge graph and the transposed matrix of the node matrix of each graph node, the adjacency matrix of each graph node is obtained.
19. The device according to claim 16, characterized in that The graph neural network model includes a graph convolution module, a fully connected layer and an activation function layer; the recognition module is specifically used for: Inputting the node matrix and the adjacency matrix of each graph node into the graph convolution module to obtain the first hidden state parameter of the plant; Performing a full average pooling process on the first hidden state parameter to obtain a second hidden state parameter; Inputting the second hidden state parameter into the fully connected layer, and performing feature extraction on the second hidden state parameter through the fully connected layer to obtain feature information of the second hidden state parameter; The characteristic information of the second hidden state parameter is input into the activation function layer, and the characteristic information of the second hidden state parameter is nonlinearly transformed through the activation function in the activation function layer, and the plant identification result information is output.
20. The device according to claim 15, characterized in that The identification module is also used for: Inputting the plant image and the morphological description information into the multimodal feature extractor, and outputting the confidence of the plant multimodal feature information; The identification module is specifically used for: When the confidence level is greater than or equal to the confidence level threshold, the plants in the plant image are identified to obtain plant identification result information.
21. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, the steps of the plant identification method according to any one of claims 1 to 10 are implemented.