Plant identification method and related device
By continuing to pre-train the plant knowledge question and answer model combining large language models and visual models on the plant field corpus, the problem of inefficient plant recognition is solved and more efficient plant recognition and maintenance suggestions are achieved.
Patent Information
- Application Number
- CN202510213405.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art has problems of inefficiency in plant recognition and difficulty in obtaining knowledge, especially when the corpus size of the plant field is small.
The plant knowledge question and answer model combining large language models and visual models is adopted, and the performance of large language models in the plant field is improved through the synthetic corpus generated based on the plant field corpus.
The recognition ability and user experience of the plant knowledge question and answer model in the plant field has been improved, and the performance of the model in plant recognition has been enhanced through a larger scale and diversified synthetic corpus.
Smart Images

Figure CN120124746A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of information processing technology, and more particularly, to a plant recognition method, an electronic device, a non-transitory storage medium, and a computer program product, and also relates to a maintenance system. Background Art
[0002] Currently, there are various applications (APPs) for recognizing plants, such as applications for recognizing plants. These applications typically receive an image and a question input by a user, and identify an object in the image through a recognition model established based on artificial intelligence technology to obtain an answer to the question, and present the answer to the user on a user interface. Summary of the Invention
[0003] A brief overview of the present disclosure is given below to provide a basic understanding of some aspects of the present disclosure. However, it should be understood that this overview is not an exhaustive overview of the present disclosure. It is not intended to identify the key or important parts of the present disclosure, nor is it intended to limit the scope of the present disclosure. Its purpose is merely to present some concepts of the present disclosure in a simplified form as a prelude to the more detailed description given later.
[0004] According to a first aspect of the present disclosure, there is provided a plant recognition method, including: obtaining a plant image and question text for recognizing the plant in the plant image; inputting the plant image and the question text into a plant knowledge-answering model, the plant knowledge-answering model including a vision model and a large language model, wherein the vision model is configured to receive the plant image to extract image features of the plant image, the large language model is configured to receive the image features and the question text to recognize the plant in the plant image, the large language model is further pre-trained on the basis of pre-training with a synthetic corpus generated based on a plant domain corpus, the plant knowledge-answering model is trained with multimodal data, the multimodal data including plant images, questions for recognizing the plant in the plant images, and answers to the questions; and outputting answer text for recognizing the plant in the plant image provided by the plant knowledge-answering model.
[0005] In some embodiments, the synthetic corpus includes a mixture of plant domain corpus and synthetic corpus.
[0006] In some embodiments, the synthetic corpus includes a mixture of general domain corpus and synthetic corpus.
[0007] In some embodiments, the quantity of the general domain corpus is in a preset ratio to the quantity of the synthetic corpus, and the preset ratio is determined through the following operations: using the general domain corpus and the synthetic corpus mixed in different ratios as training data to train a large language model respectively, and using validation data to determine the validation loss of the trained large language model; and selecting a ratio from different ratios as the preset ratio based on the validation loss.
[0008] In some embodiments, the synthetic corpus is constructed through the following operations: segmenting the plant domain corpus in the plant domain corpus into one or more segments according to semantics; for each of the one or more segments, extracting the entities in the segment through a language model to obtain an entity set, based on the entities in each entity subset of the entity set extracted from the segment, obtaining the mutual relationships between the entities, and generating synthetic corpus from the segment based on the extracted entities and / or the mutual relationships between the entities.
[0009] In some embodiments, generating synthetic corpus from the segment based on the extracted entities and / or the mutual relationships between the entities includes at least one of the following: for each entity subset of the entity set, respectively generating rewritten content of the segment with each entity in the entity subset as the focus as the synthetic corpus; or for each entity subset of the entity set, generating analysis content about the interactions between the various entities in the entity subset in the context of the segment as the synthetic corpus.
[0010] In some embodiments, for each of the one or more segments, a summary and a title of the segment are also generated through a language model, and among them, at least one of the rewritten content and the analysis content is generated based on at least one of the summary and the title.
[0011] In some embodiments, the number of entities included in each entity subset is two or three.
[0012] In some embodiments, the entities include one or more of the plant species, symptoms, diseases, and causes of diseases.
[0013] In some embodiments, the visual model is a Contrastive Language–Image Pre-training (CLIP) model, and the CLIP model is trained through contrastive learning using an image-text pair dataset.
[0014] In some embodiments, the CLIP model is first trained alone through contrastive learning using an image-text pair dataset, and then jointly trained with a large language model that has been pre-trained and continuously pre-trained using multi-modal data.
[0015] In some embodiments, the image-text pair dataset includes plant-domain image-text pairs and general-domain image-text pairs, and the number of plant-domain image-text pairs is in a preset ratio to the number of general-domain image-text pairs.
[0016] In some embodiments, the obtained question text includes question text for identifying the type of plant in a plant image, the multimodal data includes the plant image, a question asking about the type of plant in the plant image, and an answer indicating the type of plant in the plant image, and the plant-domain image-text pair includes at least one of a pair of a plant image and a plant Latin name, and a pair of a plant image and a set of plant feature labels.
[0017] In some embodiments, the obtained question text includes question text for identifying the disease of the plant in a plant image, the multimodal data includes the plant image, a question asking about the disease of the plant in the plant image, and an answer indicating the disease of the plant in the plant image, and the plant-domain image-text pair includes at least one of a pair of a plant image and a plant disease name, and a pair of a plant image and a set of disease feature labels.
[0018] In some embodiments, the answer text includes the disease of the plant, and the answer text further includes one or more of the cause of the disease, the treatment method of the disease, and the maintenance suggestions for the plant.
[0019] In some embodiments, the plant identification method includes: generating a maintenance plan according to the answer text, the maintenance plan including one or more pairs, and each pair in the one or more pairs includes one or more maintenance tasks and an identifier of a maintenance device for performing the one or more maintenance tasks; outputting the maintenance plan.
[0020] In some embodiments, the plant identification method includes: controlling the corresponding maintenance device according to the identifier of the maintenance device in each pair of the one or more pairs in the maintenance plan to complete the one or more maintenance tasks in the pair.
[0021] According to a second aspect of the present disclosure, there is provided an electronic device, including: a processor; and a memory storing computer-executable instructions, which when executed by the processor cause the processor to execute the plant identification method according to any one of the embodiments of the first aspect of the present disclosure.
[0022] According to a third aspect of the present disclosure, there is provided a non-transitory storage medium storing computer-executable instructions, which when executed by a computer cause the computer to execute the plant identification method according to any one of the embodiments of the first aspect of the present disclosure.
[0023] According to a fourth aspect of the present disclosure, there is provided a computer program product including instructions which, when executed by a processor, implement the plant recognition method according to any embodiment of the first aspect of the present disclosure.
[0024] According to a fifth aspect of the present disclosure, there is provided a maintenance system, including: an electronic device including a processor and a memory coupled to the processor and storing instructions which, when executed by the processor, cause the processor to: obtain a plant image and question text regarding the recognition of the plant in the plant image, input the plant image and the question text into a plant knowledge Q&A model, the plant knowledge Q&A model including a vision model and a large language model, wherein the vision model is configured to receive the plant image to extract image features of the plant image, the large language model is configured to receive the image features and the question text to recognize the plant in the plant image, the large language model is further pre-trained with a synthetic corpus generated based on a plant domain corpus on the basis of being pre-trained, the plant knowledge Q&A model is trained with multimodal data, the multimodal data including the plant image, the question regarding the recognition of the plant in the plant image, and the answer to the question, generate answer text regarding the recognition of the plant in the plant image provided by the plant knowledge Q&A model, generate a maintenance plan based on the answer text, the maintenance plan including one or more pairs, each pair in the one or more pairs including one or more maintenance tasks and an identifier of a maintenance device for performing the one or more maintenance tasks, and transmit a command to a corresponding maintenance device according to the identifier of the maintenance device in each pair in the one or more pairs in the maintenance plan to control the corresponding maintenance device to complete the one or more maintenance tasks in the pair; and a maintenance device communicatively coupled to the electronic device, the maintenance device being configured to perform the maintenance tasks in response to receiving a command from the electronic device.
[0025] In some embodiments, the maintenance system includes: a camera communicatively coupled to the electronic device, the camera being configured to capture a plant image and transmit the captured plant image to the electronic device, wherein the instructions include instructions which, when executed by the processor, cause the processor to perform the following operations: output answer text based on the plant image received from the camera and default question text regarding the recognition of the plant in the plant image.
[0026] In some embodiments, the default question text includes question text regarding the recognition of the species and / or diseases of the plant in the plant image. Description of the Drawings
[0027] The foregoing and other features and advantages of the present disclosure will become apparent from the following description of embodiments of the present disclosure shown in conjunction with the accompanying drawings. The accompanying drawings are incorporated herein and form a part of the specification, further for explaining the principles of the present disclosure and enabling those skilled in the art to make and use the present disclosure. Among them:
[0028] Figure 1 shows a flowchart of a plant recognition method according to some embodiments of the present disclosure;
[0029] Figure 2 shows a schematic block diagram of a plant knowledge Q&A model according to some embodiments of the present disclosure;
[0030] Figure 3 shows a schematic diagram of an exemplary user interface in which the plant recognition method according to some embodiments of the present disclosure is applied;
[0031] Figure 4 shows a flowchart of operations for generating a synthetic corpus according to some embodiments of the present disclosure;
[0032] Figure 5 shows a non-limiting example process of generating a synthetic corpus based on a plant domain corpus according to some embodiments of the present disclosure;
[0033] Figure 6 shows a schematic block diagram of an electronic device according to some embodiments of the present disclosure;
[0034] Figure 7 shows a schematic block diagram of a computer system on which embodiments of the present disclosure can be implemented;
[0035] Figure 8 shows a schematic block diagram of a maintenance system according to some embodiments of the present disclosure.
[0036] Note that in the embodiments described below, sometimes the same reference numerals are used commonly between different drawings to denote the same parts or parts having the same functions, and their repeated description is omitted. In some cases, similar reference numerals and letters are used to denote similar items. Therefore, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.
[0037] For ease of understanding, the positions, sizes, ranges, etc. of the respective structures shown in the drawings and the like sometimes do not represent the actual positions, sizes, ranges, etc. Therefore, the present disclosure is not limited to the positions, sizes, ranges, etc. disclosed in the drawings and the like. Detailed Embodiments
[0038] Hereinafter, various exemplary embodiments of the present disclosure will be described in detail with reference to the drawings. It should be noted that: Unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.
[0039] The following description of at least one exemplary embodiment is in fact merely illustrative and is in no way intended to limit the present disclosure and its application or use. That is, the structures and methods herein are shown in an exemplary manner to illustrate different embodiments of the structures and methods in the present disclosure. However, those skilled in the art will appreciate that they merely illustrate exemplary ways of the present disclosure that can be implemented, rather than exhaustive ways. In addition, the drawings need not be drawn to scale, and some features may be enlarged to illustrate the details of specific components.
[0040] In addition, technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered as part of the specification.
[0041] In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limiting. Therefore, other examples of the exemplary embodiments may have different values.
[0042] The Multimodal Large Language Model (MLLM) is a large neural network model that combines multiple modal data such as text and images. This model can not only process text information, but also other types of data such as images and audio. By learning the correlation between multiple modal data at the same time, the Multimodal Large Language Model can understand and express information more comprehensively.
[0043] In order to provide more professional understanding and suggestions on plant identification, plant maintenance, and diagnosis and treatment of pests and diseases, the multimodal large language model needs to have sufficient knowledge in the field of plants. Therefore, in order to strengthen the knowledge reserve of the multimodal large language model in the field of plants, the pre-trained multimodal large language model can be further pre-trained on the plant field corpus to improve its performance in the field of plants.
[0044] However, multimodal large language models are inefficient in acquiring knowledge from large, unstructured general corpora, and require a large number of text instances with different expressions (for example, about the growth characteristics of a certain type of plant) to learn a specific knowledge point. In the plant domain corpus, many text instances appear only once or very few times. Therefore, it is very challenging for multimodal large language models to continue pre-training on the plant domain corpus, which is smaller than the general corpus.
[0045] To this end, the present disclosure provides a plant recognition method, which utilizes a plant knowledge Q&A model that combines a large language model and a vision model, and can automatically process plant images and related questions for answering. The large language model in the plant knowledge Q&A model adopted by the present disclosure is further pre-trained using a synthetic corpus generated based on a plant domain corpus. Since the large language model is further pre-trained using a larger and more diverse synthetic corpus generated based on a plant domain corpus, the large language model performs better in the plant domain, thereby improving the performance of the plant knowledge Q&A model in the plant domain and enhancing the user experience.
[0046] The plant recognition method according to the present disclosure will be described in detail below with reference to the accompanying drawings. It should be understood that the actual plant recognition method may also include other additional steps, but to avoid obscuring the key points of the present disclosure, these other additional steps are not discussed herein and are not shown in the drawings.
[0047] Figure 1 The flowchart of a plant recognition method 100 (hereinafter simply referred to as method 100) according to some embodiments of the present disclosure is shown. As Figure 1 shown, method 100 includes step S102 to step S106.
[0048] At step S102, a plant image and question text for recognizing the plant in the plant image are obtained.
[0049] At step S104, the plant image and the question text are input into a plant knowledge Q&A model, which includes a vision model and a large language model. As a non-limiting implementation, the plant knowledge Q&A model can be a multimodal large language model.
[0050] At step S106, answer text provided by the plant knowledge Q&A model for recognizing the plant in the plant image is output.
[0051] Specifically, recognizing the plant in the plant image may include, for example, recognizing the species and / or diseases of the plant in the plant image. Correspondingly, the obtained question text may include question text regarding recognizing the species and / or diseases of the plant in the plant image. As a non-limiting example, a user interface 300 as Figure 3 shown may be provided. The user interface 300 includes a dialog box 310, a text input box 320, and an image addition button 330. The image addition button 330 is used to receive the plant image, and the text input box 320 is used to receive the question text.
[0052] The plant knowledge Q&A model adopted in method 100 can be trained with multimodal data, which includes plant images, questions about identifying the plants in the plant images, and answers to the questions. To identify the species of plants in a plant image, the multimodal data for training the plant knowledge Q&A model can include a plant image, a question asking about the species of the plant in the plant image, and an answer indicating the species of the plant in the plant image, such as {<plant image.jpg>, "What is this plant?", "This plant is Ficus hispida"}. To identify the diseases of plants in a plant image, the multimodal data for training the plant knowledge Q&A model can include a plant image, a question asking about the diseases of the plant in the plant image, and an answer indicating the diseases of the plant in the plant image, such as {<plant image.jpg>, "What's wrong with this plant?", "This plant has leaf mold disease"}.
[0053] In the multimodal data for training the plant knowledge Q&A model, the plant image and the question can be used as training samples, while the answer to the question can be used as the annotation of the sample. During the training process, the plant knowledge Q&A model will learn how to understand the relationship between images and texts, and how to generate relevant answers. By jointly training on different modal data such as images and texts, the plant knowledge Q&A model can learn the corresponding relationships between different modalities, so as to achieve cross-modal information expression and reasoning capabilities. Through the training data specific to the plant recognition field, the plant recognition ability is injected into the model.
[0054] Figure 2 Fig. shows a schematic block diagram of a plant knowledge Q&A model 200 according to some embodiments of the present disclosure. As Figure 2 shown, the plant knowledge Q&A model 200 includes a visual model 210 and a large language model 230. The visual model 210 is configured to receive a plant image to extract image features of the plant image. The large language model 230 is configured to receive the image features and question text to identify the plants in the plant image, so as to output answer text about identifying the plants in the plant image.
[0055] The large language model 230 is further pre-trained with a synthetic corpus generated based on a plant domain corpus on the basis of being pre-trained. Compared with directly further pre-training the large language model 230 on a small-scale plant domain corpus, further pre-training the large language model 230 on a larger-scale and more diverse synthetic corpus can enable the large language model 230 to more efficiently acquire knowledge from the synthetic corpus and learn the associations between knowledge points, so as to better improve the performance of the large language model 230 in the plant field.
[0056] A synthetic corpus can be generated based on a corpus in the plant domain. In some examples, tens of thousands of books in the plant domain can be obtained and parsed to obtain text content, thereby forming a corpus in the plant domain. Refer to Figure 4 , which shows a flowchart of operation 400 for generating a synthetic corpus according to some embodiments of the present disclosure. As Figure 4 shown, operation 400 may include step S402 to step S406.
[0057] At step S402, the plant domain corpus in the plant domain corpus is segmented into one or more segments according to semantics. In some examples, segmenting the plant domain corpus according to semantics can be completed by a language model for subsequent extraction of entities and entity relationships, or by other language models, or by experts.
[0058] At step S404, for each of the one or more segments obtained by segmentation, entities in the segment are extracted through a language model to obtain an entity set, and the mutual relationships between entities in each entity subset of the entity set are extracted based on the segment. For example, the language model can be a Generative Pre-trained Transformer (GPT), etc. As a non-limiting implementation, the entity may include one or more of the plant species, symptoms, diseases, and causes.
[0059] In some examples, assuming the entity set includes N entities and each entity subset includes n entities, then the entity set may include C N n types of entity subsets. For example, when the number of entities included in each entity subset is 2 (such an entity subset can be called an entity pair), the entity set may include C N 2 , that is, (n - 1) n / 2 types of entity subsets. In other examples, the number of entities included in each entity subset may be 3 (such an entity subset can be called an entity triple). Of course, the number of entities included in each entity subset can be flexibly set according to actual needs. Thus, by iteratively extracting the mutual relationships between entities in each entity subset in the entity set multiple times, comprehensive coverage of different entity relationships is ensured, thereby reducing the repetition degree of the subsequent generated synthetic corpus while maintaining the diversity of the descriptions of text instances in the synthetic corpus.
[0060] At step S406, a synthetic corpus is generated from the paragraph by a language model based on the extracted entities and / or the relationships between the entities. In some embodiments, generating a synthetic corpus from the paragraph based on the extracted entities and / or the relationships between the entities includes at least one of the following: for each entity subset of the entity set, generating rewritten content of the paragraph with each entity in the entity subset as the focus as the synthetic corpus; or for each entity subset of the entity set, generating analytical content regarding the interactions of the entities in the entity subset in the context of the paragraph as the synthetic corpus.
[0061] In some embodiments, operation 400 may further include: for each of one or more paragraphs, also generating a summary and a title of the paragraph by a language model, wherein at least one of the rewritten content and the analytical content is generated based on at least one of the summary and the title. In some embodiments, at least one of the rewritten content and the analytical content is performed based on the summary of the paragraph rather than the paragraph itself. In some embodiments, at least one of the rewritten content and the analytical content is performed based on the summary of the paragraph and the paragraph itself. In some embodiments, it may be required that at least one of the rewritten content and the analytical content includes or reflects the title of the paragraph. In some embodiments, generating the rewritten content of the paragraph with each entity in the entity subset as the focus may include generating the rewritten content of the paragraph based on the interaction between the entity and the title of the paragraph.
[0062] Thus, the synthetic corpus generated based on operation 400 may include a large number of text instances in different expression forms generated on the basis of the original plant domain corpus, facilitating the training of large language models.
[0063] Reference Figure 5 , which shows a non-limiting example process of generating a synthetic corpus based on a plant domain corpus according to some embodiments of the present disclosure. Figure 5 The first box in shows a paragraph of a plant domain corpus "raw_doc", and then a language model is used to extract the entities in the paragraph to obtain an entity set "entities" and generate a summary "summary" and a title "title" of the paragraph. Then, taking the entity subset {"A. rhodocyanea", "A. fasciata"} as an example, the rewritten content 1, 2 are generated with "A. rhodocyanea", "A. fasciata" in the entity subset as the focus respectively, and the analytical content 3 is generated based on the relationship between them, so that the synthetic corpus includes the rewritten content 1, 2 and the analytical content 3.
[0064] Specifically, the rewritten content 1 corresponding to “A. rhodocyanea” further includes “Discussion of Care and Characteristics of Bromeliads and Chinese Evergreens in relation to “A. rhodocyanea” as a sub - title or dividing line of the rewritten content 1, which includes “A. rhodocyanea” and the title “Care and Characteristics of Bromeliads and Chinese Evergreens”. The rewritten content 2 corresponding to “A. fasciata” further includes “Discussion of Care and Characteristics of Bromeliads and Chinese Evergreens in relation to A. fasciata” as a sub - title or dividing line of the rewritten content 2, which includes “A. fasciata” and the above - mentioned title. Also, the analysis content 3 corresponding to the relationship between “A. rhodocyanea” and “A. fasciata” further includes “Discussion of Interaction between A. rhodocyanea and A. fasciata in context of Care and Characteristics of Bromeliads and Chinese Evergreens” as a sub - title or dividing line of the analysis content 3, which includes “A. rhodocyanea”, “A. fasciata” and the above - mentioned title. Using sub - titles or dividing lines to segment a large number of generated text examples can make the synthetic corpus more organized.
[0065] In some embodiments, the synthetic corpus may include a mixture of plant - domain corpus and synthetic corpus, so as to be more focused on improving the plant recognition performance of the plant knowledge Q&A model 200. That is to say, the plant - domain corpus itself used to generate the synthetic corpus can be incorporated into the synthetic corpus.
[0066] In some embodiments, the synthetic corpus may include a mixture of general - domain corpus and synthetic corpus, so as to enhance the generalization of the plant knowledge Q&A model 200 while improving the plant recognition performance of the plant knowledge Q&A model 200.
[0067] In some further embodiments, the number of general domain corpora and the number of synthetic corpora are in a preset ratio. In some examples, the preset ratio can be determined by the following operations: using general domain corpora and synthetic corpora mixed in different ratios as training data to train the large language model, and using verification data to determine the verification loss of the trained large language model; and selecting a ratio from these different ratios based on the verification loss as the preset ratio. For example, the ratio corresponding to the minimum verification loss can be selected as the preset ratio. In addition, in some examples, the verification data can be general domain corpora or synthetic corpora that are not used as training data.
[0068] Since the mixing ratio of training data has a significant impact on the final performance of the large language model, by optimizing the mixing ratio of the number of general corpora to the number of synthetic corpora based on small-scale training in advance and using the optimized mixing ratio for subsequent large-scale training, it is possible to reduce training costs while ensuring excellent training results.
[0069] In some embodiments, the synthetic corpus may include a mixture of plant domain corpus, synthetic corpus, and general domain corpus.
[0070] The visual capability of the large language model 230 may mainly depend on the visual model 210, especially the image features extracted by the visual model 210. Therefore, the expressive power of the image features of the visual model 210 will directly affect the performance of the large language model on the plant recognition visual task.
[0071] The visual model 210 can be based on various suitable neural network architectures. In some embodiments, the visual model 210 can be a CLIP model. The CLIP model is a deep learning model that aims to achieve interaction between natural language processing and computer vision.
[0072] In order to improve the performance of the visual model 210 in the field of plant identification, in some embodiments, the CLIP model is trained by contrastive learning using an image-text pair dataset. For example, the CLIP model can first be trained by contrastive learning using an image-text pair dataset alone, and then jointly trained with a large language model that has been pre-trained and continues to be pre-trained using multimodal data. This is conducive to improving the overall performance of the model. Of course, the CLIP model can also be pre-trained alone first, and then when training the plant knowledge question-answering model using multimodal data, the parameters of the CLIP model are fixed and only the parameters of the large language model are updated, which can speed up the training and reduce the computing and storage resources consumed by the training.
[0073] It can be understood that the same applies to other multimodal data. For example, if audio data also needs to be processed, the plant knowledge Q&A model can be provided by adopting a large language model for processing audio features, image features, and text features either alone or in combination with one or both of a visual model for extracting image features and an auditory model for extracting audio features.
[0074] In some embodiments, the image-text pair dataset includes plant-domain image-text pairs and general-domain image-text pairs. In some embodiments, the number of plant-domain image-text pairs is in a preset ratio to the number of general-domain image-text pairs. In some examples, the ratio of the number of plant-domain image-text pairs to the number of general-domain image-text pairs is 1:1. Thus, on the basis of retaining the general visual encoding ability of the CLIP model, the observation ability of the CLIP model for plant details is improved, so that relevant information about plants can be captured from visual signals (such as images), enabling the subsequent large language model to have the ability to observe and understand both the overall vision and details.
[0075] In addition, as a non-limiting implementation, for identifying the species of plants in plant images, the plant-domain image-text pairs can include at least one of the pair of a plant image and the Latin name of the plant, and the pair of a plant image and a set of plant feature tags; for identifying the diseases of plants in plant images, the plant-domain image-text pairs can include at least one of the pair of a plant image and the name of the plant disease, and the pair of a plant image and a set of disease feature tags.
[0076] The CLIP model maps images and texts to the same embedding space through contrastive learning, making the embedding vectors of corresponding image-text pairs closer in this space, while the embedding vectors of irrelevant image-text pairs are farther apart in this space. Through training, the CLIP model can learn a common embedding space, enabling effective semantic alignment of images and texts, and thus extracting cross-modal visual representations for images.
[0077] The training process of the CLIP model based on contrastive learning of the image-text pair dataset is as follows, for example.
[0078] For the image-text pairs input into the CLIP model, it is first necessary to convert these different types of data into feature representations that are compatible within the model for subsequent processing of these different types of data. This is typically achieved by encoders corresponding to different modalities. Specifically, the image encoder of the CLIP model (such as a convolutional neural network (CNN), a recurrent neural network (RNN), etc.) extracts image features from the image, and the text encoder of the CLIP model (such as the Word2Vec model, the bag of words model (BoW), etc.) extracts text features from the text.
[0079] Suppose a training batch contains N image-text pairs. Then the image encoder will extract N image features, and the text encoder will also extract N text features. Combine the N image features and N text features pairwise to obtain N 2 samples. When the image encoder extracts image features and the text encoder extracts text features, the image features and text features can be normalized respectively. Among them, for each image feature, there is 1 positive sample and (N - 1) negative samples; for each text feature, there is 1 positive sample and (N - 1) negative samples. Generally speaking, there are N positive samples and (N 2 - N) negative samples.
[0080] The training objective of the CLIP model can be to maximize the similarity of the N positive samples (for example, the cosine similarity between the text feature and the image feature can be directly calculated) and / or minimize the similarity of the (N 2 - N) negative samples. For example, is the i-th normalized image feature, is the j-th normalized text feature, the cosine similarity between the image feature and the text feature is , that is, is and dot product. Thus, the similarity between the N image features and N text features.
[0081] A temperature parameter greater than zero can also be introduced to control the smoothness of the softmax (normalized exponential function) distribution. A smaller temperature parameter will make the softmax distribution sharper, and the CLIP model will pay more attention to the matching positive samples. And the temperature parameter is usually a learnable parameter. Thus, the loss function can be used to calculate the loss of the current training batch. Among them, the exp function is the exponential function with the real number e as the base, is a temperature parameter, and represent a similarity score or a compatibility score. Specifically, represents the similarity score between sample i and its positive sample and can be used to represent the similarity between samples of the same class, represents the similarity score between sample i and another sample j and can be used to calculate the similarity between sample i and all other samples. These similarity scores calculate a probability distribution through an exponential function and a normalization step (such as softmax), so that the similarity between different samples can be compared.
[0082] During the training process, the CLIP model can also calculate gradients through backpropagation and use an optimizer (e.g., the Adam optimizer) to update the parameters of the image encoder and the text encoder as well as the temperature parameter.
[0083] Through contrastive learning, the expression ability of the visual features of the CLIP model can be strengthened, thereby improving the performance of the large language model in visual tasks and ultimately improving the plant recognition ability of the entire plant knowledge Q&A model.
[0084] In some embodiments, the answer text includes the diseases of the plant, and the answer text further includes one or more of the cause of the disease, the treatment method of the disease, and the maintenance suggestions for the plant.
[0085] After obtaining the answer text, relevant tasks can be generated and displayed to the user according to the content in the answer text. In the case of a device connected with an executable task, the corresponding device can also be automatically controlled to execute the task.
[0086] Specifically, in some embodiments, method 100 may include: generating a maintenance plan according to the answer text, the maintenance plan including one or more pairs, each pair in the one or more pairs including one or more maintenance tasks and the identifier of the maintenance device for executing the one or more maintenance tasks; outputting the maintenance plan.
[0087] The maintenance plan may include a daily maintenance plan, a treatment maintenance plan, etc. or a combination thereof. For example, when the answer text does not involve the diseases of the plant, the daily maintenance plan of the plant can be output; when the answer text involves the diseases of the plant, the treatment maintenance plan of the plant can be output, and optionally the matching daily maintenance plan can also be output.
[0088] For illustrative purposes, a non-limiting example application of method 100 may include the diagnosis of diseases in tomato plants maintained by a user. In this example, the leaves and fruits of the tomato plants in the plant images input by the user start to rot. To further assist the user in maintaining the tomatoes, maintenance tasks can be generated based on the answer text. After generating the maintenance tasks, since the identification of the maintenance device communicatively coupled to the user terminal is also stored in the relevant database, the identification of the maintenance device associated with the maintenance task can also be determined from the relevant database based on the maintenance task, thereby generating a maintenance plan. After obtaining the maintenance plan, the maintenance plan can be displayed on the user interface. Thus, the user can know what kind of maintenance device (such as an irrigation device, a fertilization device, a pruning device, a drug administration device, a light control device, a temperature control device, a humidity control device, etc. or a combination thereof) should be used and what kind of maintenance tasks should be implemented to maintain the tomatoes. The maintenance tasks can include, for example, watering, spraying, fertilizing, pruning, weeding, repotting, sunlight exposure, shading, temperature adjustment, humidity adjustment, application of pesticides, and application of fungicides. Specifically, the maintenance tasks can also include various parameters of the tasks, such as the time, interval, and amount of watering, the dose, time, and interval of fertilization, the position of pruning, the dose and position of pesticide spraying, etc.
[0089] In some embodiments, method 100 may include: controlling the corresponding maintenance device according to the identification of the maintenance device in each of one or more pairs in the maintenance plan to complete one or more maintenance tasks in the pair.
[0090] Since the maintenance device usually has a communication function, commands can be transmitted to the maintenance device (e.g., via the Bluetooth protocol, the Zigbee protocol, etc.). For example, in the foregoing example, the maintenance plan may include two pairs: {fruit pruning task and leaf pruning task, identification of the pruning device} and {pesticide spraying task, identification of the drug administration device}. Commands instructing to perform the pruning task and the spraying task are sent to the corresponding pruning device and spraying device respectively according to the identification, so as to control the pruning device and the spraying device to automatically complete the pruning task and the spraying task, thereby reducing the user's maintenance burden and improving the maintenance efficiency.
[0091] In addition to automatically executing the maintenance plan, in some embodiments, after displaying the maintenance plan, the user may be further asked whether to confirm the execution of the maintenance plan, and after the user confirms the execution of the maintenance plan, the corresponding maintenance device is controlled according to the maintenance plan to complete the corresponding maintenance task.
[0092] The present disclosure also provides an electronic device in another aspect. Refer to Figure 6 , which shows a schematic block diagram of an electronic device 600 according to some embodiments of the present disclosure. As Figure 6As shown, the electronic device 600 includes a processor 602 and a memory 604 that stores computer-executable instructions. When the computer-executable instructions are executed by the processor 602, the processor 602 is caused to execute the plant recognition method according to any of the foregoing embodiments of the present disclosure. The processor 602 may be, for example, a central processing unit (CPU) of the electronic device 600. The processor 602 may be any type of general-purpose processor or may be a processor specifically designed for plant recognition, such as an application-specific integrated circuit (“ASIC”). The memory 604 may be coupled to the processor 602 and may include various computer-readable media accessible by the processor 602. In various embodiments, the memory 604 described herein may include volatile and non-volatile media, removable and non-removable media. For example, the memory 604 may include any combination of the following: random access memory (“RAM”), dynamic RAM (“DRAM”), static RAM (“SRAM”), read-only memory (“ROM”), flash memory, cache memory, and / or any other type of non-transitory computer-readable media. The memory 604 may store instructions that, when executed by the processor 602, cause the processor 602 to execute the method 100 according to any of the foregoing embodiments of the present disclosure.
[0093] In some embodiments, the electronic device 600 may be implemented as a smartphone, a smart tablet, a smart camera, a computer, etc.
[0094] The electronic device 600 is configured to execute the method 100 according to any of the foregoing embodiments. Therefore, reference may be made to the descriptions of the various embodiments of the method 100 above, and details are not repeated herein.
[0095] The present disclosure also provides a non-transitory storage medium having stored thereon computer-executable instructions. When the computer-executable instructions are executed by a computer, the computer is caused to execute the plant recognition method according to any of the foregoing embodiments of the present disclosure.
[0096] The present disclosure also provides a computer program product. The computer program product may include instructions that, when executed by a processor, may implement the plant recognition method according to any of the foregoing embodiments of the present disclosure. The instructions may be any set of instructions that are directly executable by one or more processors, such as machine code, or any set of instructions that are indirectly executable, such as a script. The instructions may be stored in a target code format for direct processing by one or more processors or stored in any other computer language, including a script or a collection of independent source code modules that are interpreted on demand or pre-compiled.
[0097] Figure 7FIG. 0 shows a schematic block diagram of a computer system 700 on which embodiments of the present disclosure may be implemented. The computer system 700 includes a bus 702 or other communication mechanism for transferring information, and a processing device 704 coupled to the bus 702 for processing information. The computer system 700 also includes a memory 706 coupled to the bus 702 for storing instructions to be executed by the processing device 704. The memory 706 may be a random access memory (RAM) or other dynamic storage device. The memory 706 may also be used to store temporary variables or other intermediate information during the execution of instructions to be executed by the processing device 704. The computer system 700 also includes a read-only memory (ROM) 708 or other static storage device coupled to the bus 702 for storing static information and instructions for the processing device 704. A storage device 710, such as a magnetic disk or optical disk, is provided and coupled to the bus 702 for storing information and instructions. The computer system 700 may be coupled via the bus 702 to an output device 712 for providing output to a user, such as, but not limited to, a display (such as a cathode ray tube (CRT) or liquid crystal display (LCD)), a speaker, etc. Input devices 714, such as a keyboard, a mouse, a microphone, etc., are coupled to the bus 702 for transmitting information and command selections to the processing device 704. The computer system 700 may execute embodiments of the present disclosure. Consistent with certain implementations of the present disclosure, results are provided by the computer system 700 in response to execution of one or more sequences of one or more instructions contained in the memory 706 by the processing device 704. Such instructions may be read into the memory 706 from another computer-readable medium, such as the storage device 710. Execution of the sequences of instructions contained in the memory 706 causes the processing device 704 to perform the methods described herein. Alternatively, the present teachings may be implemented using hardwired circuitry in place of, or in combination with, software instructions. Thus, implementations of the present disclosure are not limited to any specific combination of hardware circuitry and software. In various embodiments, the computer system 700 may be connected across a network to one or more other computer systems, such as the computer system 700, via a network interface 716 to form a networked system. The network may include a private network or a public network such as the Internet. In the networked system, one or more computer systems may store data and supply the data to other computer systems. As used herein, the term "computer-readable medium" refers to any medium that participates in providing instructions to the processing device 704 for execution. Such a medium may take many forms, including but not limited to non-volatile media, volatile media, and transmission media. Non-volatile media includes, for example, optical or magnetic disks such as the storage device 710. Volatile media includes dynamic memory such as the memory 706. Transmission media includes coaxial cables, copper wire, and fiber optics, including the wiring that includes the bus 702.Common forms of computer-readable media or computer program products include, for example, floppy disks, flexible disks, hard disks, magnetic tapes, or any other magnetic media, CD-ROMs, digital video discs (DVDs), Blu-ray discs, any other optical media, thumb drives, memory cards, RAM, PROM, and EPROM, flash EPROM, any other memory chip or cartridge, or any other tangible medium from which a computer can read. Various forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to the processing device 704 for execution. For example, the instructions may initially be carried on a disk of a remote computer. The remote computer may load the instructions into its dynamic memory and send the instructions over a telephone line using a modem. A modem local to the computer system 700 may receive the data over the telephone line and convert the data into an infrared signal using an infrared transmitter. An infrared detector coupled to the bus 702 may receive the data carried in the infrared signal and place the data on the bus 702. The bus 702 carries the data to the memory 706, and the processing device 704 retrieves the instructions from the memory 706 and executes the instructions. For example, the instructions received by the memory 706 may be stored on the storage device 710 before or after being executed by the processing device 704.
[0098] Figure 8 FIG. shows a schematic block diagram of a maintenance system 800 according to some embodiments of the present disclosure. The maintenance system 800 includes an electronic device 810, which includes a processor 812 and a memory 814 coupled to the processor 812 and storing instructions. The electronic device 810 may, for example, but is not limited to, take the form of the aforementioned electronic device 600 or computer system 700, etc., and may be implemented as, for example, but not limited to, a smart phone, a smart tablet, a smart camera, a computer, etc. The memory 814 may store instructions that, when executed by the processor 812, cause the processor 812 to execute the plant recognition method according to any one of the foregoing embodiments of the present disclosure. The maintenance system 800 further includes one or more maintenance devices (e.g., 820 1 、820 2 、……、820 n ) communicatively coupled to the electronic device 810 and configured to perform maintenance tasks in response to receiving a command from the electronic device 810.
[0099] Specifically, in some embodiments, the instructions stored in the memory 814, when executed by the processor 812, may cause the processor 812 to: obtain a plant image and question text regarding the identification of the plant in the plant image; input the plant image and the question text into a plant knowledge Q&A model, where the plant knowledge Q&A model includes a visual model and a large language model. The visual model is configured to receive the plant image to extract image features of the plant image, and the large language model is configured to receive the image features and the question text to identify the plant in the plant image. The large language model is further pre-trained with a synthetic corpus generated based on a plant domain corpus on the basis of being pre-trained. The plant knowledge Q&A model is trained with multimodal data, and the multimodal data includes plant images, questions regarding the identification of the plants in the plant images, and answers to the questions; generate answer text regarding the identification of the plant in the plant image provided by the plant knowledge Q&A model; generate a maintenance plan according to the answer text, where the maintenance plan includes one or more pairs, and each pair in the one or more pairs includes one or more maintenance tasks and an identifier of a maintenance device for performing the one or more maintenance tasks; transmit commands to the corresponding maintenance devices (e.g., 820 1 、820 2 、……、820 n ) according to the identifiers of the maintenance devices in each pair in the one or more pairs in the maintenance plan to control the corresponding maintenance devices (e.g., 820 1 、820 2 、……、820 n ) to complete one or more maintenance tasks in the pair.
[0100] In some embodiments, the electronic device 810 includes a user interface (not shown). For example, the answer text can be displayed on the user interface, and / or the maintenance plan can be displayed on the user interface.
[0101] In some embodiments, the maintenance system 800 may include a camera 830 communicatively coupled to the electronic device 810. The camera 830 can be any suitable imaging device for monitoring the target plant. The target plant can be located in the field of view of the camera 830. The camera 830 is configured to capture a plant image and transmit the captured plant image to the electronic device 810. Accordingly, the instructions stored in the memory 814 may include instructions that, when executed by the processor 812, cause the processor 812 to perform the following operations: output answer text based on the plant image received from the camera 830 and default question text regarding the identification of the plant in the plant image.
[0102] In some examples, the default question text may include question text regarding the identification of the type and / or disease of a plant in a plant image. For example, the default question text may be stored in a relevant database, such that in response to a plant image captured by a camera, the default question text can be extracted from the database. Information such as the name of the plant, the maintenance information of the plant, the diseases of the plant, and the corresponding treatment and prevention methods may also be stored in the database. This information about the plant can be stored in the database in association with the image features of the plant. Information can be extracted based on the degree of matching between the image features of the identified plant image and the image features of the plants stored in the database, for example when the degree of matching falls within a preset range. As a non-limiting example, the cosine similarity between a first vector representing the image features of the identified plant image and a second vector representing the image features of the plant stored in the database can be calculated, and when the calculated cosine similarity exceeds a preset threshold, it is considered that the image features of the identified plant image match the image features of the plant stored in the database, and the information stored in the database in association with the matched image features is extracted as the default question text.
[0103] In some examples, the plant image received from the camera 830 can be input into the aforementioned plant knowledge Q&A model for processing. For example, close-up images or videos of one or more characteristic parts of the target plant may be required. In this case, it is not necessary for the user to input these close-up images or videos, but the camera 830 can automatically obtain these close-up images or videos. Or, when the plant image from the user cannot be recognized due to various reasons such as insufficient clarity, it is not necessary for the user to re-enter the plant image, but the camera 830 can automatically obtain the plant image. That is, the plant image captured by the camera 830 can be used to assist in generating the answer text.
[0104] Thus, the image captured by the camera 830 can be used to independently generate the answer text and automatically execute the maintenance plan, so as to achieve full-automatic monitoring and maintenance of the target plant.
[0105] Various embodiments of the maintenance system 800 can be similarly referred to any of the foregoing embodiments of the present disclosure, and will not be elaborated herein.
[0106] One or more exemplary embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.
[0107] The systems, devices, modules or units illustrated in the above embodiments may be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a server system. Of course, with the development of future computer technologies, the computers that implement the functions of the above embodiments may include, for example, personal computers, laptop computers, in-vehicle human-machine interaction devices, cellular phones, camera phones, smart phones, personal digital assistants, media players, game consoles, tablet computers, wearable devices, or any combination thereof.
[0108] The term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, product or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, product or device. Without further limitation, there is no exclusion of additional identical or equivalent elements in the process, method, product or device comprising the said elements. For example, when terms such as "first" and "second" are used to denote names, they do not denote any particular order.
[0109] For convenience of description, the above device is described by dividing it into various modules according to functions. Of course, when implementing one or more embodiments of the present disclosure, the functions of each module may be implemented in the same or multiple software and / or hardware, or the modules implementing the same function may be realized by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the couplings, direct couplings or communication connections shown or discussed among each other may be through some interfaces. The indirect couplings or communication connections of the devices or units may be in electrical, mechanical or other forms.
[0110] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be realized by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate a device for realizing the functions specified in one or more processes of the flowchart and / or one or more blocks of the block diagram.
[0111] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the function specified in one or more blocks of a flowchart and / or one or more blocks of a block diagram. These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the function specified in one or more blocks of a flowchart and / or one or more blocks of a block diagram.
[0112] Those skilled in the art will appreciate that one or more embodiments of the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0113] One or more embodiments of the present disclosure may be described in the general context of computer-executable instructions, such as program modules, being executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. One or more embodiments of the present disclosure may also be practiced in a distributed computing environment where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media including storage devices.
[0114] For the same or similar parts among the various embodiments of the present disclosure, reference may be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the apparatus embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference may be made to the partial description of the method embodiments. In the description of the present disclosure, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", "exemplary", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In the present disclosure, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in the present disclosure and the features of the different embodiments or examples.
[0115] In addition, as used in the present disclosure, words such as "here", "above", "below", "hereinafter", "above-mentioned", and words with similar meanings shall refer to the whole of the present disclosure rather than any specific part of the present disclosure. Moreover, unless otherwise clearly stated or otherwise understood in the context in which it is used, the conditional language used herein, such as "may", "might", "for example", "such as", etc., generally intends to indicate that certain embodiments include, while other embodiments do not include certain features, elements, and / or states. Therefore, such conditional language generally does not intend to imply that one or more embodiments require in any way the features, elements, and / or states, or whether they include these features, elements, and / or states or perform these features, elements, and / or states in any particular embodiment.
[0116] In addition, the embodiments of the present disclosure may also include the following examples:
[0117] Example 1. A plant recognition method, comprising:
[0118] Obtaining a plant image and question text for recognizing the plant in the plant image;
[0119] Input the plant image and the question text into a plant knowledge Q&A model, where the plant knowledge Q&A model includes a vision model and a large language model. Among them, the vision model is configured to receive the plant image to extract image features of the plant image, and the large language model is configured to receive the image features and the question text to identify the plant in the plant image. The large language model is further pre-trained on a synthetic corpus generated based on a plant domain corpus on the basis of being pre-trained. The plant knowledge Q&A model is trained with multi-modal data, and the multi-modal data includes plant images, questions about identifying the plants in the plant images, and answers to the questions; and
[0120] Output the answer text provided by the plant knowledge Q&A model for identifying the plant in the plant image.
[0121] Example 2. The method according to Example 1, wherein the synthetic corpus includes a mixture of plant domain corpus and synthetic corpus.
[0122] Example 3. The method according to Example 1 or 2, wherein the synthetic corpus includes a mixture of general domain corpus and synthetic corpus.
[0123] Example 4. The method according to Example 3, wherein the quantity of the general domain corpus and the quantity of the synthetic corpus are in a preset ratio, and the preset ratio is determined by the following operations:
[0124] Use the general domain corpus and the synthetic corpus mixed in different ratios as training data to train the large language model respectively, and use validation data to determine the validation loss of the trained large language model; and
[0125] Select the ratio from the different ratios as the preset ratio based on the validation loss.
[0126] Example 5. The method according to Example 1, wherein the synthetic corpus is constructed by the following operations:
[0127] Segment the plant domain corpus in the plant domain corpus into one or more segments according to semantics;
[0128] For each of the one or more segments, use a language model to extract entities in the segment to obtain an entity set, extract the mutual relationships between entities in each entity subset of the entity set based on the segment, and generate synthetic corpus from the segment based on the extracted entities and / or the mutual relationships between entities.
[0129] Example 6. The method according to Example 5, wherein generating synthetic corpus from the segment based on the extracted entities and / or the interrelationships between the entities includes at least one of the following:
[0130] For each entity subset of the entity set, generating rewritten content of the segment with each entity in the entity subset as the focus as the synthetic corpus; or
[0131] For each entity subset of the entity set, generating analysis content on the interactions between the entities in the entity subset in the context of the segment as the synthetic corpus.
[0132] Example 7. The method according to Example 6, wherein for each of the one or more segments, a summary and a title of the segment are further generated by the language model, and wherein at least one of the rewritten content and the analysis content is generated based on at least one of the summary and the title.
[0133] Example 8. The method according to Example 5, wherein the number of entities included in each entity subset is two or three.
[0134] Example 9. The method according to Example 5, wherein the entity includes one or more of the type of plant, symptom, disease, cause of disease.
[0135] Example 10. The method according to Example 1, wherein the visual model is a multi-modal contrastive learning pre-trained CLIP model, and the CLIP model is trained by contrastive learning using an image-text pair dataset.
[0136] Example 11. The method according to Example 1, wherein the CLIP model is first trained by contrastive learning alone using an image-text pair dataset, and then jointly trained with the large language model that has undergone the pre-training and the continued pre-training using the multi-modal data.
[0137] Example 12. The method according to Example 10 or 11, wherein the image-text pair dataset includes plant domain image-text pairs and general domain image-text pairs, and the number of the plant domain image-text pairs is in a preset ratio to the number of the general domain image-text pairs.
[0138] Example 13. The plant recognition method according to Example 12, wherein the obtained question text includes question text for recognizing the type of the plant in the plant image, the multimodal data includes the plant image, a question for asking about the type of the plant in the plant image, and an answer indicating the type of the plant in the plant image, and the plant domain image-text pair includes at least one of a pair of a plant image and a plant Latin name, and a pair of a plant image and a set of plant feature tags.
[0139] Example 14. The plant recognition method according to Example 12, wherein the obtained question text includes question text for recognizing the disease of the plant in the plant image, the multimodal data includes the plant image, a question for asking about the disease of the plant in the plant image, and an answer indicating the disease of the plant in the plant image, and the plant domain image-text pair includes at least one of a pair of a plant image and a plant disease name, and a pair of a plant image and a set of disease feature tags.
[0140] Example 15. The plant recognition method according to Example 14, wherein the answer text includes the disease of the plant, and the answer text further includes one or more of the cause of the disease, the treatment method of the disease, and the maintenance suggestions for the plant.
[0141] Example 16. The plant recognition method according to Example 1, including:
[0142] Generating a maintenance plan according to the answer text, the maintenance plan includes one or more pairs, and each pair in the one or more pairs includes one or more maintenance tasks and an identifier of a maintenance device for performing the one or more maintenance tasks;
[0143] Outputting the maintenance plan.
[0144] Example 17. The plant recognition method according to Example 16, including:
[0145] Controlling the corresponding maintenance device according to the identifier of the maintenance device in each pair of the one or more pairs in the maintenance plan to complete the one or more maintenance tasks in the pair.
[0146] Example 18. An electronic device, including:
[0147] A processor; and
[0148] A memory storing computer-executable instructions, which when executed by the processor cause the processor to execute the plant recognition method according to any one of Examples 1 to 17.
[0149] Example 19. A non-transitory storage medium storing computer-executable instructions, which, when executed by a computer, cause the computer to execute the plant recognition method according to any one of Examples 1 to 17.
[0150] Example 20. A computer program product comprising instructions which, when executed by a processor, implement the plant recognition method according to any one of Examples 1 to 17.
[0151] Example 21. A maintenance system, comprising:
[0152] An electronic device, the electronic device comprising a processor and a memory coupled to the processor and storing instructions which, when executed by the processor, cause the processor to:
[0153] Obtain a plant image and question text regarding the recognition of the plant in the plant image,
[0154] Input the plant image and the question text into a plant knowledge Q&A model, the plant knowledge Q&A model comprising a vision model and a large language model, wherein the vision model is configured to receive the plant image to extract image features of the plant image, the large language model is configured to receive the image features and the question text to recognize the plant in the plant image, the large language model is further pre-trained on a synthetic corpus generated based on a plant domain corpus on the basis of being pre-trained, the plant knowledge Q&A model is trained with multimodal data, the multimodal data comprising plant images, questions regarding the recognition of the plants in the plant images, and answers to the questions,
[0155] Generate answer text regarding the recognition of the plant in the plant image provided by the plant knowledge Q&A model,
[0156] Generate a maintenance plan according to the answer text, the maintenance plan comprising one or more pairs, each pair in the one or more pairs comprising one or more maintenance tasks and an identifier of a maintenance device for executing the one or more maintenance tasks, and
[0157] Transmit commands to corresponding maintenance devices according to the identifiers of the maintenance devices in each pair in the one or more pairs in the maintenance plan to control the corresponding maintenance devices to complete one or more maintenance tasks in the pair; and
[0158] Maintenance devices communicatively coupled to the electronic device, the maintenance devices being configured to execute maintenance tasks in response to commands received from the electronic device.
[0159] Example 22. The maintenance system according to Example 21, comprising:
[0160] A camera communicatively coupled to the electronic device, the camera being configured to capture an image of the plant and transmit the captured plant image to the electronic device,
[0161] wherein the instructions include instructions that, when executed by the processor, cause the processor to perform the following operations:
[0162] Output the answer text based on the plant image received from the camera and the default question text regarding the identification of the plant in the plant image.
[0163] Example 23. The maintenance system according to Example 22, wherein the default question text includes question text regarding the identification of the type and / or disease of the plant in the plant image.
[0164] The above are only embodiments of one or more embodiments of the present disclosure and are not intended to limit one or more embodiments of the present disclosure. For those skilled in the art, various changes and modifications can be made to one or more embodiments of the present disclosure. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included within the scope of the claims.
Claims
1. A plant identification method, comprising: Acquire a plant image and a question text regarding identifying the plant in the plant image; Inputting the plant image and the question text into a plant knowledge question-answering model, the plant knowledge question-answering model comprising a visual model and a large language model, wherein the visual model is configured to receive the plant image to extract image features of the plant image, the large language model is configured to receive the image features and the question text to identify the plant in the plant image, the large language model is further pre-trained on the basis of pre-training with a synthetic corpus generated based on a plant field corpus, and the plant knowledge question-answering model is trained with multimodal data, the multimodal data comprising plant images, questions about identifying the plants in the plant images, and answers to the questions; as well as Output the answer text provided by the plant knowledge question-answering model regarding the identification of the plant in the plant image.
2. The method according to claim 1, wherein: Optionally, the synthetic corpus includes a mixture of plant field corpus and synthetic corpus; Optionally, the synthetic corpus includes a mixture of general domain corpus and synthetic corpus; Optionally, the amount of the general domain corpus and the amount of the synthetic corpus are in a preset ratio, and the preset ratio is determined by the following operations: Using different proportions of general domain corpora and synthetic corpora as training data to train a large language model, and using validation data to determine the validation loss of the trained large language model, and Selecting a ratio from the different ratios as the preset ratio based on the verification loss; Optionally, the synthetic corpus is constructed by the following operations: The plant domain corpus in the plant domain corpus is segmented into one or more segments according to semantics, For each of the one or more segments, extract entities in the segment by using a language model to obtain an entity set, extract relationships between entities in each entity subset of the entity set based on the segment, and generate a synthetic corpus from the segment based on the extracted entities and / or the relationships between entities; Optionally, generating a synthetic corpus from the segment based on the extracted entities and / or the relationships between entities comprises at least one of the following: For each entity subset of the entity set, generating the rewritten content of the segment as the synthetic corpus with each entity in the entity subset as the focus, or For each entity subset of the entity set, generating analysis content about the interaction of each entity in the entity subset in the context of the segment as synthetic corpus; Optionally, for each of the one or more segments, a summary and a title of the segment are also generated by the language model, and wherein at least one of the rewritten content and the analyzed content is generated based on at least one of the summary and the title; Optionally, the number of entities included in each entity subset is two or three; Optionally, the entity comprises one or more of a plant species, a symptom, a disease, a cause of disease.
3. The method according to claim 1, wherein: Optionally, the visual model is a multimodal contrastive learning pre-trained CLIP model, and the CLIP model is trained by contrastive learning using an image-text pair dataset; Optionally, the CLIP model is first trained using an image-text pair dataset by contrastive learning alone, and then jointly trained with the large language model that has undergone the pre-training and the continued pre-training using the multimodal data.
4. The method according to claim 3, wherein: Optionally, the image-text pair dataset includes plant-domain image-text pairs and general-domain image-text pairs, and the number of the plant-domain image-text pairs is in a preset ratio to the number of the general-domain image-text pairs; Optionally, the acquired question text includes question text about identifying the species of the plant in the plant image, the multimodal data includes a plant image, a question asking about the species of the plant in the plant image, and an answer indicating the species of the plant in the plant image, and the plant field image-text pair includes at least one of a pair of a plant image and a plant Latin name, and a pair of a plant image and a set of plant feature labels; Optionally, the acquired question text includes question text about identifying a disease of the plant in the plant image, the multimodal data includes a plant image, a question asking about the disease of the plant in the plant image, and an answer indicating the disease of the plant in the plant image, and the plant field image-text pair includes at least one of a pair of a plant image and a plant disease name, and a pair of a plant image and a set of disease feature labels; Optionally, the answer text includes the symptom of the plant, and the answer text further includes one or more of the cause of the symptom, the treatment method of the symptom, and maintenance suggestions for the plant.
5. The plant identification method according to claim 1, comprising: Optionally, a maintenance plan is generated according to the answer text, the maintenance plan comprising one or more pairs, each of the one or more pairs comprising one or more maintenance tasks and an identification of a maintenance device for performing the one or more maintenance tasks, outputting the maintenance plan; Optionally, according to the identification of the maintenance device in each of the one or more pairs in the maintenance plan, the corresponding maintenance device is controlled to complete one or more maintenance tasks in the pair.
6. An electronic device comprising: processor; as well as A memory storing computer executable instructions, which, when executed by the processor, cause the processor to perform the plant identification method according to any one of claims 1 to 5. 7 . A non-transitory storage medium having computer executable instructions stored thereon, which, when executed by a computer, causes the computer to perform the plant identification method according to any one of claims 1 to 5. 8 . A computer program product, comprising instructions, which, when executed by a processor, implement the plant identification method according to claim 1 .
9. A maintenance system comprising: An electronic device comprising a processor and a memory coupled to the processor and storing instructions, the instructions, when executed by the processor, causing the processor to: obtaining a plant image and a question text regarding identifying a plant in the plant image, The plant image and the question text are input into a plant knowledge question-answering model, the plant knowledge question-answering model includes a visual model and a large language model, wherein the visual model is configured to receive the plant image to extract image features of the plant image, the large language model is configured to receive the image features and the question text to identify the plant in the plant image, the large language model is further pre-trained on the basis of pre-training with a synthetic corpus generated based on a plant field corpus, the plant knowledge question-answering model is trained with multimodal data, the multimodal data includes plant images, questions about identifying the plants in the plant images, and answers to the questions, generating an answer text provided by the plant knowledge question-answering model regarding identification of the plant in the plant image, generating a maintenance plan according to the answer text, the maintenance plan comprising one or more pairs, each of the one or more pairs comprising one or more maintenance tasks and an identification of a maintenance device for performing the one or more maintenance tasks, and transmitting, according to the identification of the maintenance device in each of the one or more pairs in the maintenance plan, a command to control the corresponding maintenance device to complete one or more maintenance tasks in the pair; and A maintenance device is communicatively coupled to the electronic device, and is configured to perform a maintenance task in response to receiving a command from the electronic device.
10. The maintenance system according to claim 9, comprising: a camera communicatively coupled to the electronic device, the camera being configured to capture plant images and transmit the captured plant images to the electronic device, The instructions include instructions that, when executed by the processor, cause the processor to perform the following operations: outputting the answer text based on the plant image received from the camera and a default question text regarding identification of the plant in the plant image, The default question text includes question text regarding identifying the type and / or disease of the plant in the plant image.
Citation Information
Cited By
Method and electronic device for determining health score of plant
CN120766066A