A Zero-Shot Plant Disease Image Characterization Description Generation Method Based on Concept Enhancement
By introducing concept enhancement technology in the generation of plant disease image description, building a directed graph of dependent structures and generating descriptions using the large-scale language LLAMA3 model, the problems of single description and high computing resource consumption in the existing technology are solved, and more accurate and comprehensive phenotype description of plant disease image is achieved, which enhances the generalization ability and professionalism of the model.
Patent Information
- Application Number
- CN202411599788.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-11
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-11-11
AI Technical Summary
In the generation of pathological shape descriptions of plant disease image phenotypes, it is difficult to fully grasp the global representation of the image, resulting in the generated description text being biased towards a certain prominent point while ignoring other important features. The computing resource consumption is high, the data set collection and labeling are time-consuming and labor-intensive, and the model performance is weak, making it difficult to meet actual needs.
A zero-sample plant disease image representation description generation method based on concept enhancement is proposed. By acquiring plant disease images, building a directed graph of dependent structures, acquiring key concept pairs, and inputting a large language LLAMA3 model to generate real description text. This method incorporates concept introduction correction modules, supplement and correct generated descriptions, introduce professional terms, and enhance the accuracy and professionalism of descriptions.
The accuracy and comprehensiveness of plant disease image phenotype description is achieved, the generalization ability of the model is enhanced, and it can effectively deal with unseen crop disease types, avoid overfitting a single crop data set, and the generated description is more in line with the needs of the field and provides more reliable support for agricultural production.
Smart Images

Figure CN119540181B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of computer vision and natural language processing, and particularly relates to a method for generating zero-shot plant disease image characterization descriptions based on concept enhancement. Background Art
[0002] Crops are often threatened by various diseases during their growth process, which affects their yield and quality. Therefore, it is of great significance to timely and accurately analyze and describe the pathological characteristics of staple food crops for their high-quality growth monitoring and disease diagnosis. Through accurate pathological characteristic description, early disease warning and precise disease control can be achieved, thus ensuring the healthy growth and high yield of crops. For example, by describing in detail the lesions on rice leaves, farmers can be helped to discover and take measures in time to prevent the further spread of diseases. Similarly, for wheat rust, by accurately describing its symptoms and occurrence locations, farmers can be helped to select appropriate pesticides for control and reduce yield losses.
[0003] The technology of plant disease image characterization description has wide application potential in aspects such as crop pest and disease identification, agricultural remote sensing image analysis, agricultural robot visual navigation, and agricultural product quality assessment. By learning the shared features between different categories, this technology can generate accurate natural language descriptions to help farmers identify and respond to new pests and diseases, analyze complex agricultural scenarios, provide farmers with more comprehensive and accurate information, and assist them in making better decisions. The wide application of zero-shot image description technology in the agricultural field is expected to promote the development of smart agriculture and inject new vitality into modern agriculture.
[0004] In recent years, the research on generating pathological shape descriptions of plant disease image phenotypes mainly constructs the features of the input image by using a pre-trained visual encoder, and uses an adapter to map these features to embeddings, and then processes them as the input part of a language model. However, although these methods have great potential in theory, there are still some problems in practical applications:
[0005] 1. Although the visual encoder based on the attention mechanism can effectively extract local significant features, it is difficult to comprehensively grasp the global representation of the image, resulting in the generated description text often focusing on a certain significant point and ignoring other important features; secondly, this limitation makes the description text appear single and information-deficient when dealing with complex disease features and unable to provide a comprehensive and accurate disease description.
[0006] 2. High computing resource consumption. To adapt a pre-trained model to a specific task, a large amount of computing resources and time are usually required for fine-tuning. This process not only requires high-performance computing devices but also a high-quality dataset of image-text pairs to ensure accurate learning of image features and generation of descriptions. However, the collection and annotation of the dataset are time-consuming and laborious, and the data quality directly affects the model performance. Even with a large investment of resources, the improvement in the model performance after fine-tuning may still be marginal and difficult to meet the actual needs. These problems pose great challenges to research teams and individuals with limited resources and limit the wide application of zero-shot learning.
[0007] In summary, the generation of pathological shape descriptions for plant disease image phenotypes is a topic with important research value and development prospects. Summary of the Invention
[0008] To solve the above technical problems, the present invention will propose a method for generating zero-shot plant disease image characterization descriptions based on concept enhancement. The present invention aims to improve the accuracy and comprehensiveness of plant pathological shape descriptions, make up for the deficiencies of many current methods, and thus construct an efficient and accurate method for describing plant disease image phenotypes.
[0009] The present invention provides a method for generating zero-shot plant disease image characterization descriptions based on concept enhancement, including:
[0010] Obtain plant disease images;
[0011] Based on the plant disease images, obtain a dataset of plant disease image characterization descriptions and original texts;
[0012] Based on the dataset of characterization descriptions, obtain supporting texts;
[0013] Based on the supporting texts and the original texts, construct a dependency structure directed graph, where the dependency structure directed graph is used to indicate the corresponding dependency relationship between the supporting texts and the original texts;
[0014] According to the dependency structure directed graph, obtain key concept pairs;
[0015] Input the key concept pairs into the large language LLAMA3 model to obtain real description texts.
[0016] Optionally, obtaining a dataset of plant disease image characterization descriptions includes:
[0017] Collect plant disease images and perform category annotation on the plant disease images;
[0018] Perform data processing on the annotated plant disease images to obtain a dataset of plant disease image characterization descriptions.
[0019] Optionally, perform data processing on the labeled plant disease images to obtain a plant disease image characterization description dataset, including:
[0020] Perform target picture removal processing on the labeled plant disease images and correct the corresponding disease types;
[0021] Based on the corrected plant disease images, use a deep learning algorithm to generate label information;
[0022] Construct a description template according to the label information, and based on the description template, obtain a plant disease image characterization description dataset.
[0023] Optionally, based on the plant disease images, obtain the original text, including:
[0024] Input the plant disease images into a second pre-trained model that uses a frozen image encoder and a large language model to guide language images to obtain the original text, where the second pre-trained model consists of a visual encoder, a query generator, and a language decoder;
[0025] The visual encoder is used to extract high-dimensional features of the image;
[0026] The query generator is used to generate a joint multimodal representation according to the high-dimensional features;
[0027] The language decoder is used to generate a text description.
[0028] Optionally, based on the characterization description dataset, obtain supporting text, including:
[0029] Use the text encoder of the contrastive language-image pre-training model to encode the plant disease image characterization description dataset and the plant disease images respectively;
[0030] Calculate the similarity between the encoded characterization description dataset and the plant disease images, and based on the calculation results, screen the target texts to obtain the supporting text.
[0031] Optionally, based on the supporting text, construct a dependency structure directed graph, including:
[0032] Use the spaCy library to perform syntactic analysis on the original text and the supporting text to obtain a dependency structure;
[0033] Add a directed edge between two nodes in the dependency structure and add attributes to the directed edge to obtain the dependency structure directed graph.
[0034] Optionally, according to the dependency structure directed graph, obtain key concept pairs, including:
[0035] Concept filtering is performed on the dependency structure directed graph to obtain concept types;
[0036] The similarity score between the concept type and the plant disease image is calculated using a contrastive language-image pre-trained model;
[0037] Concepts are screened according to the similarity score to obtain the optimal key concepts;
[0038] The optimal key concepts are mapped to generate the key concept pairs.
[0039] Optionally, the large language LLAMA3 model is trained using the cross-entropy loss function.
[0040] Compared with the prior art, the present invention has the following advantages and technical effects:
[0041] The present invention solves the problem that traditional methods are limited by the over-reliance on a single crop, resulting in insufficient generalization ability of the model when facing other crops and difficulty in adapting to the diverse crop disease recognition needs. Based on a zero-shot learning image description model, the present invention ensures the generalization ability of the model for different crop diseases, can effectively handle unseen crop disease types, and avoids overfitting to a single crop dataset. Moreover, the present invention incorporates a concept introduction and correction module into the architecture, which corrects the original text generated by other methods. This module not only supplements the key details that may be missed in the description generation process but also specifically introduces professional terms in the field of plant pathology, making the generated description not only include the basic characteristics of the disease but also cover specific morphological and development degree information of the disease symptoms, making the generated description more in line with the field requirements and enhancing the accuracy and professionalism of the description. Through these improvements, the present invention realizes a method for generating phenotypic descriptions of plant disease images that has both good generalization and can generate descriptions with detailed information. This will help to more accurately identify and describe plant diseases and provide more reliable support for agricultural production. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:
[0043] Figure 1 is a flowchart of a method for generating a zero-shot plant disease image characterization description based on concept enhancement according to an embodiment of the present invention;
[0044] Figure 2 is a schematic diagram of obtaining a dataset according to an embodiment of the present invention;
[0045] Figure 3 is an exemplary display diagram of a dataset according to an embodiment of the present invention;
[0046] Figure 4 It is a schematic diagram of the overall framework of an embodiment of the present invention. Detailed implementation manners
[0047] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will detail this application with reference to the accompanying drawings and in combination with the embodiments.
[0048] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.
[0049] The present invention proposes a method for generating a zero-shot plant disease image characterization description based on concept enhancement, as Figure 1 shown, which specifically includes the following steps:
[0050] Obtain plant disease images;
[0051] Based on the plant disease images, obtain a plant disease image characterization description dataset and the original text;
[0052] Based on the characterization description dataset, obtain supporting texts;
[0053] Based on the supporting texts and the original text, construct a dependency structure directed graph, where the dependency structure directed graph is used to indicate the corresponding dependency relationship between the supporting texts and the original text;
[0054] According to the dependency structure directed graph, obtain key concept pairs;
[0055] Input the obtained key concept pairs into the large language LLAMA3 model to obtain the true description text.
[0056] Specifically, the method in this embodiment is elaborated in detail:
[0057] S1 Dataset construction: Collect plant images with diseases, store them in different folders according to species and disease categories, generate labels through an automatic annotation method and conduct manual review;
[0058] S2 Supporting text acquisition: Use the text encoder of CLIP to encode the support set, calculate the similarity with the input image, and screen the text with the highest score as the supporting text;
[0059] S3 Concept graph construction: Use the spaCy library to perform syntactic analysis on the original text and the supporting text, obtain the dependency structure, and construct a directed graph;
[0060] S4 Concept Fusion Optimization: Filter and retain the "amod" and "compound" relationships, calculate the similarity between concepts and images using the CLIP model, and extract key concept pairs;
[0061] S5 Sentence Reorganization: Fine-tune the LLAMA3 model, use the key concept mapping results and the real description text as input and output, and optimize the model through the cross-entropy loss function.
[0062] Furthermore, obtaining the plant disease image characterization description dataset includes:
[0063] Collect plant disease images and perform category annotation on the plant disease images;
[0064] Perform data processing on the annotated plant disease images to obtain the plant disease image characterization description dataset.
[0065] Furthermore, performing data processing on the annotated plant disease images to obtain the plant disease image characterization description dataset includes:
[0066] Perform removal processing on the annotated plant disease images and correct the removal processing of the annotated plant disease images;
[0067] Based on the corrected plant disease images, use deep learning algorithms to generate label information;
[0068] Construct a description template according to the label information. Based on the description template, obtain the plant disease image characterization description dataset.
[0069] Specifically, elaborate on obtaining the plant disease image characterization description dataset:
[0070] Step 1.1 Data Acquisition: Collect datasets on the network and scattered plant disease images, and record the category corresponding to each image.
[0071] Step 1.2 Data Preprocessing: Process the diseases obtained from different data sources, remove damaged and blurred pictures through image processing techniques, and correct the disease types corresponding to each image.
[0072] Step 1.3 Generate Labels: Use deep learning algorithms to generate disease characterization information for each image, such as labels for the color, texture, location, and distribution of disease symptoms.
[0073] Step 1.4 Label Processing: Since the labels generated by the machine are incorrect, manual review is performed, and two additional labels are generated manually for the image: "Leaf / Plant Color" and "Leaf / Plant Shape".
[0074] Step 1.4 Construct Templates: Description templates designed according to scientific methods. Ensure that each description template can contain all label information.
[0075] Step 1.5 description generation: To ensure the repeatability and consistency of random selection, the present invention sets the random seed to 42. In this way, the same random number sequence will be generated each time the program runs, and it also ensures that each tag information can randomly generate a unique description according to the preset template style.
[0076] Step 1.6 dataset construction: Organize the images and corresponding descriptions, and divide them into a training set, a support set, and a test set according to a ratio of 7:2:1.
[0077] Furthermore, based on the plant disease images, the original text obtained includes:
[0078] Input the plant disease images into the BLIP-2 model to obtain the original text. Among them, the BLIP-2 model consists of a visual encoder, a query generator, and a language decoder;
[0079] The visual encoder is used to extract the high-dimensional features of the images;
[0080] The query generator is used to generate a joint multi-modal representation according to the high-dimensional features;
[0081] The language decoder is used to generate text descriptions.
[0082] Furthermore, based on the characterization description dataset, the support text obtained includes:
[0083] Use the text encoder of the CLIP model to encode the plant disease image characterization description dataset and the plant disease images respectively;
[0084] Calculate the similarity between the encoded characterization description dataset and the plant disease images, and obtain the support text based on the calculation results.
[0085] Specifically, elaborate on the obtained support text in detail:
[0086] Step 2.1 construct a support database: Use the text encoder of CLIP to encode the support set in the dataset made in Step 1.
[0087] Step 2.2 similarity calculation: Encode the image using CLIP, and calculate the similarity score between the image and the support database obtained in Step 2.1.
[0088] Step 2.3 obtain support text: Screen the top few texts with the highest scores as the support text;
[0089] Furthermore, based on the support text, construct a dependency structure directed graph including:
[0090] Use the spaCy library to perform syntactic analysis on the original text and the supporting text to obtain the dependency structure;
[0091] Add a directed edge between two nodes in the dependency structure, and for the attributes of the added directed edge, obtain the directed graph of the dependency structure.
[0092] Specifically, elaborate on obtaining the directed graph of the dependency structure:
[0093] Step 3.1 Dependency Structure Obtaining: Use the spaCy library to perform syntactic analysis on the original text and the supporting text to obtain the dependency structure.
[0094] Step 3.2 Concept Graph Construction: For each pair of dependency relationships, add a directed edge between the corresponding two nodes. The direction of this edge is from the governor to the dependent, and the specific type of this dependency relationship is attached as the attribute of the edge. In this way, the analyzed syntax is constructed into a directed graph.
[0095] Furthermore, according to the directed graph of the dependency structure, the obtained key concept pairs include:
[0096] Perform concept filtering on the directed graph of the dependency structure to obtain the concept types;
[0097] Use the CLIP model to calculate the similarity score between the concept type and the plant disease image;
[0098] Filter the concepts according to the similarity score to obtain the optimal key concepts;
[0099] Map the optimal key concepts to generate key concept pairs.
[0100] Specifically, elaborate on obtaining the key concept pairs:
[0101] Step 4.1 Concept Filtering: When constructing the concept graph, the present invention retains all dependency structures, including various relationships such as root (sentence center word), nsubj (nominal subject), dobj (direct object), etc. After analysis, it is found that the two relationships of "amod" (adjective modification) and "compound" (compound word) have the greatest impact on the quality of plant disease description generation. Therefore, the present invention performs concept filtering and only retains the relationships of the "amod" and "compound" types.
[0102] Step 4.2 Similarity Calculation: Use the CLIP model to calculate the similarity score between the concept and the image.
[0103] Step 4.3 Key Concept Extraction: Rank the concepts according to the similarity score and select the one with the highest score.
[0104] Step 4.4 Key Concept Mapping: Map the selected relationships back to the modified nouns to generate key concept pairs.
[0105] Furthermore, the LLAMA3 model is trained using the cross - entropy loss function.
[0106] Specifically, for the fine - tuned model: Use the key concept mapping results obtained in Step 4 and the real description text as the input and output of the model to fine - tune the model, calculate the cross - entropy loss, and obtain the model with the best reconstruction result after multiple loop iterations.
[0107] A zero - shot plant disease image characterization description generation method based on concept enhancement proposed in this embodiment is elaborated in detail:
[0108] First, the dataset used in the training and testing of the present invention is self - made. It contains 20,943 images from PDDD, PlantDOC, and PlantVillage, covering 60 crop types and more than 300 disease characteristics. To better match the research objectives of the present invention, the images are classified and organized according to crop types and disease manifestations and stored in named folders in an appropriate format. The method of making the dataset is as shown in the appendix Figure 3 shown. The present invention uses a deep - learning model to generate labels for each image, specifying the color, texture, location, and distribution of disease symptoms. To ensure accuracy, these automatically generated labels are reviewed and corrected by domain experts, who also add two additional labels: "leaf / plant color" and "leaf / plant shape". Following these steps, each image is associated with 8 specific labels. Then, the present invention invites domain experts to create descriptive texts for each type of plant disease image, combining these eight labels. This process produces ten descriptive templates, integrating optimized and enhanced labels to generate detailed text descriptions for each picture. A total of 20,943 descriptions are created. (As shown in the appendix Figure 2 shown) Finally, the 20,943 image - text pairs are divided into a training set, a support set, and a test set in a ratio of 7:2:1 to support model training and testing, thereby enhancing the model's ability to identify newly emerging diseases in practical applications. (As shown in the appendix Figure 3 shown)
[0109] The present invention develops a framework based on the existing zero - shot image captioning model (as shown in the appendix Figure 4As shown, it can be generalized to various crop diseases and effectively handle unseen disease types, thus preventing overfitting of a specific crop dataset. To improve the model's ability to capture details, the present invention introduces three key components: Support Text Acquisition (STA), which uses CLIP to retrieve some texts most relevant to the image from the augmented database as support texts; Concept Graph Constructor (CGC), which uses spaCy to extract complex data relationships from the texts and then converts these relationships into a graphical representation for subsequent processing; and Concept Injection Refinement (CIR) module, which is a customized component for supplementing missing details and correcting inaccuracies while incorporating professional phytopathological terms. This dual approach ensures the comprehensiveness, accuracy, and domain relevance of the description, while using text-dependent analysis and graph-based representation to improve plant disease diagnosis. In addition, considering the computational cost of large-scale model training, the present invention optimizes the design of the CIR module to effectively integrate domain-specific knowledge without significantly increasing the computational overhead. This design strategy not only reduces the overall training cost but also ensures that the model can understand and utilize domain-specific professional vocabulary, thereby improving the quality and efficiency of description generation.
[0110] Step 1 Data preprocessing: The present invention uses CLIP to preprocess the validation set of the self-made dataset to enhance its richness in visual concept representation. The preprocessed validation set serves as the support database and contains visually relevant sentences related to key detail information. These sentences are crucial for removing the illusion of no training cases and reducing knowledge forgetting in pure text training cases. The whole process can be expressed as:
[0111]
[0112] SD = T'
[0113] where T represents the support set in the dataset, and E T refers to the process of encoding the text using the text encoder in the CLIP model. The formula represents that for each text t in T i , the vector representation of the text t i is calculated and the preprocessed text vectors are stored in the support database SD.
[0114] Step 2 To accurately retrieve the text description closely related to the input image, extract key concepts from it, and inject them into the zero-shot description model to ensure the description is accurate and detailed, avoiding "description illusion". Specifically, given an image I, the present invention uses CLIP to evaluate the image-text similarity and determine the relevant descriptions of the Top-N images to ensure the description quality and avoid wasting resources. Finally, the most relevant description obtained is used as the support text for subsequent operations.
[0115] I' = E I (I i )
[0116] S ( I,d ) = CLIP(I', v j )
[0117] D topN = argsord d∈D S(I, v j )[1:N]
[0118] Among them, I i represents the current image, while E I refers to the process of encoding the image I using the image encoder in the CLIP model. The present invention utilizes the mechanism of CLIP to calculate the similarity score between the image and the text, and selects the top N items with the highest scores as the supporting texts.
[0119] Step 3: Obtaining the original text. Use the pre-trained BLIP-2 to obtain the original text that needs to be corrected. BLIP-2 combines language and visual information and can perform excellently in image caption generation tasks. This model is pre-trained with a large-scale image-text pair data and learns rich semantic representations and image features. It consists of a visual encoder, a query generator, and a language decoder. The visual encoder is based on a powerful visual backbone network (such as ViT or Swin Transformer) to extract high-dimensional feature representations of the image. The query generator is a Transformer model responsible for processing the input text query and fusing it with the image features generated by the visual encoder through the self-attention mechanism to generate a joint multi-modal representation. The language decoder is based on a powerful language model (such as OPT or T5) and generates a coherent and accurate text description word by word according to the multi-modal representation generated by Q-Former.
[0120] T ori = BLIP_2(I i )
[0121] T ori represents obtaining the original text using BLIP2.
[0122] Step 4: Obtaining the dependency structure. Use the general dependency parsing tool provided by spaCy to extract the dependency structure from the original text description and the supporting materials. It can identify the grammatical associations between words, such as subject-predicate agreement, the modifying effect of adjectives on nouns, and noun-noun combinations, and transform these associations into a series of word pairs, each consisting of a dominant word and a corresponding subordinate word.
[0123] W(T) = {ω1, ω2, …, ω n}
[0124] R(T) = {Dep(ω i , ω j , r)|ω i , ω j ∈W(T), r∈R}
[0125] Among them, W(T) is all the words extracted by parsing the text using the dependency analysis tool of spaCy for the given text T, and the dependency relationship between each word is analyzed. R refers to the set of all possible types of dependency relationships. Then all the dependency relationships are collected to form a dependency relationship set.
[0126] Step 5 The present invention also designs a visualization component to verify model interpretability. This component uses graph visualization technology to generate a visual representation of the graph and highlights the relationships between nodes. In addition, the visual display can be customized to highlight the key elements of the diagram, such as the most important nodes or the closest relationships. For each pair of dependency relationships, a directed edge is added between the corresponding two nodes. The direction of this edge is from the governing word to the subordinate word, and the specific type of this dependency relationship is attached as an attribute of the edge. This graph is the visual representation of the text semantic structure.
[0127] Step 6 To further improve the quality of the description, the present invention introduces a Concept Integration Optimization (CIR) module based on CGC. This module focuses on adding missing details, correcting inaccuracies in the description, and integrating specialized plant pathology terms. Specifically, the present invention first inputs the reference text extracted from the model and the database into spaCy for dependency parsing. This process includes three main steps:
[0128] Step 6.1: Relationship filtering, CIR first analyzes the relationships extracted by CGC. Through empirical tests, the present invention finds that the 'amod' (adjective modification) and 'compound' (noun compound) relationships are particularly informative for attribute optimization. Therefore, CIR filters the extracted relationships and only retains the 'amod' and 'compound' type relationships because they are more likely to provide accurate descriptive attributes.
[0129] R filter = {r∈R(T)|type(r}∈{'amod', 'compound'}
[0130] Step 6.2: Similarity calculation. CIR uses the CLIP (Contrastive Language-Image Pretraining) model to calculate the similarity score between each filtered relationship and the corresponding image. CLIP is particularly effective in this case because it is trained to align text and visual representations and can accurately evaluate the relevance of each attribute to the image. CIR ranks the relationships based on the similarity scores and selects those with the highest scores to be included in the description. This step ensures that the attributes included in the description are not only grammatically correct but also semantically relevant to the image.
[0131] S(r, I i ) = CLIP(r, I i )
[0132]
[0133] R select ={r ∈ R filter | S r is among the top k socres}
[0134] where S(r, I i ) refers to the calculated similarity score between the dependency relationship and the image, and only the top k relationship groups with higher scores are retained as key concepts to guide subsequent operations.
[0135] Step 6.3: Mapping to key concepts: Finally, CIR maps the selected relationships back to the nouns (objects) they modify, generating a set of key concept pairs. Based on the grammatical structure of the text and the visual content of the image, these concept pairs represent the most prominent and relevant attributes in the image.
[0136] C key ={mapToConcepts(r) | r ∈ R select}
[0137] According to the dependency graph constructed in Step 5, map the key concepts back to the nouns they modify to generate key concept pairs.
[0138] Step 7 To recombine the key concepts extracted from the image into a semantically rich and smoothly expressed vision-related caption, the present invention uses the pre-trained large language model LLAMA3. LLAMA3 adopts a decoder-only Transformer architecture and uses an innovative Rotary Position Encoding (RoPE) scheme instead of traditional absolute or relative position embeddings. This position embedding method encodes absolute positions using a rotation matrix and directly integrates relative position information into the self-attention operation, thus showing significant extrapolation advantages in long text tasks. Since the pre-trained LLAMA3 model lacks expertise in the field of plants, the present invention needs to fine-tune it first. It is a unimodal model, so the input is set to the key concepts extracted in Step 6, and the output is the true description text of the corresponding image. The loss function uses cross-entropy loss, and after multiple loop iterations, the model with the best reconstruction result is obtained.
[0139]
[0140]
[0141]
[0142]
[0143] In the inference stage, first, key concepts are extracted from the image, and then these concepts are input into LLAMA3 in a specific format. The model then generates description text related to these concepts.
[0144] Through the above steps, the present invention realizes a zero-shot plant disease image characterization description generation method based on concept enhancement, which effectively alleviates the problem of single description generation commonly existing in current smart agriculture practices by using concept enhancement technology, and significantly improves the interpretability of disease recognition and diagnosis. This innovative framework provides strong support for multi-modal disease recognition and precise drug use recommendations in modern smart agriculture, significantly improves the overall diagnosis and management efficiency, and makes important contributions to the sustainable development and efficient management of modern agriculture.
[0145] The above is only the preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A zero-shot plant disease image representation description generation method based on concept enhancement, characterized in that: include: Obtain plant disease images; Based on the plant disease image, obtaining a plant disease image representation description dataset and original text; Based on the representation description dataset, obtaining supporting text; Based on the representation description dataset, obtaining supporting text includes: Using a text encoder of a contrastive language-image pre-training model, respectively encoding the plant disease image representation description dataset and the plant disease image; Calculating the similarity between the encoded representation description dataset and the plant disease image, and based on the calculation result, screening the target text to obtain the supporting text; Based on the supporting text and the original text, construct a dependency structure directed graph, wherein the dependency structure directed graph is used to indicate the corresponding dependency relationship between the supporting text and the original text; According to the dependency structure directed graph, obtaining key concept pairs; According to the dependency structure directed graph, obtaining key concept pairs includes: Performing concept filtering on the dependency structure directed graph to obtain concept types; Calculating a similarity score between the concept type and the plant disease image using a contrastive language-image pre-trained model; Screening the concepts according to the similarity scores to obtain the optimal key concepts; Mapping the optimal key concepts to generate the key concept pairs; The obtained key concept pairs are input into the Large Scale Language Architecture (LLAMA) 3 model to obtain the real description text.
2. The method for generating zero-shot plant disease image representation description based on concept enhancement according to claim 1, characterized in that: Obtaining plant disease image representation description datasets includes: Collecting plant disease images and labeling the plant disease images by category; Data processing is performed on the labeled plant disease images to obtain a plant disease image representation description data set.
3. The method for generating zero-shot plant disease image representation description based on concept enhancement according to claim 2, characterized in that: Performing data processing on the labeled plant disease images to obtain a plant disease image representation description data set includes: Performing target image removal processing on the labeled plant disease image and correcting the corresponding disease type; Based on the corrected plant disease image, generating label information using a deep learning algorithm; A description template is constructed according to the label information, and a plant disease image representation description dataset is obtained based on the description template.
4. The method for generating zero-shot plant disease image representation description based on concept enhancement according to claim 1, characterized in that: Based on the plant disease image, obtaining the original text includes: Inputting the plant disease image into a pre-trained model using a frozen image encoder and a large language model to guide a language image, to obtain the original text, wherein the pre-trained model is composed of a visual encoder, a query generator, and a language decoder; The visual encoder is used to extract high-dimensional features of the image; The query generator is used to generate a joint multimodal representation based on the high-dimensional features; The language decoder is used to generate a text description.
5. The method for generating zero-shot plant disease image representation description based on concept enhancement according to claim 1, characterized in that: Based on the supporting text, constructing a dependency structure directed graph includes: Using the spaCy library to perform grammatical analysis on the original text and the supporting text to obtain the dependency structure; A directed edge is added between two nodes in the dependency structure, and an attribute is added to the directed edge to obtain the dependency structure directed graph.
6. The method for generating zero-shot plant disease image representation description based on concept enhancement according to claim 1, characterized in that: The large language LLAMA3 model is trained using a cross entropy loss function.
Citation Information
Patent Citations
Text inference method based on limited semantic dependency analysis
CN102360346A
Method for constructing educational concept map with multiple relations from multi-source data
CN111428052A