A Large Language Model-Driven Crop Leaf Disease Diagnosis Method and System
Through a large language model-driven method, a crop disease classification agent model is constructed, which solves the problems of large demands and insufficient scalability of crop disease identification data in the prior art, and realizes accurate disease diagnosis and efficient identification in zero samples.
Patent Information
- Application Number
- CN202510262677.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-06
AI Technical Summary
The prior art has problems with high data demand in crop disease identification, uncertain performance of well-trained models in new scenarios, and insufficient potential for zero-sample application and scalability.
Using a large language model-driven method, a crop disease classification agent model is constructed by extracting leaf disease images and text semantic information from basic crop management files, and a crop disease knowledge database is constructed using logical thinking chain guidance framework and prompt engineering technology to build a crop disease classification agent model to realize the diagnosis of crop disease images.
Without large-scale image training, crop diseases can be accurately diagnosed, adapted to the disease image diagnosis needs of different crops, improved zero-sample application and scalability, simplified the disease identification operation process, and reduced the demand for data and computing resources.
Smart Images

Figure CN119785223B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and system for diagnosing crop leaf diseases, and particularly to a method and system for diagnosing crop leaf diseases driven by a large language model. Background Art
[0002] Accurate and timely disease identification is crucial for effective disease management to reduce yield losses. However, due to the complexity and increasing number of crop species and existing phytopathological problems, crop disease identification is challenging. Establishing accurate and general crop disease identification models is of great significance for ensuring food security.
[0003] Deep learning has shown high classification accuracy in image-based disease identification. As a typical multi-layer feature extraction model in the image domain, convolutional neural networks have achieved remarkable success in various visual classification tasks for crop disease detection. However, these deep learning models usually require a large amount of data for training to achieve satisfactory performance in specific situations. The performance of trained models in new scenarios still has great uncertainty, especially in identifying various crop diseases.
[0004] The development of artificial intelligence technology has made it possible to train large language models (LLMs) with hundreds of billions of parameters, such as the Generative Pretrained Transformer 4 (GPT-4). These models demonstrate their powerful generative ability and text understanding ability and have achieved success in various natural language processing tasks, such as reasoning, summarization, and translation. Advanced LLMs have been rapidly studied and evaluated in many fields, such as medical diagnosis, financial reports, manuscript review, etc. In the agricultural field, recent research has also evaluated their performance in crop disease classification. However, existing research usually focuses on pre-trained vision-language models and does not consider the potential of LLMs in zero-shot applications and scalability. Summary of the Invention
[0005] To solve the problems and needs in the background art, the present invention proposes a method and system for diagnosing crop leaf diseases driven by a large language model. Without large-scale image training, the present invention can accurately diagnose crop disease images and can adapt to the diagnostic requirements of disease images of different crops, improving the problem of insufficient potential in zero-shot applications and scalability in existing work.
[0006] The technical solutions adopted by the present invention to solve its technical problems are as follows:
[0007] I. A method for diagnosing crop leaf diseases driven by a large language model
[0008] Step 1: Extract the leaf disease images of the target crop and the corresponding text semantic information from the basic crop management file to obtain the initial crop disease knowledge database;
[0009] Step 2: Optimize the initial crop disease knowledge database by combining the visual understanding of the large language model to obtain an updated crop disease knowledge database;
[0010] Step 3: Build a large language model under the guidance framework of the logical thinking chain and denote it as the crop disease classification agent model. Input the leaf disease image to be diagnosed into the crop disease classification agent model. The crop disease classification agent model diagnoses the crop disease according to the crop disease knowledge database and outputs the disease recognition result.
[0011] The specific content of Step 2 is as follows:
[0012] Randomly select multiple leaf disease images corresponding to each crop disease type in the initial crop disease knowledge database as the example images of this crop disease type. Use the large language model to generate the picture semantic information of this crop disease type, and then combine the picture semantic information of the current crop disease type to streamline the data of all text semantic information of this crop disease type, so that each text semantic information is aligned with the picture semantic information to obtain the optimized text semantic information. Then, generate the corresponding prompt information according to all the optimized text semantic information of the current crop disease type. After traversing and processing the text semantic information corresponding to other crop disease types in the initial crop disease knowledge database, obtain the optimized text semantic information and prompt information corresponding to all crop disease types, so as to obtain an updated crop disease knowledge database.
[0013] The specific content of Step 3 is as follows:
[0014] The process of the large language model in crop disease classification includes a visual understanding stage and an image classification reasoning stage; for the visual understanding stage, use prompt engineering based on chain of thought (CoT) to drive the large language model to capture the key visual features of the leaf state in the leaf disease image; for the image classification reasoning stage, use prompt engineering based on chain of thought (CoT) to remind the large language model to pay attention to the general features of the disease when calculating the matching degree between the image semantic information and the text semantic information, and regard the inconsistency and conflict between the picture semantic information and the text semantic information as an important deviation of negative samples, so that the large language model can classify the leaf disease image into the correct disease type from the perspective of reasoning, thus obtaining the crop disease classification agent model; then input the leaf disease image to be diagnosed into the crop disease classification agent model. The crop disease classification agent model diagnoses the crop disease according to the crop disease knowledge database and takes the disease type with the highest matching degree between the image semantic information and the text semantic information as the disease recognition result.
[0015] The key visual features of the leaf state include the color and shape of the leaf and the distribution pattern of the disease spots.
[0016] The general features of the disease include the spot size, color, and leaf state.
[0017] In the image classification inference stage, the prompt engineering based on Chain of Thought (CoT) generates reference inference examples based on the crop disease knowledge database. When there are semantic conflicts in the image semantic information or text semantic information is lost in the large language model, the reference inference examples are used to generate the recommended scoring tendency.
[0018] II. A Crop Leaf Disease Diagnosis System Driven by a Large Language Model
[0019] A crop disease knowledge database construction unit, which is used to extract the leaf disease images of the target crop and the corresponding text semantic information from the crop management file, so as to obtain the initial crop disease knowledge database;
[0020] A database update unit, which is used to optimize the initial crop disease knowledge database by combining the visual understanding of the large language model to obtain an updated crop disease knowledge database;
[0021] A disease diagnosis unit, which is used to combine the latest crop disease knowledge database, and the crop disease classification proxy model diagnoses the crop disease according to the leaf disease image to be diagnosed, so as to output the disease recognition result.
[0022] The crop disease classification proxy model is specifically a large language model under the guidance framework of the chain of logical thinking.
[0023] III. A Computer Device
[0024] The device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method for diagnosing crop leaf diseases driven by a large language model are implemented.
[0025] IV. A Computer Readable Storage Medium
[0026] The storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for diagnosing crop leaf diseases driven by a large language model are implemented.
[0027] V. A Computer Program Product
[0028] The product includes a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the method for diagnosing crop leaf diseases driven by a large language model are implemented.
[0029] The beneficial effects of the present invention are:
[0030] The present invention constructs a crop disease knowledge database and constructs a crop disease classification agent model in a large language model through techniques such as the chain of thought guidance strategy and prompt engineering. Without large-scale image training, it can accurately diagnose crop disease images, adapt to the disease image diagnosis needs of different crops, and improve the problem of insufficient potential in zero-shot application and scalability in the existing work.
[0031] Compared with the traditional disease recognition method that relies on large-scale image data training, the present invention only needs disease description text to identify crop diseases by inputting unlabeled images, significantly simplifies the operation process of disease recognition, reduces the requirements for data and computing resources, and also improves the accuracy of the large language model in diagnosing crop disease images.
[0032] Due to the scalability of the pathological text, the present invention can directly apply to the diagnosis of different disease images of different crops without an additional training process, which improves the generalization of the workflow and can complete zero-shot diagnosis.
[0033] In addition, the method and system proposed by the present invention have high adaptability and scalability and are easy to migrate to new disease classification scenarios. This crop disease diagnosis framework based on a large language model is especially suitable for the variable disease discrimination in the agricultural production environment, greatly improving the response efficiency of farmers and managers to different crop diseases. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a flowchart for the construction and update of the crop disease knowledge database.
[0035] Figure 2 It is a schematic diagram of the text semantic information of a tomato early blight image before and after simplification, where (a) is the original pathological description text of the tomato early blight image, and (b) is the simplified pathological description text of the tomato early blight image.
[0036] Figure 3 It is a flowchart of a method for diagnosing crop leaf diseases driven by a large language model.
[0037] Figure 4 It is the accuracy graph of tomato disease classification in this embodiment.
[0038] Figure 5 It is a comparison graph of the classification accuracy and scalability performance of newly added diseases in this embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] The present invention will be further described below with reference to the drawings and embodiments.
[0040] The present invention proposes a method for diagnosing crop leaf diseases driven by a large language model, as Figure 1 and Figure 3 shown, which specifically includes the following steps:
[0041] Step 1: Extract the text semantic information corresponding to the leaf disease images of the target crop from the basic crop management files, so as to obtain an initial crop disease knowledge database;
[0042] Step 1 is specifically as follows:
[0043] The crop disease knowledge database is to provide knowledge of crop pathology. The crop disease description text is the basis for disease classification. This information includes the visual characteristics of the leaves under the corresponding diseases, such as color, shape, and leaf condition. For example, leaves infected with bacterial leaf spot usually have small, round spots that appear dark brown to black. Especially in the first infection, the spots have a water-soaked appearance. As the disease progresses, the number of spots increases, making the leaves appear more spotted, and the infected leaves also turn yellow. The disease description text data is collected from crop management manuals available on university portals, such as the University of Minnesota Extension. These manuals mainly provide text-based disease information descriptions and related management strategies. To improve the quality of the disease description text, ChatGPT is used to process these manuals, collect clear and concise text information related to the visual characteristics of the infected leaves, and record it as the text semantic information corresponding to each leaf disease image. The selected information is then stored in JSON files. These JSON files store all the key information of the diseases in the form of key-value pairs, facilitating the subsequent invocation and analysis of methods and systems.
[0044] Step 2: Optimize the initial crop disease knowledge database by combining the visual understanding of the large language model to obtain an updated crop disease knowledge database;
[0045] Step 2 is specifically as follows:
[0046] Considering the possible gaps between the disease descriptions in the crop management manual and the visual understanding of the GPT-4o model, such as "mosaic patterns", "burn lesions", and "root and fruit characteristics". However, in the visual feature extraction stage, only relatively simple attributes such as shape and color can be accurately captured. Therefore, five leaf disease images corresponding to each crop disease type in the initial crop disease knowledge database are randomly selected and used as example images of the crop disease type, and the GPT-4o model is used to generate the picture semantic information of the crop disease type. Then, combined with the picture semantic information of the current crop disease type, data reduction is performed on all the text semantic information of the crop disease type, so that each text semantic information is aligned with the picture semantic information, and the optimized text semantic information is obtained, thereby deleting information unrelated to the visual features of the infected leaves, such as information about the roots, stems, and fruits of the infected plants. Taking the text semantic information of tomato early blight images as an example, as Figure 2 shown in (a) of Figure 2 and (b) of
[0047] After the text semantics are optimized, the score is improved, and the GPT-4o model can better understand the image. Then, according to all the optimized text semantic information of the current crop disease type, corresponding prompt information (prompt) is generated, so that the GPT-4o model can identify the diseases of crop leaves; after traversing and processing the text semantic information corresponding to other crop disease types in the initial crop disease knowledge database, the optimized text semantic information and prompt information corresponding to all crop disease types are obtained, thereby obtaining an updated crop disease knowledge database. Due to the scalability of the text, it is easy to obtain and store the text descriptions of different diseases and store them uniformly in the database of the same crop disease description.
[0048] Step 3 is specifically as follows:
[0049] The process of the large language model in crop disease classification includes task definition, visual understanding stage, and image classification reasoning stage.
[0050] In the task definition stage, the task content that GPT-4o needs to execute is clarified, including the definition of the general framework identity of GPT-4o, as well as the input, output, and detailed task content of the work process. In this embodiment, GPT-4o is defined as an experienced botanist, and the task is to evaluate whether the leaf disease in the image conforms to the semantic meaning of the text prompt according to the scoring criteria.
[0051] For the visual understanding stage, prompt engineering based on Chain of Thought (CoT) is used to drive the GPT-4o model to capture the key visual features of the leaf state in the leaf disease image, enabling the GPT-4o model to capture these key features related to crop diseases; the key visual features of the leaf state include the color and shape of the leaf and the distribution pattern of the disease spots. For example, infected leaves usually have some disease spots and vary in color and shape. The visual understanding process enables the GPT-4o model to capture these key features related to crop diseases.
[0052] For the image classification and reasoning stage, prompt engineering based on Chain of Thought (CoT) is used to remind the GPT-4o model to pay attention to the general features of the disease and the important deviation of using the inconsistency and conflict between the image semantic information and the text semantic information as negative samples when calculating the matching degree between the image semantic information and the text semantic information, enabling the GPT-4o model to classify the leaf disease image into the correct disease type from the perspective of reasoning, thereby obtaining a crop disease classification proxy model. The general features of the disease include spot size, color, and leaf state. The guidance of the present invention in the two stages enhances the reasoning ability of the GPT-4o model in crop disease classification. Among them, the inconsistency and conflict between the image semantic information and the text semantic information as an important deviation of the negative sample specifically means that when there is an inconsistency and / or conflict between the image semantic feature and the text semantic feature, it is prompted that the text semantic information tends to be a negative sample, otherwise it is prompted that the text semantic information tends to be a positive sample; in this stage, the prompt engineering based on Chain of Thought (CoT) generates reference reasoning examples based on the crop disease knowledge database, and when there is a semantic conflict in the image semantic information or the text semantic information is lost in the GPT-4o model, the reference reasoning examples are used to generate a suggested scoring tendency. Compared with traditional prompts, the developed CoT prompts enable the disease classification proxy to follow the logical reasoning chain and classify the diseases step by step.
[0053] Scoring is used to evaluate the consistency between the visual content of the disease image and the disease text description. Specifically, the GPT-4o is guided to autonomously assign a score from 0 to 100 according to the matching degree between the visual features of the diseased leaf and the semantic information in the disease description. The scoring criteria are as follows:
[0054] Low score: When there is a conflict between the text description and the image features, the overall score should tend to be a lower value (for example, the lesion is described as round, while the image shows an irregular shape of the lesion).
[0055] Medium score: When some features of the image are consistent with the corresponding disease description, the score should be a medium value.
[0056] High score: When the text description highly matches the image features, a high score should be given. The higher the matching degree, the higher the score.
[0057] The leaf disease image to be diagnosed is then input into the crop disease classification agent model. The crop disease classification agent model performs crop disease diagnosis based on the crop disease knowledge database and takes the disease type with the highest matching degree between the image semantic information and the text semantic information as the disease recognition result.
[0058] Take the four common tomato diseases - bacterial spot, late blight, yellow leaf curl virus, mosaic virus and healthy leaves as examples.
[0059] Image diagnosis example: Input a tomato mosaic virus image and different tomato pathology text descriptions in the crop disease knowledge database "Typical symptoms of tomato mosaic virus (ToMV) include mottled or mosaic patterns on leaves, in which irregular light green and dark green areas form obvious uneven colors. Leaf mottled or mosaic patterns are an important distinction from yellow leaf curl virus; tomato bacterial spot disease manifests as multiple irregular yellow-brown areas, mainly concentrated at the edges of the leaves, with many small spots or perforations scattered on the surface; late blight manifests as large areas of brown or gray-green lesions with a waterlogged appearance and blurred edges; tomato yellow leaf curl virus (TYLCV) infection usually causes the leaves to turn yellow overall, Yellowing is most noticeable in the interveinal areas, making the green veins stand out more. Leaves infected with TYLCV often curl upward and may be deformed, with a stunted or twisted appearance. Yellow leaves with prominent green veins, accompanied by obvious curling and deformation, are important distinguishing features from mosaic virus; healthy tomato leaves are uniformly dark green, with no yellow or brown spots on the surface, pinnately compound leaves, with serrated or wavy edges, and the main veins extend clearly from the base to the tip of the leaf. The leaf surface is smooth, with no obvious wrinkles or depressions. "In the tomato image disease diagnosis task, the tomato mosaic virus is matched with all the above descriptions through the logical thinking chain reasoning matching rules and the corresponding matching degree is obtained.
[0060] The proxy model outputs the disease type that best matches the disease description. The output of the LLM has a specified format and detailed instructions for handling missing data, scoring, and other issues to ensure the standardization of the framework output. The consistency of the proxy model when matching images and text does not require that the image and text are exactly the same, but a higher score is obtained when the image display and text description are relatively consistent.
[0061] The present invention also proposes a large language model driven crop leaf disease diagnosis system, comprising:
[0062] A crop disease knowledge database construction unit is used to extract leaf disease images and corresponding text semantic information of target crops from crop management files using ChatGPT, thereby obtaining an initial crop disease knowledge database, which is used to provide knowledge of crop diseases;
[0063] A database update unit for optimizing the initial crop disease knowledge database in combination with the visual understanding of large language models to obtain an updated crop disease knowledge database;
[0064] A disease diagnosis unit for combining the latest crop disease knowledge database, and the crop disease classification proxy model diagnoses crop diseases based on the leaf disease image to be diagnosed, thereby outputting a disease recognition result. Among them, the crop disease classification proxy model is used to understand the pattern of the infected leaf and identify the disease type
[0065] Next, the performance effectiveness of the method and system (ChatLeafDisease, ChatLD) proposed by the present invention is evaluated.
[0066] The GPT-4o and the Contrastive Language-Image Pretraining (CLIP) model are used as baseline methods. The original GPT-4o model is used to evaluate the contribution of the CoT prompt developed by the present invention in crop disease classification. The parameters of GPT-4o are uniformly set, the temperature is 0.1, and the value of the sampling strategy (Top-p) is 0.9. The CLIP model is a typical neural network trained on various image-text pairs. The CLIP model can predict the most relevant text segment according to the given image under natural language instructions without directly optimizing for the task, similar to the zero-shot ability of GPT-2 and GPT-3. The CLIP model shows good performance in cross-modal retrieval, image classification, and text-to-image generation.
[0067] First, the classification performance of the GPT-4o, CLIP, and ChatLD models was independently evaluated on the tomato disease dataset. Each model was evaluated 5 times to ensure the reliability of the results. Considering the domain difference, five images of each tomato disease were randomly selected to fine-tune the CLIP model. The input sample batch size was 8, and 50 rounds of fine-tuning training were carried out until the model reached the fitting state. The tomato disease classification accuracy graphs of the present invention and the GPT-4o and CLIP models are as Figure 4 shown.
[0068] As can be seen from the figure, the ChatLD framework achieved the highest classification accuracy in tomato disease classification, with an average accuracy of 88.9% (see Table 1). When directly applying the original GPT-4o model to classify tomato disease leaf images, the classification accuracy was low, with an average accuracy of 45.9%. These results indicate that it is necessary to optimize and adjust large language models (LLMs) in the corresponding fields. As a typical method for connecting text and images, the CLIP model outperformed the GPT-4o model, with an average classification accuracy of 64.3%. However, it should be noted that the CLIP model was fine-tuned with a small number of samples (five samples per disease category) to achieve this performance, while the GPT-4o and ChatLD models did not require training. The poor performance of the CLIP model indicates that crop disease classification faces difficulties in the case of scarce training data, which is a significant challenge for traditional data-driven models. The ChatLD model showed the most stable performance in five tests, with an overall accuracy range of 85.6% to 91.6%. In contrast, the accuracy of the GPT-4o and CLIP models fluctuated significantly in some categories (such as healthy leaves and yellow leaf curl virus). Notably, prompt engineering based on the chain of thought (CoT) significantly improved the classification accuracy and stability compared to the two baseline models. This significant improvement can be attributed to the logical reasoning ability conferred by the CoT prompt.
[0069] Table 1 shows the experimental results of the classification accuracy of tomato disease samples
[0070]
[0071] When facing the new demand for crop disease diagnosis, the knowledge database was updated according to the method process, and the scalability of the framework was tested on new crop disease datasets such as grapes, strawberries, and peppers (100 images per disease category). To emphasize the powerful zero-shot ability of ChatLD, the amount of training data required for the CLIP model to achieve similar performance was also tested, and the experimental results are as Figure 5 shown.
[0072] The results show that after updating the knowledge database according to the specific process, the ChatLD framework achieved high accuracy in zero-shot classification of new crop diseases, with an overall accuracy of 94.4%, showing its high stability on the new dataset (see Figure 5). The ChatLD framework provides average classification accuracies of 97.3%, 88.8%, and 100% for grape, strawberry, and pepper diseases, respectively. In traditional neural network diagnosis, as the amount of fine-tuning data increases, the performance of the CLIP model improves significantly. When the amount of data per class increases from five to 50, the average disease classification accuracy increases from 74.3% to 92.1%. However, when there are only five samples per class for fine-tuning, the performance of the CLIP model is limited. For some disease classes, such as black spot of grape (21%) and angular leaf spot of strawberry (37%), the classification accuracy is low. After increasing the number of samples per class to 50, the CLIP model provides a classification accuracy of over 80% for most classes. The results show that even with fine-tuning data, the zero-shot performance of the ChatLD framework is better than that of the CLIP model. Compared with the CLIP(50) model, the biggest improvement of the ChatLD framework is in the classification of grape diseases, with the overall classification accuracy increasing from 90.5% to 97.3%. It is worth noting that the ChatLD framework proposed in the present invention does not require training and outperforms the retrained CLIP model in the classification accuracy of three new crops. This result shows that the ChatLD framework has strong scalability and adaptability, which is crucial for crop disease classification in complex agricultural production scenarios.
[0073] From the results of the diagnosis accuracy of strawberry, grape, and pepper diseases using traditional methods and the method proposed in the present invention, it can be seen that using the method proposed in the present invention enables the large language model to accurately diagnose different crop diseases without training data, only by updating the knowledge database. Constructing the disease diagnosis method and system according to the present invention can reduce or even avoid the cost of training deep learning and improve the generalization ability of identifying new crop diseases.
[0074] Finally, it should be noted that the above embodiments and descriptions are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the disclosure of the technical solutions of the present invention, and they should all be covered by the protection scope of the claims of the present invention.
Claims
1. A large language model driven crop leaf disease diagnosis method, characterized in that: The following steps are involved: Step 1: Extract the leaf disease images and corresponding text semantic information of the target crop from the basic crop management files to obtain the initial crop disease knowledge database; Step 2: Optimize the initial crop disease knowledge database by combining the visual understanding of the large language model to obtain an updated crop disease knowledge database; Step 3: Construct a large language model under the framework of logical thinking chain guidance and record it as the crop disease classification agent model, input the leaf disease image to be diagnosed into the crop disease classification agent model, and the crop disease classification agent model performs crop disease diagnosis according to the crop disease knowledge database and outputs the disease recognition result; The step 3 is specifically as follows: The process of crop disease classification by the large language model includes the visual understanding stage and the image classification reasoning stage; for the visual understanding stage, the prompt engineering based on chain thinking is used to drive the large language model to capture the key visual features of the leaf status in the leaf disease image; for the image classification reasoning stage, the prompt engineering based on chain thinking is used to remind the large language model to pay attention to the common characteristics of the disease when calculating the matching degree between the image semantic information and the text semantic information, and to take the inconsistency and conflict between the image semantic information and the text semantic information as the important deviation of the negative sample, so that the large language model can classify the leaf disease image into the correct disease type from the perspective of reasoning, thereby obtaining the crop disease classification proxy model; then the leaf disease image to be diagnosed is input into the crop disease classification proxy model, and the crop disease classification proxy model performs crop disease diagnosis according to the crop disease knowledge database, and takes the disease type with the highest matching degree between the image semantic information and the text semantic information as the disease recognition result.
2. The method for diagnosing crop leaf diseases driven by a large language model according to claim 1, characterized in that: The step 2 is specifically as follows: A plurality of leaf disease images corresponding to each crop disease type in the initial crop disease knowledge database are randomly selected as example images of the crop disease type, and the image semantic information of the crop disease type is generated using a large language model; then, combined with the image semantic information of the current crop disease type, all text semantic information of the crop disease type is data simplified so that each text semantic information is aligned with the image semantic information to obtain optimized text semantic information; then, corresponding prompt information is generated according to all optimized text semantic information of the current crop disease type; After traversing and processing the text semantic information corresponding to other crop disease types in the initial crop disease knowledge database, the optimized text semantic information and prompt information corresponding to all crop disease types are obtained, thereby obtaining an updated crop disease knowledge database.
3. The method for diagnosing crop leaf diseases driven by a large language model according to claim 1, characterized in that: The key visual features of leaf condition include leaf color and shape and the distribution pattern of lesions.
4. The method for diagnosing crop leaf diseases driven by a large language model according to claim 1, characterized in that: Common characteristics of the disease include spot size, color and leaf condition.
5. The method for diagnosing crop leaf diseases driven by a large language model according to claim 1, characterized in that: In the image classification reasoning stage, the chain thinking-based prompt engineering generates reference reasoning examples based on the crop disease knowledge database, and uses the reference reasoning examples to generate recommended scoring tendencies when there is a semantic conflict in the image semantic information or the text semantic information is lost in the large language model.
6. A crop leaf disease diagnosis system driven by a large language model, characterized in that: include: A crop disease knowledge database construction unit is used to extract leaf disease images and corresponding text semantic information of target crops from crop management files, thereby obtaining an initial crop disease knowledge database; A database updating unit, used to optimize the initial crop disease knowledge database in combination with the visual understanding of the large language model to obtain an updated crop disease knowledge database; The disease diagnosis unit is used to combine the latest crop disease knowledge database and the crop disease classification agent model to diagnose crop diseases based on the leaf disease images to be diagnosed, thereby outputting the disease identification results; The disease diagnosis unit comprises: The construction process of the crop disease classification agent model is as follows: The process of crop disease classification by the large language model includes the visual understanding stage and the image classification reasoning stage. In the visual understanding stage, the prompt engineering based on chain thinking is used to drive the large language model to capture the key visual features of the leaf state in the leaf disease image. In the image classification reasoning stage, the prompt engineering based on chain thinking is used to remind the large language model to pay attention to the common features of the disease when calculating the matching degree between the image semantic information and the text semantic information, and to take the inconsistency and conflict between the image semantic information and the text semantic information as the important deviation of the negative sample, so that the large language model can classify the leaf disease image into the correct disease type from the perspective of reasoning, thereby obtaining a crop disease classification proxy model. The leaf disease image to be diagnosed is input into the crop disease classification agent model. The crop disease classification agent model diagnoses crop diseases according to the crop disease knowledge database and takes the disease type with the highest matching degree between the image semantic information and the text semantic information as the disease recognition result.
7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the crop leaf disease diagnosis method driven by a large language model as described in any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a crop leaf disease diagnosis method driven by a large language model as described in any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Disease and pest knowledge generation type question answering method based on big and small model collaboration
CN118551845A
Plant identification method and related device
CN118608957A