Knowledge injection-based multi-modal agricultural disease and insect pest large model prediction method
By building a multimodal agricultural pest and disease model based on knowledge injection, the problems of scarcity of data and professional knowledge requirements in agricultural pest and disease management are solved, efficient pest and disease identification and management are achieved, and professional agricultural pest and disease intelligent assistants are provided.
Patent Information
- Application Number
- CN202510270987.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-04
AI Technical Summary
The existing technology has scarce data, unstable data quality and the need for professional knowledge in agricultural pest management, which leads to increased difficulty in identifying and controlling pests and diseases, and lacks a robust optimization model for segmentation uncertainty.
Build a multimodal agricultural pest and disease large model based on knowledge injection, and use data collection from multiple data sources, preprocessing, build feature alignment and instruction fine-tuning data sets, train and adjust the multimodal large language model, and evaluate its detection results.
It realizes the efficient identification and understanding of models in agricultural pest management, improves the accuracy and adaptability of intelligent pest management, and provides professional intelligent agricultural pest assistants.
Smart Images

Figure CN120256812A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of electric energy trading and optimal dispatching of multi - microgrid and service provider systems, and particularly relates to a multi - modal agricultural pest and disease large - model prediction method based on knowledge injection. Background Art
[0002] The present invention focuses on the content related to pest and disease knowledge in agricultural scenarios, that is, knowledge about agricultural pests and diseases such as the types, symptoms, and transmission methods of pests and diseases. This problem belongs to the cross - field of computer vision and agriculture. With the progress of smart agriculture, people are increasingly concerned about the combination of computer vision and agriculture. Driven by economic needs, it is particularly important to achieve high - quality production. For farmers, the knowledge reserve of agricultural pests and diseases is scarce. Therefore, it is particularly important to analyze and understand pests and diseases in agricultural vision problems, and to detect and prevent pests and diseases in a timely and accurate manner.
[0003] In recent years, single - modal large language models (LLMs) have demonstrated powerful capabilities. In contrast, multi - modal large language models (LMMs) better simulate the human cognitive process and understand the world through the collaboration and integration of various "sensory" (sensor) inputs. The goal of the present invention is to enable the model to imitate human - like context understanding and proficiently handle various tasks with little or no guidance. Recent research has focused on visual instruction fine - tuning. Through carefully crafted multi - modal instruction - following data, LMMs have demonstrated excellent task - completion capabilities in general domains.
[0004] Although LMMs have made significant progress in general domains, their application in specific domains, especially in the agricultural field, still faces many challenges. Agricultural pests and diseases are major problems in agricultural production and a severe challenge currently faced. Compared with general images, agricultural images are inherently more complex, containing more types of environmental variables and biological characteristics. In addition, identifying and controlling agricultural pests and diseases requires extensive domain - specific knowledge. Factors such as rapid spread, strong resistance, and complex environments further exacerbate the control difficulty. Although the success of LMMs in the medical field demonstrates the feasibility of fine - tuning for specific domains, agriculture faces challenges such as data scarcity, unstable data quality, and the need for professional knowledge. These challenges seriously hinder the development of agricultural pest and disease management mechanisms. To sum up, there is currently no piece - wise uncertainty robust optimization model. Therefore, it is necessary to propose a multi - modal agricultural pest and disease large - model prediction method based on knowledge injection to solve the above problems. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a multi - modal agricultural pest and disease large - model prediction method based on knowledge injection, aiming to apply the multi - modal large language model to agricultural scenarios to solve the problem of intelligent management of pests and diseases in smart agriculture.
[0006] To achieve the above technical effects, the technical solution adopted by the present invention is as follows: A multi-modal agricultural pest and disease large model prediction method based on knowledge injection, comprising the following steps: S1. Collect agricultural pest and disease data from multiple data sources; S2. Perform data preprocessing on the collected agricultural pest and disease data to obtain raw data in a unified format; S3. Based on the raw data in the unified format, construct an agricultural pest and disease feature alignment dataset and an agricultural pest and disease instruction fine-tuning dataset, and construct a multi-modal large language model for agricultural pests and diseases; S4. Train and adjust the constructed multi-modal large language model for agricultural pests and diseases; S5. Evaluate the detection results of the constructed multi-modal large language model for agricultural pests and diseases.
[0007] Preferably, the method for performing data preprocessing on the collected agricultural pest and disease data is as follows: Perform a preliminary screening on the collected data, delete the data containing watermark information, split the images containing multiple sub-images, and delete the abstract and irrelevant data; For the finally screened pest and disease data, standardize its naming, and the unified format is: species_pest and disease name_ordinal number.jpg.
[0008] Preferably, in step S3, constructing the agricultural pest and disease feature alignment dataset includes: Use the pest and disease category labels in the raw data, search for the corresponding knowledge online, and retain the category name and detailed symptom description as associated knowledge; Design different instruction templates for various pests and diseases so that the model can associate image features with specific categories; For each pest and disease image and its corresponding knowledge , randomly select two questions from the template and , require the model to briefly describe the pest and disease features in the image, query the category and detailed symptoms of the pest and disease corresponding to the image in the image; construct two rounds of feature alignment dialogue examples based on the image, question, and knowledge triple: ; wherein, represents the i-th pest and disease image, represents the knowledge corresponding to the pest and disease image; and represent two questions randomly selected from the template.
[0009] Preferably, in step S3, constructing an agricultural pest and disease instruction fine-tuning dataset includes: Extracting agricultural knowledge from the web, segmenting the text according to keywords constructed based on symptoms, pathogens, transmission conditions, and control methods; then organizing this structured knowledge into a standardized format and storing it in a JSON file to obtain paired images and corresponding agricultural knowledge texts; Using a structured agricultural knowledge base to guide GPT-4 to generate multi-round knowledge conversations about images; Manually creating instruction data samples, constructing compliant instruction-following data based on GPT-4, and finally obtaining multiple high-quality agricultural multi-modal conversation data.
[0010] Preferably, in step S3, constructing an agricultural pest and disease multi-modal large language model includes setting an agricultural pest and disease Chatbot benchmark and an agricultural pest and disease VQA benchmark.
[0011] Preferably, step S4 includes: S401, pre-training feature alignment data: Based on the agricultural pest and disease feature comparison data obtained in step S3, keeping the weights of the visual encoder and the LLM unchanged, and only training the projection matrix within the model to enable the model to establish the correspondence between agricultural pest and disease image features, detailed symptom descriptions, and their respective categories, thereby endowing the model with the ability to identify agricultural pests and diseases; S402, end-to-end instruction fine-tuning of the large model: Based on the agricultural pest and disease instruction tuning data obtained in step S3, freezing the weights of the visual encoder during training, while updating the weights of the projection matrix and the LLM, and fine-tuning the model on different session data.
[0012] Preferably, step S5 includes performing a VQA benchmark evaluation on the model, specifically including: The VQA evaluation metrics include closed-set questions and open-set questions. For closed-set questions, accuracy is used to measure the ability of the model to provide correct answers within the range of known questions; for open-set questions, the F1 score is used to measure the accuracy of the answers; ; ; ; ; Where represents a sample that is actually a positive class and is predicted as a positive class, represents a sample that is actually a positive class and is predicted as a negative class, represents a sample that is actually a negative class and is predicted as a negative class, Samples that are actually negative but predicted as positive; the larger the calculated value, the better the detection effect.
[0013] The beneficial effects of the present invention are as follows: The present invention proposes the first large-scale vision-language model specifically tailored for the agricultural field. To train this model, the present invention designs and constructs a large-scale agricultural multi-modal instruction-following dataset, combining extensive agricultural pest and disease knowledge with high-quality agricultural dialogue data. In addition, to comprehensively evaluate the model's capabilities in instruction following and visual reasoning, the present invention introduces the first agricultural multi-modal pest and disease benchmark. Experiments show that compared with other general models, the model proposed by the present invention demonstrates good capabilities in agricultural pest and disease dialogue and reasoning tasks. Brief Description of the Drawings
[0014] Figure 1 is a schematic diagram of the model framework of the present invention; Figure 2 is a statistical chart of pest and disease data in the embodiments of the present invention. Detailed Embodiments
[0015] Embodiment 1: As Figure 1 shown, a multi-modal agricultural pest and disease large model prediction method based on knowledge injection includes the following steps: S1. Collect agricultural pest and disease data from multiple data sources; S2. Perform data preprocessing on the collected agricultural pest and disease data to obtain raw data in a unified format; S3. Based on the raw data in a unified format, construct an agricultural pest and disease feature alignment dataset and an agricultural pest and disease instruction fine-tuning dataset, and construct a multi-modal large language model for agricultural pest and disease; S4. Train and adjust the constructed multi-modal large language model for agricultural pest and disease; S5. Evaluate the detection results of the constructed multi-modal large language model for agricultural pest and disease.
[0016] Preferably, the method for performing data preprocessing on the collected agricultural pest and disease data is: Perform preliminary screening on the collected data, delete the data containing watermark information, split the images containing multiple sub-images, and delete the abstract and irrelevant data; For the finally screened pest and disease data, standardize its naming, and the unified format is: type_pest and disease name_ordinal number.jpg.
[0017] Preferably, in step S3, constructing the agricultural pest and disease feature alignment dataset includes: Using the pest and disease category labels in the original data, search for corresponding knowledge online, and retain the category names and detailed symptom descriptions as associated knowledge; Design different instruction templates for various pests and diseases so that the model can associate image features with specific categories; For each pest and disease image and its corresponding knowledge , randomly select two questions from the template and , require the model to briefly describe the pest and disease features in the image, query the category and detailed symptoms of the pest and disease corresponding to the image in the image; construct two-round feature alignment dialogue examples based on the image, question, and knowledge triple: ; wherein, represents the i-th pest and disease image, represents the knowledge corresponding to the pest and disease image; and represent two questions randomly selected from the template.
[0018] Preferably, in step S3, constructing an agricultural pest and disease instruction fine-tuning data set includes: Extract agricultural knowledge from the web page, construct keyword pairs according to symptoms, pathogens, transmission conditions, and control methods to segment the text; then organize this structured knowledge into a standardized format and store it in a JSON file to obtain paired images and corresponding agricultural knowledge texts; Use the structured agricultural knowledge base to guide GPT-4 to generate multi-round knowledge dialogues about images; Manually create instruction data samples, construct compliant instruction-following data based on GPT-4, and finally obtain multiple high-quality agricultural multi-modal dialogue data.
[0019] Preferably, in step S3, constructing an agricultural pest and disease multi-modal large language model includes setting an agricultural pest and disease Chatbot benchmark and an agricultural pest and disease VQA benchmark.
[0020] Preferably, step S4 includes: S401, pre-train the feature alignment data: Based on the agricultural pest and disease feature comparison data obtained in step S3, keep the weights of the visual encoder and the LLM unchanged, and only train the projection matrix in the model so that the model can establish the correspondence between agricultural pest and disease image features, detailed symptom descriptions, and their respective categories, thereby endowing the model with the ability to identify agricultural pests and diseases; S402, end-to-end instruction fine-tune the large model: Based on the optimized data of agricultural pest and disease instructions obtained in step S3, during the training process, freeze the weights of the visual encoder, and at the same time update the weights of the projection matrix and the LLM, and fine-tune the model on different session data.
[0021] Preferably, step S5 includes performing a VQA benchmark evaluation on the model, specifically including: The VQA evaluation metrics include closed-set questions and open-set questions. For closed-set questions, accuracy is used to measure the ability of the model to provide correct answers within the known question range; for open-set questions, the F1 score is used to measure the accuracy of the answers. ; ; ; ; where represents a sample that is actually a positive class and is predicted as a positive class, represents a sample that is actually a positive class and is predicted as a negative class, represents a sample that is actually a negative class and is predicted as a negative class, represents a sample that is actually a negative class and is predicted as a positive class; the larger the calculated value, the better the detection effect.
[0022] Example Two: The technical solution of this example constructs an agricultural pest and disease intelligent assistant by combining agricultural pest and disease images, pest and disease knowledge, and a multimodal large language model, including the following steps: (1) Collect agricultural pest and disease data.
[0023] (2) Data preprocessing - screening the data set.
[0024] (3) Design professional agricultural pest and disease knowledge to train the model according to the complex challenges of agricultural pests and diseases.
[0025] (4) Fine-tune the multimodal large language model.
[0026] (5) Evaluate the model detection results.
[0027] Steps (1) and (2) specifically include: 1.1 Collect agricultural pest and disease-related data on major websites, such as kaggle and Baidu PaddlePaddle.
[0028] 1.2 Conduct a preliminary screening on the collected data, delete the data containing watermark information, split the images containing multiple sub-images, and delete some relatively abstract data.
[0029] 2.2 For the data of pests and diseases finally screened, standardize their naming, and the unified format is: type_pest and disease name_ordinal number.jpg.
[0030] Step (3) specifically includes: 3.1 Agricultural Pest and Disease Feature Alignment Dataset: To enable the LMM to adapt from the general domain to the agricultural domain, this embodiment uses GPT for data construction; according to steps (1) and (2), this embodiment creates an agricultural pest and disease feature alignment dataset consisting of approximately 400,000 samples. Specifically, this embodiment downloads and preprocesses 391,785 images from 16 datasets. Using the pest and disease category labels in the dataset, this embodiment searches for corresponding knowledge online and retains the category names and detailed symptom descriptions as associated knowledge; the data statistics are as Figure 2 shown.
[0031] To enable the model to associate image features with specific categories, this embodiment designs different instruction templates for various pests and diseases. The dual objectives of this embodiment are: First, enable the model to recognize image features and thus identify pest and disease categories. Second, input symptom knowledge related to pests and diseases and let the model learn the detailed symptoms of each pest and disease. This method establishes a connection among images, categories, and symptoms. For each pest and disease image and its corresponding knowledge , this embodiment randomly selects two questions and from the template. The model is required to simply describe the pest and disease features in the image, query the category and detailed symptoms corresponding to the pest and disease in the image. Based on the (image, question, knowledge) triple, this embodiment constructs two rounds of feature alignment dialogue examples: ; wherein, represents the i-th pest and disease image, represents the knowledge corresponding to this pest and disease image; and represent two questions randomly selected from the template.
[0032] 3.2 Agricultural Pest and Disease Instruction Fine-tuning Dataset: To train a professional agricultural pest and disease intelligent assistant, simply recognizing agricultural pests and diseases is not enough. It must also have domain-specific conversation capabilities. To achieve this goal, this embodiment collects 5,813 pest and disease crop images, as well as the corresponding agricultural knowledge. Then, this embodiment uses GPT-4 to generate professional knowledge conversations about these images.
[0033] Specifically, in this embodiment, agricultural knowledge is extracted from web pages, and the text is segmented according to keywords such as symptoms, pathogens, transmission conditions, and control methods. Then, this structured knowledge is organized into a standardized format and stored in a JSON file. In this way, this embodiment obtains paired images and corresponding agricultural knowledge texts. Given the highly knowledge-centric nature of the agricultural field, this embodiment uses a structured agricultural knowledge base to guide GPT-4 to generate multi-round knowledge conversations about images. This method helps to reduce knowledge-related errors in the generated conversation data. In addition, this embodiment also manually creates instruction data samples to help GPT-4 understand how to generate compliant instruction-following data. Through these processes, this embodiment finally obtains 6,000 high-quality agricultural multi-modal conversation data.
[0034] 3.3 Agricultural Pest and Disease Chatbot Benchmark: To evaluate the model's instruction-following ability, this embodiment randomly selects 30 images of various pests and diseases from Baidu Encyclopedia and the World Agrochemicals Network, including 6 pest images and 24 disease images. To test the model's performance on more challenging tasks and its generalization ability in unknown scenarios, this embodiment specifically selects 25 pests and diseases that have not been encountered during the training process. Using the same data generation method for following instructions in the second stage, this embodiment generates 4-6 rounds of conversations for each image, and these conversations cover all aspects of pest and disease knowledge, including symptoms, pathogens, transmission, and control, aiming to comprehensively evaluate the model's understanding and execution ability. Finally, this embodiment generates 151 rounds of conversations, providing sufficient data to support the evaluation of the model's performance.
[0035] 3.4 Agricultural Pest and Disease VQA Benchmark: To test the model's visual reasoning ability for plant diseases and pests, in this embodiment, 49 diseases, 50 pests, and some healthy samples were randomly selected from existing public datasets, totaling 482 images, including 6 healthy samples and 476 images of plant diseases and pests. When selecting images, this embodiment follows the following principles: First, this embodiment preferentially selects types of plant diseases and pests that do not appear during the training process. Second, when this embodiment selects plant diseases and pests that appear in the training, it ensures that the selected pictures do not appear in the training data to ensure the fairness and effectiveness of the evaluation. Finally, in the dataset collected in this embodiment, 21 diseases and 3 pests did not appear during the training. After careful selection, this embodiment manually annotates each picture to generate corresponding question-and-answer pairs. To thoroughly evaluate the model's visual reasoning ability, this embodiment designs 4-5 rounds of conversations for each image, generating a total of 2268 pairs of question answers. These questions cover multiple aspects such as the parts of the crop organs damaged by plant diseases and pests, abnormal symptoms, related attributes, potential hazards, names, causes of occurrence, control methods, and transmission routes. Through these questions, this embodiment comprehensively tests the model's understanding and reasoning ability of images of plant diseases and pests.
[0036] Step (4) specifically includes: 4.1 Pre-training feature alignment data: In this stage, this embodiment mainly uses the agricultural pest and disease feature comparison data introduced in Section 3.1. Throughout the training process, this embodiment keeps the weights of the visual encoder and the LLM unchanged and only trains the projection matrix within the model. Given an agricultural pest and disease image, this embodiment requires the model to accurately predict a specific type of pest and disease and provide a detailed symptom description of the identified pest and disease. The goal of this stage is to enable the model to establish the correspondence between the features of agricultural pest and disease images, detailed symptom descriptions, and their respective categories, thereby endowing the model with the ability to identify agricultural pests and diseases.
[0037] 4.2 End-to-end instruction fine-tuning of the large model: In this stage, this embodiment uses the agricultural pest and disease instruction tuning data introduced in Section 3.2. During the training process, this embodiment only freezes the weights of the visual encoder while updating the weights of the projection matrix and the LLM. After pre-training, the model has obtained a certain degree of domain-specific knowledge in the agricultural field but lacks the ability to answer questions. By fine-tuning the model on different session data, this embodiment aligns the model with human intentions, enabling it to process and respond to relevant domain-specific questions. The result of this process is the development of an interactive agricultural multi-modal assistant that can interact with users.
[0038] Step (5) specifically includes: 5.1 Agricultural pest and disease Chatbot benchmark evaluation: To evaluate and understand the multi-modal dialogue ability of the model, this embodiment uses GPT-4 to quantify the accuracy of the model when answering questions. Specifically, this embodiment creates (image, question, knowledge) triples, and GPT4 answers the questions based on the provided knowledge. This embodiment uses its answer as the reference answer to the question. Then, this embodiment allows the candidate model to answer the same question based on the image. After obtaining the responses of the candidate model and GPT-4 to the same image and question, the image, question, knowledge, and responses of the two assistants are input into GPT-4. Then, this embodiment asks it to evaluate the usefulness, relevance, accuracy, and detail of the answers of the two assistants and give a relative score from 1 to 10. The higher the score, the better the response, indicating better model performance. In addition, this embodiment asks GPT-4 to provide a detailed explanation of the evaluation to help better understand the model's performance in this task. This evaluation method comprehensively and objectively evaluates the model's ability in multi-modal conversation tasks and provides important insights for model improvement; the specific test results are shown in Table 1 below:
[0039] Table 1: Benchmark test results of agricultural pest and disease Chatbot; 5.2 Agricultural pest and disease VQA benchmark evaluation: The VQA evaluation metrics of this embodiment consist of two main components: closed-set questions and open-set questions. For closed-set questions, this embodiment uses accuracy to measure the model's ability to provide correct answers within the known question range. For open-set questions, we use the F1 score to measure the accuracy of the answers. The specific results are shown in Table 2 below:
[0040] Table 2: Benchmark test results of agricultural pest and disease VQA; Open-set questions involve answering queries in unknown domains, so the F1 score better reflects the coverage and accuracy of the model in different queries. These two metrics together comprehensively evaluate the model's visual reasoning ability.
[0041] ; ; ; ; where represents the samples that are actually positive and predicted as positive, represents the samples that are actually positive and predicted as negative, represents the samples that are actually negative and predicted as negative, represents the samples that are actually negative and predicted as positive; the calculated The larger the value is, the better the detection effect indicates.
Claims
1. A multi-modal agricultural pest and disease large model prediction method based on knowledge injection, characterized in that, It includes the following steps: S1. Collect agricultural pest and disease data from multiple data sources; S2. Perform data preprocessing on the collected agricultural pest and disease data to obtain raw data in a unified format; S3. Based on the raw data in a unified format, construct an agricultural pest and disease feature alignment dataset and an agricultural pest and disease instruction fine-tuning dataset, and construct an agricultural pest and disease multi-modal large language model; S4. Train and adjust the constructed agricultural pest and disease multi-modal large language model; S5. Evaluate the detection results of the constructed agricultural pest and disease multi-modal large language model.
2. The method for predicting multi-modal agricultural pest and disease large models based on knowledge injection according to claim 1, wherein In step S2, the method for performing data preprocessing on the collected agricultural pest and disease data is as follows: Perform preliminary screening on the collected data, delete the data containing watermark information, split the images containing multiple subgraphs, and delete the abstract and irrelevant data.
3. The method for predicting a multi-modal agricultural pest and disease large model based on knowledge injection according to claim 2, wherein In step S2, for the finally screened pest and disease data, standardize its naming, and the unified format is: species_pest and disease name_ordinal number.jpg.
4. The method for predicting a multi-modal agricultural pest and disease large model based on knowledge injection according to claim 1, wherein, In step S3, constructing the agricultural pest and disease feature alignment dataset includes: Use the pest and disease category labels in the raw data, search for corresponding knowledge online, and retain the category name and detailed symptom description as associated knowledge.
5. The method for predicting a multi-modal agricultural pest and disease large model based on knowledge injection according to claim 4, wherein Design different instruction templates for various pests and diseases so that the model can associate image features with specific categories.
6. The method for predicting a multi-modal agricultural pest and disease large model based on knowledge injection according to claim 5, wherein, For each pest and disease image and its corresponding knowledge , randomly select two questions and from the template and require the model to briefly describe the pest and disease characteristics in the image and query the category and detailed symptoms of the pests and diseases corresponding to the image in the image; construct two rounds of feature alignment dialogue examples based on the image, question, and knowledge triples: ; In the formula, represents the i-th pest and disease image, represents the knowledge corresponding to the pest and disease image; and represent two questions randomly selected from the template.
7. The method for predicting a multi-modal agricultural pest and disease large model based on knowledge injection according to claim 6, wherein In step S3, constructing the agricultural pest and disease instruction fine-tuning dataset includes: Extract agricultural knowledge from the web, construct keyword pairs according to symptoms, pathogens, transmission conditions, and control methods to segment the text; then organize this structured knowledge into a standardized format and store it in a JSON file to obtain paired images and corresponding agricultural knowledge texts; Use the structured agricultural knowledge base to guide GPT-4 to generate multi-round knowledge conversations about images; Manually create instruction data samples, construct compliant instruction-following data based on GPT-4, and finally obtain multiple high-quality agricultural multi-modal conversation data.
8. The method for predicting a multi-modal agricultural pest and disease large model based on knowledge injection according to claim 7, wherein, In step S3, constructing the agricultural pest and disease multi-modal large language model includes setting the agricultural pest and disease Chatbot benchmark and the agricultural pest and disease VQA benchmark.
9. The method for predicting a multi-modal agricultural pest and disease large model based on knowledge injection according to claim 8, wherein, Step S4 includes: S401. Pre-train the feature alignment data: Based on the agricultural pest and disease feature comparison data obtained in step S3, keep the weights of the visual encoder and the LLM unchanged, and only train the projection matrix in the model, so that the model can establish the corresponding relationship between agricultural pest and disease image features, detailed symptom descriptions, and their respective categories, thereby endowing the model with the ability to identify agricultural pests and diseases; S402. End-to-end instruction fine-tuning of the large model: Based on the agricultural pest and disease instruction tuning data obtained in step S3, freeze the weights of the visual encoder during the training process, and at the same time update the weights of the projection matrix and the LLM, and fine-tune the model on different session data.
10. The method for predicting a multi-modal agricultural pest and disease large model based on knowledge injection according to claim 1, wherein, Step S5 includes performing a VQA benchmark evaluation on the model, specifically including: The VQA evaluation metrics include closed-set questions and open-set questions. For closed-set questions, accuracy is used to measure the ability of the model to provide correct answers within the known question range; for open-set questions, the F1 score is used to measure the accuracy of the answers; ; ; ; ; Among them represents the samples that are actually positive and predicted as positive, represents the samples that are actually positive and predicted as negative, represents the samples that are actually negative and predicted as negative, represents the samples that are actually negative and predicted as positive; the calculated The larger the value of, the better the detection effect.
Citation Information
Cited By
Road disease intelligent identification and evaluation method based on multi-modal large model and instance segmentation algorithm
CN120808112A
Whole-crop explainable disease and pest diagnosis method and system based on multi-modal large model
CN121303328A
Large-model multi-modal and multi-dimensional data augmentation method and device based on man-machine interaction
CN121996338A