Method and device for detecting agricultural diseases and insect pests based on layered LoRA architecture
By employing a hierarchical LoRA architecture, multimodal images and agricultural expert knowledge are used to enhance pest and disease detection, addressing the issues of universality and efficiency in existing methods and achieving efficient and accurate pest and disease detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHONGQING THREE GORGES VOCATIONAL COLLEGE
- Filing Date
- 2026-01-22
- Publication Date
- 2026-05-01
AI Technical Summary
Existing agricultural pest and disease detection methods suffer from poor versatility, low detection efficiency, weak small-target detection capability, insufficient environmental adaptability, lack of professional knowledge integration, and complex parameter adjustment, resulting in high detection costs and poor timeliness.
A hierarchical LoRA architecture is adopted. Multimodal agricultural images are acquired and a LoRA adapter is constructed based on a pre-trained SAM model. The LoRA branch rank parameter is dynamically adjusted by combining a meta-learning network and agricultural expert knowledge to generate a semantic bias matrix. The attention function of the ViT network is used to calculate semantically enhanced disease features to achieve disease detection.
It improves the generalizability and efficiency of agricultural pest and disease detection, enhances the accuracy and adaptability of pest and disease detection, reduces training time and resource consumption, and can accurately identify the specific location and type of pests and diseases.
Smart Images

Figure CN121962834A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image processing technology, and in particular relates to a method and apparatus for detecting agricultural pests and diseases using a hierarchical LoRA architecture. Background Technology
[0002] Agricultural pest and disease detection involves real-time monitoring, analysis, and early warning of the occurrence of crop diseases and pests, thereby providing a scientific basis for prevention and control decisions. It is a core link in ensuring yield and quality in modern agriculture.
[0003] Traditional agricultural pest and disease detection mainly relies on manual inspection and identification, which has the following inherent limitations: (a) low efficiency: large-scale farmland requires a lot of manpower and time costs; (b) strong subjectivity: it depends on personal experience, the identification standards are not uniform, and the accuracy is limited; (c) high cost: professional and technical personnel are scarce, and labor costs continue to rise; (d) poor timeliness: it cannot achieve real-time monitoring and miss the best prevention and control opportunity. While deep learning-based pest and disease detection methods can solve the problem of manual inspection and identification, they still have five major shortcomings: (a) Poor generality: Existing models are usually trained for specific crops or pest and disease types, and their generalization ability is limited when facing new agricultural scenarios, requiring data collection and retraining; (b) Weak small target detection ability: Early pest and disease symptoms often manifest as small-scale features (such as tiny lesions, insect eggs, etc.), and existing methods are insufficient in multi-scale feature extraction and fusion; (c) Insufficient environmental adaptability: Complex environmental factors such as light changes, leaf shading, and background interference in the field seriously affect detection accuracy, and there is a lack of robust environmental adaptation mechanisms; (d) Lack of professional knowledge fusion: Existing methods do not fully utilize professional knowledge in the agricultural field (such as crop growth patterns, disease occurrence mechanisms, etc.), and lack semantic understanding capabilities for agricultural scenarios; (e) Low parameter efficiency: Traditional fine-tuning methods require adjusting a large number of model parameters, resulting in high training costs, difficult deployment, and unsuitability for resource-constrained edge devices.
[0004] Therefore, there is an urgent need for a hierarchical LoRA architecture-based method for detecting agricultural pests and diseases to address the issues of limited generalization and low detection efficiency in agricultural pest and disease detection. Summary of the Invention
[0005] This invention provides a hierarchical LoRA architecture method and apparatus for detecting agricultural pests and diseases, which can improve the generalizability and efficiency of agricultural pest and disease detection.
[0006] To achieve the above objectives, the present invention provides a hierarchical LoRA architecture method for detecting agricultural pests and diseases, comprising: Multimodal agricultural images are acquired, and LoRA adapters corresponding to each modality in the multimodal agricultural images are constructed based on pre-trained SAM models to obtain a multimodal adaptation architecture. The multimodal agricultural images include RGB agricultural images, near-infrared agricultural images, and thermal infrared agricultural images. The pre-trained SAM model includes an image encoder and a mask decoder. By inserting LoRA branches for different regions into the image encoder of the multimodal adaptation architecture, and dynamically adjusting the rank parameter of the LoRA branches through a meta-learning network, a fused multimodal adaptation architecture is obtained. Acquire agricultural expert knowledge, generate a semantic bias matrix based on the agricultural expert knowledge, and calculate semantically enhanced disease features based on the attention of the ViT network that integrates the image encoder in the multimodal adaptation architecture. The semantically enhanced disease features are input into the mask decoder in the fusion multimodal adaptation architecture to obtain disease detection results.
[0007] To address the aforementioned problems, the present invention also provides an agricultural pest and disease detection device with a hierarchical LoRA architecture, the device comprising: A multimodal adaptation architecture building module is used to acquire multimodal agricultural images and construct LoRA adapters for each modality in the multimodal agricultural images based on a pre-trained SAM model to obtain the multimodal adaptation architecture. The multimodal agricultural images include RGB agricultural images, near-infrared agricultural images, and thermal infrared agricultural images. The pre-trained SAM model includes an image encoder and a mask decoder. LoRA branches for different regions are inserted into the image encoder of the multimodal adaptation architecture, and the rank parameter of the LoRA branches is dynamically adjusted through a meta-learning network to obtain the fused multimodal adaptation architecture. The agricultural expert knowledge processing module is used to acquire agricultural expert knowledge, generate a semantic bias matrix based on the agricultural expert knowledge, and calculate semantically enhanced disease features based on the attention of the ViT network that integrates the image encoder in the multimodal adaptation architecture. The disease detection module is used to input semantically enhanced disease features into the mask decoder in the fusion multimodal adaptation architecture to obtain disease detection results.
[0008] To address the above problems, the present invention also provides an electronic device, the electronic device comprising: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the hierarchical LoRA architecture agricultural pest and disease detection method described above.
[0009] To address the aforementioned problems, the present invention also provides a computer-readable storage medium storing at least one computer program, which is executed by a processor in an electronic device to implement the hierarchical LoRA architecture agricultural pest and disease detection method described above.
[0010] This invention acquires multimodal agricultural images and constructs a LoRA adapter for each modality in the multimodal agricultural images based on a pre-trained SAM model, resulting in a multimodal adaptation architecture. The multimodal agricultural images provide more comprehensive information on agricultural pests and diseases, improving the accuracy of subsequent pest and disease detection. By using the image encoder and mask decoder of the pre-trained SAM model, training time and resource consumption can be reduced. Furthermore, LoRA branches targeting different regions are inserted into the image encoder of the multimodal adaptation architecture, and the rank parameter of the LoRA branches is dynamically adjusted through a meta-learning network, resulting in a fused multimodal adaptation architecture. This architecture can efficiently adjust model parameters while retaining most of the pre-trained model's parameters unchanged. By inserting LoRA branches targeting different regions, more accurate feature extraction of agricultural images can be performed, and the dynamic adjustment of the rank parameter through the meta-learning network further enhances the multimodal adaptation architecture. The rank parameter of the oRA branch can further improve the model's adaptability and flexibility. Furthermore, by acquiring agricultural expert knowledge and generating a semantic bias matrix based on this knowledge, semantically enhanced disease features are calculated using the attention mechanism of the ViT network within the image encoder of the fusion multimodal adaptation architecture. Integrating agricultural expert knowledge provides the fusion multimodal adaptation architecture with more domain-specific information. By generating the semantic bias matrix and utilizing the attention mechanism of the ViT network for semantic enhancement, the model can understand the specific semantics in agricultural scenarios, further improving the detection accuracy of pests and diseases. Finally, the semantically enhanced disease features are input into the mask decoder within the fusion multimodal adaptation architecture to obtain the disease detection results. By utilizing the semantically enhanced disease features, the decoder can more accurately identify the specific location and type of pests and diseases, improving the generalization and detection efficiency of agricultural pest and disease detection. Attached Figure Description
[0011] Figure 1 This is a flowchart illustrating an agricultural pest and disease detection method based on a hierarchical LoRA architecture, provided in an embodiment of the present invention. Figure 2 This is a schematic flowchart illustrating an example of an agricultural pest and disease detection method based on a hierarchical LoRA architecture provided in an embodiment of the present invention. Figure 3 This is a schematic diagram of the hierarchical LoRA architecture of an agricultural pest and disease detection method according to an embodiment of the present invention. Figure 4 A functional block diagram of an agricultural pest and disease detection device with a hierarchical LoRA architecture provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of an electronic device that implements the hierarchical LoRA architecture agricultural pest and disease detection method according to an embodiment of the present invention.
[0012] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0013] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0014] This application provides a hierarchical LoRA architecture-based method for detecting agricultural pests and diseases. The executing entity of this hierarchical LoRA architecture-based method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the hierarchical LoRA architecture-based method for detecting agricultural pests and diseases can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0015] Reference Figure 1 The diagram shown is a schematic flowchart of an agricultural pest and disease detection method based on a hierarchical LoRA architecture provided in an embodiment of the present invention. In this embodiment, the agricultural pest and disease detection method based on a hierarchical LoRA architecture includes: S1. Acquire multimodal agricultural images and construct LoRA adapters for each modality in the multimodal agricultural images based on the pre-trained SAM model to obtain a multimodal adaptation architecture. The multimodal agricultural images include RGB agricultural images, near-infrared agricultural images, and thermal infrared agricultural images. The pre-trained SAM model includes an image encoder and a mask decoder.
[0016] Understandably, multimodal agricultural images include RGB agricultural images, near-infrared agricultural images, and thermal infrared agricultural images. RGB agricultural images refer to ordinary color images (red, green, and blue), which are mainly used to identify the appearance characteristics of plants, such as shape, color, and texture. Near-infrared agricultural images are images acquired based on the principle that plant leaves strongly reflect near-infrared light, and are mainly used to determine the vigor of crop growth, water pressure, and vegetation cover. Thermal infrared agricultural images are images that use thermal infrared sensors to detect the thermal radiation emitted by objects, and are often used in agriculture for drought monitoring and irrigation management.
[0017] Understandably, the pre-trained SAM (Segment Anything Model) model refers to a general image segmentation model that has been pre-trained, a model proposed by Meta AI.
[0018] It is understood that an image encoder is an encoder responsible for converting an input image into a high-dimensional feature representation. In this embodiment of the invention, the Vision Transformer (ViT) network is used as the basic model of the SAM image encoder.
[0019] Understandably, a mask decoder is a decoder that calculates and outputs an accurate outline of an object's edge based on features provided by the encoder and user prompts.
[0020] Specifically, a LoRA adapter corresponding to each modality in a multimodal agricultural image is constructed based on a pre-trained SAM model, including: Keep the original weights of the pre-trained SAM model frozen for RGB agricultural images, near-infrared agricultural images, and thermal infrared agricultural images. Independent LoRA adapters are constructed separately in the pre-trained SAM model, where, for each modality... The corresponding independent LoRA adapter is obtained by linearly transforming the weight matrix. as well as For modes The corresponding independent LoRA adapter weights are adapted, and the adaptation formula is as follows: ; Among them, modality This represents RGB agricultural images, near-infrared agricultural images, or thermal infrared agricultural images. for Modal adaptation weights, The original weights of the pre-trained SAM model. For modality The scaling factor for the corresponding independent LoRA adapter, for The first LoRA low-rank parameter matrix of the mode. for The second LoRA low-rank parameter matrix of the mode, where, , ,in, for The correlation rank parameter of the mode, For the embedding dimension of the image encoder, This represents the output dimension of the mask decoder.
[0021] Furthermore, in the process of applying a linear transformation to the weight matrix of the modes... The weights of the corresponding independent LoRA branches are adapted, including: Acquire current environmental sensor data, shooting conditions, and current detection task information to construct environmental feature vectors. ; Based on environmental feature vectors Calculate the modal gating factor : ; in, for Global pooling features of modalities For the Sigmoid function, For learnable weight vectors, This is the transpose symbol.
[0022] Furthermore, after constructing the LoRA adapter corresponding to each modality in the multimodal agricultural image based on the pre-trained SAM model, it also includes cross-modal LoRA feature fusion, where the features of each modality are adapted by LoRA. By using a multi-head cross-modal attention mechanism, fused features are obtained. The fusion formula is as follows: ; MHCA stands for Multi-headed Cross-modal Attention Module, used for explicitly modeling the complementation of near-infrared agricultural images to the texture of RGB agricultural images and the enhancement of latent water stress regions by thermal infrared agricultural images. For different modes modality .
[0023] S2. Insert LoRA branches for different regions into the image encoder of the multimodal adaptation architecture, and dynamically adjust the rank parameter of the LoRA branches through the meta-learning network to obtain the fused multimodal adaptation architecture.
[0024] Understandably, LoRA branches refer to low-rank adapter modules inserted into an image encoder or mask decoder to enhance the expressive power of a specific layer.
[0025] Understandably, the rank parameter refers to the low-rank dimension in the LoRA decomposition, which determines the capacity of the adapter. The larger the rank, the stronger the expressive power.
[0026] Understandably, a meta-learning network refers to a prediction network that dynamically adjusts the rank parameter of LoRA based on the complexity of the input features and environmental conditions to achieve sample-level adaptive adjustment.
[0027] Understandably, the fusion multimodal adaptation architecture refers to the final model formed by adding dynamic rank adjustment and cross-modal feature fusion mechanisms to the multimodal adaptation architecture.
[0028] Specifically, LoRA branches for different regions are inserted into the image encoder of the multimodal adaptation architecture, including: The LoRA branch is divided into a global LoRA branch and a local LoRA branch; The global LoRA branch is inserted into the first few layers of the image encoder and configured with a first rank parameter that is greater than or equal to a preset rank parameter threshold to extract the overall health status of the crop canopy and background environmental features. The local LoRA branch is inserted into the last few layers of the image encoder and the last few layers of the mask decoder, and is configured with a second rank parameter that is less than the preset rank parameter threshold, in order to extract early micro lesions and abnormal leaf texture features.
[0029] Understandably, the global LoRA branch is used to model the overall scene, while the local LoRA branch is used to depict fine-grained lesions.
[0030] Furthermore, by dynamically adjusting the rank parameter of the LoRA branch through a meta-learning network, a multimodal adaptation architecture is obtained, including: The meta-learning network is used to predict the rank configuration of each LoRA branch in real time based on the intrinsic dimension of the input features and the task complexity. The prediction formula is as follows: in, For the first The global rank parameter of the layer, For the first The local rank parameter of the layer, For environmental feature vectors, For the first Layer fusion characteristics, For meta-learning networks, These are the parameters of the meta-learning network.
[0031] For example, after dynamically adjusting the rank parameter of the LoRA branch through a meta-learning network to obtain a fused multimodal adaptation architecture, the system further includes cross-layer LoRA residual connections and multi-scale feature aggregation to obtain aggregated features. : For the same mode The following formula is used for feature fusion: in, For learnable fusion coefficients, For modality global features For modality Local features.
[0032] Understandably, the learnable fusion coefficient As learnable fusion coefficients, they are calculated by the gated / MLP based on input features (optionally including environmental feature vectors), for example, for global features. and local features After global pooling, the concatenation is performed and then fed into the Sigmoid-gated output. The source of this information lies in the backpropagation of the segmentation loss during training, which automatically learns the gating parameters to obtain appropriate values. During inference, the output is calculated directly from the forward direction.
[0033] Understandably, the meta-learning network is used to predict the rank configuration of each LoRA branch, and its parameters must be learned through backpropagation of task / data loss during the training phase; after training is completed, it is no longer trained when deploying inference, and is only used to output rank parameters.
[0034] Understandably, in this embodiment of the invention, the rank parameter of the LoRA branch is dynamically adjusted by the meta-learning network in conjunction with feature pyramid-style upsampling and lateral connections to achieve progressive processing from coarse-grained lesion localization to fine-grained region segmentation to high-resolution detail enhancement.
[0035] S3. Acquire agricultural expert knowledge, generate a semantic bias matrix based on the agricultural expert knowledge, and calculate semantically enhanced disease features based on the attention of the ViT network that integrates the image encoder in the multimodal adaptation architecture.
[0036] Understandably, agricultural expert knowledge refers to domain knowledge derived from agricultural research literature, plant protection manuals, and expert databases. This knowledge includes crop classification, disease types, symptoms, environmental influencing factors, etc., and is represented in the form of a knowledge graph.
[0037] Understandably, the semantic bias matrix is a matrix dynamically generated based on agricultural knowledge graphs and environmental information. It is used to introduce domain knowledge bias into the Transformer attention mechanism, which can enhance the attention weight of disease-prone areas and suppress background interference.
[0038] Specifically, acquiring knowledge from agricultural experts includes: The method utilizes a multi-level agricultural knowledge graph to acquire knowledge from agricultural experts. The construction process of the multi-level agricultural knowledge graph includes a data acquisition stage, an ontology modeling stage, a relation extraction stage, a graph construction stage, and an embedding learning stage. Among these, the entity types defined in the ontology modeling stage include at least crop entities, disease entities, symptom entities, organ entities, and environmental entities. The embedding learning stage uses a graph attention network to learn the knowledge graph and maps knowledge nodes into embedding vectors.
[0039] Understandably, organ entities refer to specific parts of a plant, such as roots, stems, leaves, flowers, and fruits.
[0040] Furthermore, a semantic bias matrix is generated based on agricultural expert knowledge, including: Crop organ category prediction is performed on the features of each image patch output by the image encoder in the fusion multimodal adaptation architecture to obtain the organ category prediction results; Based on the organ category prediction results, retrieve the corresponding entity set in the agricultural knowledge graph and construct the image patch-entity association matrix; Based on the entity relationships in the agricultural knowledge graph and combined with the environmental weight vector of the current environment, a comprehensive relationship matrix is constructed; Based on the image patch-entity association moments and the comprehensive relationship matrix, an agricultural semantic bias matrix is generated through low-rank decomposition. Generate using the following formula: ; in, This refers to the transpose of the comprehensive relation matrix. This refers to the transpose of the image patch-entity association matrix. For image patch-entity association matrix, Refers to matrix and Intermediate results of multiplication, It is the total number of image patches. For entity embedding matrix, , To retrieve the corresponding set of entities in the agricultural knowledge graph. is the dimension of the entity embedding matrix.
[0041] Understandably, the entity embedding matrix is obtained by stacking all relevant entity embeddings column-wise. .
[0042] For example, image patch-entity association matrix The features of each image patch output by the ViT encoder are denoted as follows: First, a lightweight organ classification network is used to predict the organ category for each patch, thus obtaining the organ category distribution. Then, based on the predicted crop type and organ category, the corresponding entity set is retrieved from the agricultural knowledge graph. (Includes entities such as "rice leaves", "corn cob", and "rice blast lesions"). Construct an image patch-entity association matrix. ,in ; is the embedding vector of entity e in the knowledge graph, obtained by pre-training a graph attention network.
[0043] For example, the construction of the comprehensive relation matrix includes learning relation embedding matrices for each relation in the knowledge graph that is related to the occurrence of diseases (such as "organ-susceptible disease", "disease-typical symptoms", "environment-promoting disease"). An environmental weight vector is constructed by combining current environmental sensor data (temperature, humidity, reproductive stage). The comprehensive relation matrix is obtained as follows: .
[0044] Furthermore, before computing semantically enhanced disease features using the attention of the ViT network that integrates the image encoder in the multimodal adaptation architecture, knowledge-guided location encoding is also included: the input image is divided into multiple image patches, and the knowledge-guided location encoding for each image patch is computed. ;in This indicates the position index of the patch in the image sequence (values range from 0 to 4095). This represents the dimension index of the embedded vector. 10000 is the total number of dimensions of the embedded vector, and 10000 is the temperature parameter that controls the frequency range of the sine function (from the original design of Transformer, so that different dimensions correspond to a geometric series distribution from high frequency to low frequency). The exponential term is used to normalize the dimension index to the [0,2] interval to calculate the frequency decay. It is a sine function that maps positions to continuous values in [-1, 1] (even dimensions are used in the complete implementation). Odd-numbered dimensions , These are semantic embedding vectors of crop organs learned from an agricultural knowledge graph (the dimension is also [missing information]). ).
[0045] Furthermore, based on the semantic bias matrix, semantically enhanced disease features are computed using the attention of the ViT network that integrates the image encoder in the multimodal adaptation architecture, including: Attention is calculated using the following formula: The following formula is used: , among which, among which Represents the query matrix, Represents the key matrix, The value matrix represents "what information is sought at the current location", "what searchable information is provided at each location", and "the actual feature content of each location", with each dimension being [missing information]. N is the total number of image patches. The dimension of the key matrix is the agricultural semantic bias matrix. It is used to weight the semantic associations between image blocks based on agricultural knowledge graphs, assigning positive weights to regions with disease associations and negative weights to background regions.
[0046] S4. Input the semantically enhanced disease features into the mask decoder in the fusion multimodal adaptation architecture to obtain the disease detection results.
[0047] Understandably, semantically enhanced disease features refer to the disease-related feature representations extracted by the model from multimodal images, including lesion morphology, color changes, texture anomalies, etc.
[0048] This invention acquires multimodal agricultural images and constructs a LoRA adapter for each modality in the multimodal agricultural images based on a pre-trained SAM model, resulting in a multimodal adaptation architecture. The multimodal agricultural images provide more comprehensive information on agricultural pests and diseases, improving the accuracy of subsequent pest and disease detection. By using the image encoder and mask decoder of the pre-trained SAM model, training time and resource consumption can be reduced. Furthermore, LoRA branches targeting different regions are inserted into the image encoder of the multimodal adaptation architecture, and the rank parameter of the LoRA branches is dynamically adjusted through a meta-learning network, resulting in a fused multimodal adaptation architecture. This architecture can efficiently adjust model parameters while retaining most of the pre-trained model's parameters unchanged. By inserting LoRA branches targeting different regions, more accurate feature extraction of agricultural images can be performed, and the dynamic adjustment of the rank parameter through the meta-learning network further enhances the multimodal adaptation architecture. The rank parameter of the oRA branch can further improve the model's adaptability and flexibility. Furthermore, by acquiring agricultural expert knowledge and generating a semantic bias matrix based on this knowledge, semantically enhanced disease features are calculated using the attention mechanism of the ViT network within the image encoder of the fusion multimodal adaptation architecture. Integrating agricultural expert knowledge provides the fusion multimodal adaptation architecture with more domain-specific information. By generating the semantic bias matrix and utilizing the attention mechanism of the ViT network for semantic enhancement, the model can understand the specific semantics in agricultural scenarios, further improving the detection accuracy of pests and diseases. Finally, the semantically enhanced disease features are input into the mask decoder within the fusion multimodal adaptation architecture to obtain the disease detection results. By utilizing the semantically enhanced disease features, the decoder can more accurately identify the specific location and type of pests and diseases, improving the generalization and detection efficiency of agricultural pest and disease detection.
[0049] Reference Figure 2 The diagram shown is a flowchart illustrating an example of an agricultural pest and disease detection method based on a hierarchical LoRA architecture provided in an embodiment of the present invention.
[0050] Reference Figure 3 The diagram shown is a schematic diagram of the hierarchical LoRA architecture of an agricultural pest and disease detection method provided in an embodiment of the present invention.
[0051] like Figure 4 The diagram shown is a functional block diagram of an agricultural pest and disease detection device with a hierarchical LoRA architecture provided in an embodiment of the present invention.
[0052] The hierarchical LoRA architecture agricultural pest and disease detection device 100 described in this invention can be installed in an electronic device. Depending on the functions implemented, the hierarchical LoRA architecture agricultural pest and disease detection device 100 may include a multimodal adaptation architecture construction module 101, an agricultural expert knowledge processing module 102, and a disease detection module 103.
[0053] The module described in this invention can also be called a unit, which refers to a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, and are stored in the memory of the electronic device.
[0054] In this embodiment, the functions of each module / unit are as follows: The fusion multimodal adaptation architecture construction module 101 is used to acquire multimodal agricultural images and construct LoRA adapters corresponding to each modality in the multimodal agricultural images based on a pre-trained SAM model to obtain a multimodal adaptation architecture. The multimodal agricultural images include RGB agricultural images, near-infrared agricultural images, and thermal infrared agricultural images. The pre-trained SAM model includes an image encoder and a mask decoder. LoRA branches for different regions are inserted into the image encoder of the multimodal adaptation architecture, and the rank parameter of the LoRA branches is dynamically adjusted through a meta-learning network to obtain the fusion multimodal adaptation architecture.
[0055] The agricultural expert knowledge processing module 102 is used to acquire agricultural expert knowledge, generate a semantic bias matrix based on the agricultural expert knowledge, and calculate semantically enhanced disease features based on the attention of the ViT network that integrates the image encoder in the multimodal adaptation architecture according to the semantic bias matrix.
[0056] The disease detection module 103 is used to input semantically enhanced disease features into the mask decoder in the fusion multimodal adaptation architecture to obtain disease detection results.
[0057] like Figure 5 The diagram shown is a schematic representation of an electronic device that implements a hierarchical LoRA architecture for detecting agricultural pests and diseases, according to an embodiment of the present invention.
[0058] The electronic device may include a processor 10, a memory 11, a communication bus 12, and a communication interface 13. It may also include a computer program stored in the memory 11 and capable of running on the processor 10, such as a hierarchical LoRA architecture agricultural pest and disease detection method program.
[0059] In some embodiments, the processor 10 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control unit of the electronic device, connecting various components of the entire electronic device through various interfaces and lines. It executes programs or modules stored in the memory 11 (e.g., executing a hierarchical LoRA architecture agricultural pest and disease detection method program) and calls data stored in the memory 11 to perform various functions of the electronic device and process data.
[0060] The memory 11 includes at least one type of readable storage medium, including flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, disk, optical disk, etc. In some embodiments, the memory 11 can be an internal storage unit of an electronic device, such as a portable hard drive. In other embodiments, the memory 11 can be an external storage device of the electronic device, such as a plug-in portable hard drive, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. Furthermore, the memory 11 can include both internal and external storage units of the electronic device. The memory 11 can be used not only to store application software and various types of data installed on the electronic device, such as the code of an agricultural pest and disease detection method program with a hierarchical LoRA architecture, but also to temporarily store data that has been output or will be output.
[0061] The communication bus 12 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is configured to enable communication between the memory 11 and at least one processor 10, etc.
[0062] The communication interface 13 is used for communication between the aforementioned electronic device and other devices, including a network interface and a user interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a Wi-Fi interface, Bluetooth interface, etc.), typically used to establish communication connections between the electronic device and other electronic devices. The user interface may be a display, an input unit (such as a keyboard), or optionally, a standard wired or wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen, etc. The display may also be appropriately referred to as a screen or display unit, used to display information processed in the electronic device and to display a visual user interface.
[0063] Figure 5 Only electronic devices with components are shown; it will be understood by those skilled in the art that... Figure 5 The structure shown does not constitute a limitation on the electronic device and may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0064] For example, although not shown, the electronic device may also include a power supply (such as a battery) to power the various components. Preferably, the power supply can be logically connected to the at least one processor 10 through a power management device, thereby enabling functions such as charging management, discharging management, and power consumption management. The power supply may also include one or more DC or AC power supplies, recharging devices, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The electronic device may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described in detail here.
[0065] It should be understood that the embodiments described are for illustrative purposes only and are not limited to this structure in the scope of the patent application.
[0066] The program for an agricultural pest and disease detection method based on a hierarchical LoRA architecture, stored in the memory 11 of the electronic device, is a combination of multiple instructions. When run in the processor 10, it can achieve the following: Multimodal agricultural images are acquired, and LoRA adapters corresponding to each modality in the multimodal agricultural images are constructed based on pre-trained SAM models to obtain a multimodal adaptation architecture. The multimodal agricultural images include RGB agricultural images, near-infrared agricultural images, and thermal infrared agricultural images. The pre-trained SAM model includes an image encoder and a mask decoder. By inserting LoRA branches for different regions into the image encoder of the multimodal adaptation architecture, and dynamically adjusting the rank parameter of the LoRA branches through a meta-learning network, a fused multimodal adaptation architecture is obtained. Acquire agricultural expert knowledge, generate a semantic bias matrix based on the agricultural expert knowledge, and calculate semantically enhanced disease features based on the attention of the ViT network that integrates the image encoder in the multimodal adaptation architecture. The semantically enhanced disease features are input into the mask decoder in the fusion multimodal adaptation architecture to obtain disease detection results.
[0067] Specifically, the specific implementation method of the processor 10 for the above instructions can be referred to the description of the relevant steps in the corresponding embodiment of the accompanying drawings, and will not be repeated here.
[0068] Furthermore, if the modules / units integrated in the electronic device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, or a read-only memory (ROM).
[0069] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor of an electronic device, can perform the following: Multimodal agricultural images are acquired, and LoRA adapters corresponding to each modality in the multimodal agricultural images are constructed based on pre-trained SAM models to obtain a multimodal adaptation architecture. The multimodal agricultural images include RGB agricultural images, near-infrared agricultural images, and thermal infrared agricultural images. The pre-trained SAM model includes an image encoder and a mask decoder. By inserting LoRA branches for different regions into the image encoder of the multimodal adaptation architecture, and dynamically adjusting the rank parameter of the LoRA branches through a meta-learning network, a fused multimodal adaptation architecture is obtained. Acquire agricultural expert knowledge, generate a semantic bias matrix based on the agricultural expert knowledge, and calculate semantically enhanced disease features based on the attention of the ViT network that integrates the image encoder in the multimodal adaptation architecture. The semantically enhanced disease features are input into the mask decoder in the fusion multimodal adaptation architecture to obtain disease detection results.
[0070] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0071] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0072] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0073] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0074] Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the invention. No appended diagram markings in the claims should be construed as limiting the scope of the claims.
[0075] The blockchain referred to in this invention is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.
[0076] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0077] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a system claim may also be implemented by a single unit or device through software or hardware. The terms "first," "second," etc., are used to indicate names and do not indicate any specific order.
[0078] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A hierarchical LoRA architecture method for detecting agricultural pests and diseases, characterized in that, The method includes: Multimodal agricultural images are acquired, and LoRA adapters corresponding to each modality in the multimodal agricultural images are constructed based on pre-trained SAM models to obtain a multimodal adaptation architecture. The multimodal agricultural images include RGB agricultural images, near-infrared agricultural images, and thermal infrared agricultural images. The pre-trained SAM model includes an image encoder and a mask decoder. By inserting LoRA branches for different regions into the image encoder of the multimodal adaptation architecture, and dynamically adjusting the rank parameter of the LoRA branches through a meta-learning network, a fused multimodal adaptation architecture is obtained. Acquire agricultural expert knowledge, generate a semantic bias matrix based on the agricultural expert knowledge, and calculate semantically enhanced disease features based on the attention of the ViT network that integrates the image encoder in the multimodal adaptation architecture. The semantically enhanced disease features are input into the mask decoder in the fusion multimodal adaptation architecture to obtain disease detection results.
2. The method for detecting agricultural pests and diseases using a hierarchical LoRA architecture as described in claim 1, characterized in that, The pre-trained SAM model is used to construct a LoRA adapter for each modality in a multimodal agricultural image, including: While keeping the original weights of the pre-trained SAM model frozen, separate LoRA adapters are constructed in the pre-trained SAM model for RGB agricultural images, near-infrared agricultural images, and thermal infrared agricultural images, respectively. Specifically, for each modality... The corresponding independent LoRA adapter is obtained by linearly transforming the weight matrix. as well as For modes The corresponding independent LoRA adapter weights are adapted, and the adaptation formula is as follows: ; Among them, modality This represents RGB agricultural images, near-infrared agricultural images, or thermal infrared agricultural images. for Modal adaptation weights, The original weights of the pre-trained SAM model. For modality The scaling factor for the corresponding independent LoRA adapter, for The first LoRA low-rank parameter matrix of the mode. for The second LoRA low-rank parameter matrix of the mode, where, , ,in, for The correlation rank parameter of the mode, For the embedding dimension of the image encoder, This represents the output dimension of the mask decoder.
3. The method for detecting agricultural pests and diseases using a hierarchical LoRA architecture as described in claim 2, characterized in that, The mode is subjected to a linear transformation of the weight matrix. The weights of the corresponding independent LoRA branches are adapted, including: Acquire current environmental sensor data, shooting conditions, and current detection task information to construct environmental feature vectors. ; Based on environmental feature vectors Calculate the modal gating factor : ; in, for Global pooling features of modalities For the Sigmoid function, For learnable weight vectors, This is the transpose symbol.
4. The agricultural pest and disease detection method based on a hierarchical LoRA architecture as described in claim 1, characterized in that, The insertion of LoRA branches for different regions in the image encoder of the multimodal adaptation architecture includes: The LoRA branch is divided into a global LoRA branch and a local LoRA branch; The global LoRA branch is inserted into the first few layers of the image encoder and configured with a first rank parameter that is greater than or equal to a preset rank parameter threshold to extract the overall health status of the crop canopy and background environmental features. The local LoRA branch is inserted into the last few layers of the image encoder and the last few layers of the mask decoder, and is configured with a second rank parameter that is less than the preset rank parameter threshold, in order to extract early micro lesions and abnormal leaf texture features.
5. The method for detecting agricultural pests and diseases using a hierarchical LoRA architecture as described in claim 1, characterized in that, The method of dynamically adjusting the rank parameter of the LoRA branch through a meta-learning network to obtain a fused multimodal adaptation architecture includes: The meta-learning network is used to predict the rank configuration of each LoRA branch in real time based on the intrinsic dimension of the input features and the task complexity. The prediction formula is as follows: ; in, For the first The global rank parameter of the layer, For the first The local rank parameter of the layer, For environmental feature vectors, For the first Layer fusion characteristics, For meta-learning networks, These are the parameters of the meta-learning network.
6. The method for detecting agricultural pests and diseases using a hierarchical LoRA architecture as described in claim 1, characterized in that, The acquisition of agricultural expert knowledge includes: Agricultural expert knowledge is acquired using a multi-level agricultural knowledge graph. The construction process of the multi-level agricultural knowledge graph includes a data acquisition stage, an ontology modeling stage, a relation extraction stage, a graph construction stage, and an embedding learning stage. The entity types defined in the ontology modeling stage include at least crop entities, disease entities, symptom entities, organ entities, and environmental entities; The embedding learning stage uses a graph attention network to learn the knowledge graph and maps knowledge nodes into embedding vectors.
7. The method for detecting agricultural pests and diseases using a hierarchical LoRA architecture as described in claim 1, characterized in that, The generation of the semantic bias matrix based on agricultural expert knowledge includes: Crop organ category prediction is performed on the features of each image patch output by the image encoder in the fusion multimodal adaptation architecture to obtain the organ category prediction results; Based on the organ category prediction results, retrieve the corresponding entity set in the agricultural knowledge graph and construct the image patch-entity association matrix; Based on the entity relationships in the agricultural knowledge graph and combined with the environmental weight vector of the current environment, a comprehensive relationship matrix is constructed; Based on the image patch-entity association moments and the comprehensive relationship matrix, an agricultural semantic bias matrix is generated through low-rank decomposition. It is generated using the following formula: ; in, This refers to the transpose of the comprehensive relation matrix. This refers to the transpose of the image patch-entity association matrix. For image patch-entity association matrix, Refers to matrix and Intermediate results of multiplication, It is the total number of image patches. This is an entity embedding matrix.
8. A hierarchical LoRA architecture agricultural pest and disease detection device, characterized in that, The device can implement the agricultural pest and disease detection method with a hierarchical LoRA architecture as described in any one of claims 1 to 7, and the device includes: A multimodal adaptation architecture building module is used to acquire multimodal agricultural images and construct LoRA adapters for each modality in the multimodal agricultural images based on a pre-trained SAM model to obtain the multimodal adaptation architecture. The multimodal agricultural images include RGB agricultural images, near-infrared agricultural images, and thermal infrared agricultural images. The pre-trained SAM model includes an image encoder and a mask decoder. LoRA branches for different regions are inserted into the image encoder of the multimodal adaptation architecture, and the rank parameter of the LoRA branches is dynamically adjusted through a meta-learning network to obtain the fused multimodal adaptation architecture. The agricultural expert knowledge processing module is used to acquire agricultural expert knowledge, generate a semantic bias matrix based on the agricultural expert knowledge, and calculate semantically enhanced disease features based on the attention of the ViT network that integrates the image encoder in the multimodal adaptation architecture. The disease detection module is used to input semantically enhanced disease features into the mask decoder in the fusion multimodal adaptation architecture to obtain disease detection results.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the hierarchical LoRA architecture agricultural pest and disease detection method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the agricultural pest and disease detection method with a hierarchical LoRA architecture as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Whole-crop explainable disease and pest diagnosis method and system based on multi-modal large model
CN121303328A