Multi-modal method and system for pest control based on visual pixel level

By segmenting pest images into multiple initial images and combining image features and text features, the problems of low pest recognition accuracy and poor generalization ability in the prior art are solved, and higher recognition accuracy and effectiveness of prevention and control suggestions are achieved.

CN120070883APending Publication Date: 2025-05-30HUAZHI RICE BIO TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510003162.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art has problems of low accuracy and poor generalization ability in pest recognition, and it is particularly difficult to deal with micro- or blurred images.

Method used

A multimodal method based on pest control at the visual pixel level is adopted. By segmenting the pest and disease images into multiple initial images, combining image features and text features, a pre-trained pest and disease control multimodal model is used for identification.

Benefits of technology

It improves the accuracy and generalization ability of pest identification, can identify pests and diseases more accurately and provide effective prevention and control suggestions, improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070883A_ABST
    Figure CN120070883A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal method and system for disease and pest control based on a visual pixel level. The method comprises the following steps: segmenting a disease and pest image into a plurality of initial images; inputting the first initial image into a coding layer to obtain a first coded image; the first initial image is any one initial image in a part of initial images randomly selected from the plurality of initial images; shielding the second initial image, and converting the second initial image into uniform embedded representation to obtain a second coded image; the second initial image is any one of the plurality of initial images except the first initial image; splicing the first coded image and the second coded image according to the position of the initial image so as to extract image features of the disease and insect pest image from the spliced image; inputting the image text into a text coding layer to obtain text features of the image text; the recognition result of the disease and pest image is obtained through the image features and the text features, and the recognition accuracy and generalization can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular to a multi-modal method and system for pest control based on visual pixel level. Background Art

[0002] In the process of agricultural production, in order to ensure the production of green, healthy and high-yield crops, it is necessary to detect and deal with pests and diseases at an early stage, prevent pathogens or pests from leaving obvious damage or pollution on the crops, and then ensure the quality and appearance of agricultural products to meet the market and consumers' demand for high-quality products.

[0003] Some existing pest control automation systems mainly rely on a large number of formed pest and disease images for training and recognition. They can not only accurately identify micro or fuzzy images, but also have the problem of misjudgment only relying on image recognition. Therefore, the existing technology has the technical problems of low accuracy and poor generalization ability in identifying pests and diseases. Summary of the Invention

[0004] The following is an overview of the subject matter described in detail in this document. This overview is not intended to limit the scope of protection of the claims.

[0005] The main purpose of the embodiments of the present disclosure is to propose a multi-modal method and system for pest control based on visual pixel level, which can combine text features and image features to improve the multi-modal accuracy and generalization ability of pest control based on visual pixel level, and then give pest control suggestions according to the recognition results to improve the user experience.

[0006] The first aspect of the embodiments of this application provides a multi-modal method for pest control based on visual pixel level, which is used for a central controller. The method includes:

[0007] Obtain a pest and disease image and its corresponding image text;

[0008] Input the pest and disease image into a pre-trained multi-modal pest control model to obtain the recognition result output by the multi-modal pest control model;

[0009] Among them, the multi-modal pest control model outputs a recognition result, including:

[0010] Segment the pest and disease image into multiple initial images;

[0011] Input the first initial image into the encoding layer to obtain the first encoded image; the first initial image is any one of a part of the initial images randomly selected from the multiple initial images;

[0012] Occlude the second initial image and convert it into a unified embedded representation to obtain a second encoded image; the second initial image is any one of the initial images other than the first initial image among the multiple initial images;

[0013] According to the position embedding of the initial image, splice the first encoded image and the second encoded image to extract the image features of the pest and disease image from the spliced image;

[0014] Input the image text into the text encoding layer to obtain the text features of the image text;

[0015] Obtain the recognition result of the pest and disease image through the image features and the text features.

[0016] The embodiment of the present application provides a multi-modal method for pest and disease control based on visual pixel level. By segmenting the pest and disease image into multiple initial images; inputting the first initial image into the encoding layer to obtain the first encoded image; the first initial image is any one of a part of the initial images randomly selected from the multiple initial images; occluding the second initial image and converting it into a unified embedded representation to obtain the second encoded image; the second initial image is any one of the multiple initial images other than the first initial image; according to the position embedding of the initial image, splicing the first encoded image and the second encoded image to extract the image features of the pest and disease image from the spliced image; inputting the image text into the text encoding layer to obtain the text features of the image text; obtaining the recognition result of the pest and disease image through the image features and the text features, the recognition accuracy and generalization ability can be improved.

[0017] In some embodiments of the present application, the training process of the pest and disease control multi-modal model includes:

[0018] Segment the pest and disease training image into multiple initial training images;

[0019] Input the first initial training image into the encoding layer to obtain the first encoded training image; the first initial training image is any one of a part of the initial training images randomly selected from the multiple initial training images;

[0020] Occlude the second initial training image and convert it into a unified embedded representation to obtain the second encoded training image; the second initial training image is any one of the multiple initial training images other than the first initial training image;

[0021] Input the first encoded training image and the second encoded training image into the decoding layer in the order of the first initial training image to obtain the restored training image;

[0022] Calculate a loss function based on the second-encoded training image and the region in the restored training image corresponding to the second-encoded training image to optimize the pest control multi-modal model.

[0023] In some embodiments of the present application, the text encoding layer is the LLaMA-2-13b-chat large language model;

[0024] The step of inputting the image text into the text encoding layer to obtain the text features of the image text includes:

[0025] Input the image text into the LLaMA-2-13b-chat large language model to obtain the text features of the image text output by the LLaMA-2-13b-chat large language model.

[0026] In some embodiments of the present application, the step of obtaining the recognition result of the pest image through the image features and the text features includes:

[0027] Obtain the modal features of the pest image according to the image features and the text features;

[0028] Input the modal features into a linear layer to generate the recognition result of the pest image.

[0029] In some embodiments of the present application, the step of obtaining the modal features of the pest image according to the image features and the text features includes:

[0030] Map the image features and the text features to a preset shared semantic space to obtain visual language features;

[0031] Perform modal separation on the visual language features to obtain the modal features.

[0032] In some embodiments of the present application, the step of performing modal separation on the visual language features includes:

[0033] Input the visual language features into a modal adaptive layer to perform modal separation operations on the visual language features;

[0034] The modal separation operation includes:

[0035] φ(X,M,m)=X⊙| {M=m} ;

[0036]

[0037] n∈[1,n];

[0038] Wherein, X is the visual language feature, M is the modality indicator of the modality adaptation layer, m is the modality indication value, ⊙ is the element-wise multiplication operation, is the query vector, is the hidden state of the previous layer given by the decoding layer, W l Q is the linear transformation matrix for calculating the query vector, is the key vector, is the feature with the modality indication value of 0, is the feature with the modality indication value of n, and is the linear transformation matrix for calculating the key vector, is the value vector, and is the linear transformation matrix for calculating the value vector, and n is the number of preset language encoder layers.

[0039] In some embodiments of the present application, after obtaining the recognition result of the pest and disease image through the image feature and the text feature, the method further includes:

[0040] Constructing a pest and disease resource database;

[0041] Generating a prevention and control suggestion corresponding to the recognition result of the pest and disease image according to the pest and disease resource database and the recognition result of the pest and disease image.

[0042] To achieve the above object, a second aspect of the embodiments of the present invention provides a multi-modal system for pest and disease control based on visual pixel level, and the system includes:

[0043] An acquisition module, configured to acquire a pest and disease image and its corresponding image text;

[0044] A recognition module, configured to input the pest and disease image into a pre-trained pest and disease control multi-modal model to obtain a recognition result output by the pest and disease control multi-modal model;

[0045] Wherein, the pest and disease control multi-modal model outputs a recognition result, including:

[0046] Segmenting the pest and disease image into multiple initial images;

[0047] Inputting the first initial image into the encoding layer to obtain a first encoded image; the first initial image is any one of a part of the initial images randomly selected from the multiple initial images;

[0048] Occlude the second initial image and convert it into a unified embedded representation to obtain a second encoded image; the second initial image is any one of the multiple initial images other than the first initial image;

[0049] According to the position embedding of the initial image, splice the first encoded image and the second encoded image to extract the image features of the pest and disease image from the spliced image;

[0050] Input the image text into the text encoding layer to obtain the text features of the image text;

[0051] Obtain the recognition result of the pest and disease image through the image features and the text features.

[0052] To achieve the above object, a third aspect of the embodiments of the present invention provides an electronic device, including: at least one control processor and a memory communicatively connected to the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to execute the above-mentioned multimodal method for pest and disease control based on visual pixel level.

[0053] To achieve the above object, a fourth aspect of the embodiments of the present invention provides a computer-readable storage medium, the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute the above-mentioned multimodal method for pest and disease control based on visual pixel level.

[0054] It can be understood that the beneficial effects of the above-mentioned second aspect to the fourth aspect compared with the related art are the same as the beneficial effects of the above-mentioned first aspect compared with the related art. For relevant descriptions, reference can be made to the relevant descriptions in the above-mentioned first aspect, and details will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] The above and / or additional aspects and advantages of the present application will become apparent and easy to understand from the description of the embodiments in conjunction with the following drawings, where:

[0056] Figure 1 is a schematic flowchart of a multimodal method for pest and disease control based on visual pixel level provided by an embodiment of the present application;

[0057] Figure 2 is a schematic application diagram provided by an embodiment of the present application;

[0058] Figure 3 is a diagram of model comparison results provided by an embodiment of the present application;

[0059] Figure 4It is a model evaluation result diagram provided by an embodiment of the present application;

[0060] Figure 5 It is an implementation schematic diagram provided by an embodiment of the present application;

[0061] Figure 6 It is a model schematic diagram provided by an embodiment of the present application;

[0062] Figure 7 It is a schematic structural diagram of a multimodal system for pest control based on visual pixel level provided by an embodiment of the present application;

[0063] Figure 8 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0064] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present application and should not be construed as a limitation to the present application.

[0065] In the description of the present application, if the first, second, etc. are described only for the purpose of distinguishing technical features, they should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence of the indicated technical features.

[0066] In the description of the present application, it should be understood that the orientation descriptions such as up, down, etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present application.

[0067] In the description of the present application, it should be noted that unless otherwise clearly defined, words such as setting, installation, connection, etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above words in the present application in combination with the specific content of the technical solution.

[0068] First, several nouns involved in the present application are analyzed:

[0069] Pests and diseases refer to the combination of diseases and pests, which often have an adverse impact on agriculture, forestry, animal husbandry, etc. In the field of agricultural science and technology, there are many animal species that harm medicinal plants, mainly insects, and in addition, there are mites, snails, rats, etc. Although many insects are pests, there are also beneficial insects, and beneficial insects should be protected, bred and utilized. Therefore, it is necessary to accurately identify pests and diseases and protect beneficial insects.

[0070] In the process of agricultural production, to ensure the production of green, healthy and high-yield crops, it is necessary to detect and deal with pests and diseases at an early stage, prevent pathogens or pests from leaving obvious damage or contamination on the crops, and then ensure the quality and appearance of agricultural products to meet the market and consumers' demand for high-quality products. Traditional multi-modal pest and disease control based on visual pixel level mainly relies on the experience of farmers or agricultural experts for observation and judgment. This method highly depends on personal experience and is easily affected by subjective factors, resulting in inconsistent recognition results. In addition, manual recognition requires a large amount of time and manpower, especially in large-scale farmland, and the efficiency of this method is extremely low.

[0071] Some existing multi-modal automation systems for pest and disease control based on visual pixel level mainly train and recognize through a large number of formed pest and disease images. The existing technology recognizes through a single modality (such as image data), ignores other information sources such as text descriptions, limits the comprehensive utilization of information, or lacks effective multi-modal fusion technology and cannot organically combine image and text information, resulting in insufficient ability of the system to process complex information. In addition, due to insufficient model complexity or inadequate training, it is unable to accurately extract and generalize the characteristics of pests and diseases, and may lack sufficient context understanding and expression ability when generating natural language suggestions, resulting in suggestions that are not natural and relevant enough.

[0072] Therefore, there are problems such as being unable to accurately identify micro or fuzzy images, low recognition accuracy and poor generalization ability, and misjudgment simply relying on image recognition, and thus it is unable to efficiently and accurately identify and control pests and diseases and give effective control suggestions.

[0073] Based on this, the embodiments of the present application provide a multi-modal method and system for pest and disease control based on visual pixel level, aiming to be able to process image and text information simultaneously, thereby improving the accuracy of model recognition, and then improving the recognition accuracy and control efficiency of pests and diseases, and giving efficient and accurate control suggestions.

[0074] The multi-modal method, system, electronic device and medium for pest and disease control based on visual pixel level provided by the embodiments of the present application are specifically described through the following embodiments. First, the multi-modal method for pest and disease control based on visual pixel level in the embodiments of the present application is described.

[0075] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is a theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0076] The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0077] The multi-modal method for pest control based on visual pixel level provided by the embodiments of the present application relates to the field of computer vision technology. The multi-modal method for pest control based on visual pixel level provided by the embodiments of the present application can be applied to a terminal, or to a server side, or can also be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application for implementing the multi-modal method for pest control based on visual pixel level, etc., but is not limited to the above forms.

[0078] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet-type devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0079] It should be noted that in each specific embodiment of the present application, when it comes to relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of the present application will be obtained.

[0080] For this reason, referring to Figure 1 , the embodiments of the present application provide a multimodal method for pest control based on visual pixel level. This method is applied to a central controller, and the controller can be a server, an electronic device, a mobile terminal, etc., which is not specifically limited here. The method includes the following steps S110 to S120.

[0081] Step S110: Obtain a pest image and its corresponding image text.

[0082] Specifically, the pest image is not limited to the photos directly taken by the image processing device (such as a visual intelligent monitor, an intelligent terminal device) equipped with a multimodal system for pest control based on visual pixel level, and also includes the pictures taken by other external devices (such as a network camera, a portable camera) and then imported or synchronized to the image processing device. In addition, the pest image may also come from a cloud storage platform and directly reach the image processing device for processing through network transmission. Similarly, these images may also be taken by independent imaging devices, such as high-end digital single-lens reflex cameras or professional video cameras in different environments, and then uploaded to the image processing device through wireless transmission (Wi-Fi, Bluetooth, etc.), data cable transmission, etc.

[0083] In this step, obtain the pest image that the user needs to identify and its corresponding image text. The pest image can be uploaded by the user independently to the device equipped with a multimodal system for pest control based on visual pixel level, or can be an image obtained by shooting through an imaging system (such as a camera). Among them, the imaging system can be configured for the image processing device itself or for other devices. For example, the pest image is a monitoring object in a specific scene photographed by an intelligent vision device equipped with an image acquisition device; or the pest image is a specific monitoring target photographed by a fixed network camera and sent to the image processing device.

[0084] Further, the image text can be descriptive text of the pest and disease image (such as symptoms, environmental conditions, etc.), or it can be demand text expressing the user's need to identify the pest and disease image, which is not limited here.

[0085] In some embodiments, as Figure 2 shown, after the user logs in to their personal account on the intelligent application equipped with the multi-modal system for pest and disease control based on visual pixel level, the user can select "Select Picture" to select the already taken image from the picture library or the image taken by other devices and transmitted to the user's current device, or select "Take Photo" to directly take an image through the imaging device of the user's current device. Finally, the user independently selects the image to be analyzed and recognized for uploading, and inputs or selects image tags according to the image to describe the image.

[0086] Step S120: Input the pest and disease image into the pre-trained multi-modal model for pest and disease control to obtain the recognition result output by the multi-modal model for pest and disease control.

[0087] In this step, the obtained pest and disease image is input into the pre-trained multi-modal model for pest and disease control, and through the analysis of the multi-modal model for pest and disease control, the recognition result output by the multi-modal model for pest and disease control is obtained.

[0088] The following explains the specific training steps of the multi-modal model for pest and disease control:

[0089] Segment the pest and disease training image into multiple initial training images;

[0090] Input the first initial training image into the encoding layer to obtain the first encoded training image; the first initial training image is any one of a part of the initial training images randomly selected from the multiple initial training images;

[0091] Occlude the second initial training image and convert it into a unified embedded representation to obtain the second encoded training image; the second initial training image is any one of the multiple initial training images other than the first initial training image;

[0092] Input the first encoded training image and the second encoded training image into the decoding layer in the order of the first initial training image to obtain the restored training image;

[0093] Calculate the loss function according to the region corresponding to the second encoded training image in the second encoded training image and the restored training image to optimize the multi-modal model for pest and disease control.

[0094] Specifically, in the specific training process of the multi-modal model for pest and disease control, first, it is necessary to obtain an image data set containing various pests and diseases and corresponding text descriptions (such as symptoms, environmental conditions, etc.). Specifically, a large amount of existing data can be obtained through existing information databases, or training images can be directly collected. Then, the obtained image data is pre-processed, including but not limited to screening or pre-processing according to the model requirements, so as to obtain image data more suitable for model training.

[0095] Furthermore, during the training process, the text data needs to be classified and labeled with symptom descriptions. Therefore, it is necessary to label the processed image data to mark the types of pests and diseases and their related characteristics.

[0096] Furthermore, the image data and its text data used for model training are input into the multi-modal model for pest and disease control to train the multi-modal model for pest and disease control. Among them, the multi-modal model for pest and disease control is preferably a Vision Transformer (ViT) model pre-trained based on Masked Autoencoders (MAE). Furthermore, by introducing a deep learning model Vision Transformer (ViT) for image recognition and segmentation tasks, it is possible to more accurately identify pests and diseases and capture detailed information in the images.

[0097] Specifically, first, the pest and disease training images input into the multi-modal model for pest and disease control are segmented into multiple initial training images. Any one of the initial training images randomly selected from a part of the initial training images among the multiple initial training images is used as the first initial training image, and any one of the initial training images other than the first initial training image among the multiple initial training images is used as the second initial training image.

[0098] Furthermore, by inputting the first initial training image into the encoding layer of the ViT (Vision Transformer) architecture in the multi-modal model for pest and disease control for encoding learning, a first encoded training image is obtained, and the original image representation of the second initial training image is blocked. Then, the blocked second initial training image is converted into a unified embedding representation to obtain a second encoded training image.

[0099] Furthermore, the first encoded training image and the second encoded training image are input into the decoding layer of the ViT architecture in the order of the first initial training image to obtain a restored training image. Then, according to the region corresponding to the second encoded training image in the second encoded training image and the restored training image, a loss function is calculated to optimize the multi-modal model for pest and disease control.

[0100] In some embodiments, a large number of field pest and disease pictures are first collected, covering 16 kinds of crops, including two categories of diseases and pests, with a total of 41 diseases and 16 pests. And corresponding pest and disease corpora are compiled based on the field pest and disease pictures, which can also be generated by using GPT-assisted dialogue. Finally, a dataset containing 35,195 image-text paired data is used as the training images, providing rich instance data support for the training of the model.

[0101] Furthermore, a Vision Transformer (ViT) pre-trained based on MAE is adopted to process high-resolution inputs. In MAE pre-training, the original training images are cut into multiple non-overlapping patches, and a part of the patches are retained and fed into the encoder of the ViT architecture for learning to obtain the learned patches, while the remaining images are masked as masks.

[0102] Furthermore, the learned patches and masks (all masks use a unified embedding, but with different position embeddings) are input into the decoder of the ViT architecture in the order of the corresponding patches in the original training images to obtain the restored images. Then, by calculating the L2 loss before and after the restoration of the masked part, the multi-modal pest and disease control model is optimized.

[0103] Furthermore, the sparse prompts (such as points, boxes, text) and dense prompts (such as masks) are processed by the prompt encoder to help the multi-modal pest and disease control model accurately locate and segment the targets, improving the fineness and accuracy of the segmentation. Finally, the mask decoder combines the image embedding, prompt embedding, and output tokens, maps them to the final mask, realizes the pixel-level fine segmentation of the multi-modal pest and disease control model, provides high-quality image segmentation results, trains the multi-modal pest and disease control model to more accurately locate and segment the target objects, and improves the fineness and accuracy of the segmentation.

[0104] In step 120, the multi-modal pest and disease control model outputs the recognition results, including the following steps:

[0105] The pest and disease image is segmented into multiple initial images;

[0106] The first initial image is input into the encoding layer to obtain the first encoded image; the first initial image is any one of a part of the initial images randomly selected from the multiple initial images;

[0107] The second initial image is masked and converted into a unified embedding representation to obtain the second encoded image; the second initial image is any one of the multiple initial images other than the first initial image;

[0108] Stitch the first encoded image and the second encoded image according to the position embedding of the initial image, so as to extract the image features of the pest and disease image from the stitched image;

[0109] Input the image text into the text encoding layer to obtain the text features of the image text;

[0110] Obtain the recognition result of the pest and disease image through the image features and the text features.

[0111] In this step, after inputting the pest and disease image into the pest and disease control multimodal model, the pest and disease control multimodal model divides the pest and disease image into multiple initial images, which helps to localize the pest and disease features and may improve the model's detection ability for different parts of the disease. Randomly select any one of the initial images from a part of the initial images selected from the multiple initial images as the first initial image, input it into the encoding layer to extract its feature representation to obtain the first encoded image, then take any one of the initial images other than the first initial image from the multiple initial images as the second initial image, and occlude the second initial image, and convert the occluded second initial image into a unified embedding representation to obtain the second encoded image, thereby increasing the robustness of the model and enabling it to better process incomplete or occluded inputs.

[0112] Specifically, occluding the second initial image can specifically be achieved through an image mask, including using a selected image, graphic, or object to occlude the processed image (entirely or partially) to control the area or process of image processing, or it can also occlude the image through other image processing methods.

[0113] Furthermore, determine the position embeddings of the first encoded image and the second encoded image according to the relative position of the initial image in the pest and disease image, and then stitch the first encoded image and the second encoded image according to the position embedding of the initial image to form a new image that combines the two encoding information, so as to extract the image features of the pest and disease image from the new image, thereby further extracting richer image features through the new image.

[0114] Furthermore, input the image text into the text encoding layer to obtain text features, and additional context can be obtained according to the text information to help improve the recognition accuracy. Finally, project the image features and the text features into a shared semantic space, and use a linear layer or an attention mechanism to achieve the transformation and fusion of the features to obtain the recognition result of the pest and disease image.

[0115] In the step of inputting the image text into the text encoding layer to obtain the text features of the image text, the following steps are included:

[0116] Input the image text into the LLaMA-2-13b-chat large language model to obtain the text features of the image text output by the LLaMA-2-13b-chat large language model.

[0117] In this step, the text encoding layer is the LLaMA-2-13b-chat large language model. Its core principle is based on the Transformer architecture, which is a deep learning model widely used in natural language processing tasks. Among them, Transformer captures long-range dependencies in the input text through the self-attention mechanism to generate more coherent and accurate outputs. The LLaMA-2-13b-chat large language model has been further optimized on this basis. Especially in the dialogue generation task, by introducing specific prompt templates and dialogue context management mechanisms, the dialogue generation ability of the model has been improved.

[0118] Furthermore, by using a large language model based on the LLaMA-2-13b-chat architecture to process the image text, the text features of the image text are obtained, and then more subtle semantic features and context information in the text are captured, thus providing more accurate text analysis.

[0119] In some embodiments, by using a large language model (LLaMA-2-13b-chat) based on the Transformer architecture as the text encoder to process the user's text description of pests and diseases, an embedding vector of the text features is generated. Specifically, the embedding vector can be used as a feature input to subsequent analysis modules for multi-modal, classification of pest and disease control at the visual pixel level, or fusion analysis with other modal data.

[0120] In the step of obtaining the recognition result of the pest and disease image from the image features and text features, the following steps are included:

[0121] Obtain the modal features of the pest and disease image according to the image features and text features;

[0122] Input the modal features into the linear layer to generate the recognition result of the pest and disease image.

[0123] In this step, by mapping the image features and text features to a preset shared semantic space, visual language features are obtained. Then, the visual language features are input into the modal adaptation layer to perform a modal separation operation on the visual language features to obtain modal features. Modal features refer to features extracted from different data types (or modalities). Finally, the modal features are input into the linear layer to generate the recognition result of the pest and disease image, realizing a multi-modal pest and disease image recognition system based on image and text information, improving the recognition accuracy and the practicality of the system.

[0124] In the step of obtaining the modal features of the pest and disease image based on the image features and text features, the following steps are included:

[0125] Map the image features and text features to a preset shared semantic space to obtain visual-language features;

[0126] Perform modal separation on the visual-language features to obtain modal features.

[0127] In this step, first, the image features and text features need to be mapped to a preset shared semantic space, which can be specifically implemented through specific encoders in the multi-modal model for pest and disease control. For example, an image encoder (such as a convolutional neural network CNN) and a text encoder (such as a language model based on the Transformer architecture). Different encoders process different types of data respectively and convert them into vectors with similar structures and meaning representations, so as to map to a preset shared semantic space, obtaining visual-language features that not only contain information from the image but also the context information provided by the text, thus expressing the content of the pest and disease image more richly and comprehensively, enabling the image features and text features to be compared and combined in the same dimension.

[0128] Furthermore, input the visual-language features into the modal adaptation layer to separate the original different modal features (i.e., image features and text features), while retaining the interaction information between the two.

[0129] In the step of performing modal separation on the visual-language features, the following steps are included:

[0130] Input the visual-language features into the modal adaptation layer to perform modal separation operations on the visual-language features;

[0131] Among them, the modal separation operations include:

[0132] φ(X,M,m)=X⊙| {M=m} ;

[0133]

[0134] n∈[1,n];

[0135] Among them, X is the visual-language feature, M is the modal indicator of the modal adaptation layer, m is the modal indicator value, ⊙ is the element-wise multiplication operation, is the query vector, is the hidden state of the previous layer given by the decoding layer, W l Q is the linear transformation matrix for calculating the query vector, is the key vector, is the feature with a modal indicator value of 0, is the feature with modal indicator value n, and To calculate the linear transformation matrix of the key vector, is a value vector, and is the linear transformation matrix for calculating the value vector, and n is the number of preset language encoder layers.

[0136] In this step, the visual language features are input into the modal adaptation layer for modal separation, and the unique information of each original modality is restored or emphasized from the fused feature representation, which helps improve the model's ability to understand different data types and better utilize information from each modality.

[0137] In some embodiments, a modality separation operation is performed by introducing an attention mechanism, specifically including separating features of a specific modality from mixed features (visual language features), φ(X,M,m)=X⊙| {M=m} It means that from the visual language feature X, the features belonging to a specific modality are selected through the modality indicator M and the specific modality indicator value m, where m∈{0,1} represents the type of modality (i.e., visual or language).

[0138] Specifically, a Boolean vector is created in which the position with a value of 1 corresponds to the position in M ​​where the value is equal to m, and then this Boolean vector is multiplied element-by-element with the visual language feature X, so that only the elements in X corresponding to the position where the value is equal to m in M ​​will be retained, and the elements at other positions will be multiplied by 0 and thus filtered out.

[0139] Furthermore, the separated features of different modalities are normalized to the same level by linear transformation W l Q From the hidden state of the previous layer Calculate the query vector Through mode separation operation and linear transformation and From the hidden state of the previous layer Calculate the key vector Through mode separation operation and linear transformation and From the hidden state of the previous layer Calculate the value vector

[0140] Furthermore, the cross-interaction of different modalities can be achieved by reconstructing the self-attention operation, while retaining the unique characteristics of each modality as much as possible. Specifically:

[0141]

[0142] Among them, C l represents the attention output, and calculates the attention weights through the dot product of the query vector and the key vector , then multiplies with the value vector to obtain the final attention output. Softmax is used to normalize the attention weights to ensure that the sum of all weights is 1. represents calculating the similarity between the query vector and the key vector, where d is the dimension of the vector, used to scale the dot product result to prevent numerical overflow.

[0143] Finally, use the output result (attention output) of the pest and disease control multimodal model to make a decision and identify the most likely pest and disease types.

[0144] After obtaining the recognition result of the pest and disease image through image features and text features in this step, it further includes:

[0145] Construct a pest and disease resource database;

[0146] According to the pest and disease resource database and the recognition result of the pest and disease image, generate prevention and control suggestions corresponding to the recognition result of the pest and disease image.

[0147] In this step, construct a pest and disease resource database through existing information, and generate prevention and control suggestions corresponding to the recognition result of the pest and disease image according to the pest and disease resource database and the recognition result of the pest and disease image.

[0148] In some embodiments, based on agricultural expert knowledge and literature, establish a knowledge base of pests and diseases and corresponding control measures, then match the recognition result with the knowledge base to generate corresponding control suggestions. For Figure 2 as shown, after the user uploads the picture and description, they can view the recognition result and suggestions after being recognized by the pest and disease control multimodal model.

[0149] In some embodiments, as Figure 3 shown, in order to verify the accuracy of the diagnosis of the pest and disease control multimodal model, the recognition accuracy test of the pest and disease control multimodal model was carried out on the test set. The average recognition accuracy rate of the pest and disease control multimodal model reached 86.7%. At the same time, twelve pests and diseases were randomly selected from the test set and compared with the current mainstream multimodal models Qwen, Zhipu, and LLaVA-13b. The average recognition accuracy rate of each type of the pest and disease control multimodal model is above 85%, and the recognition of each type is more accurate than the other three models, which is sufficient to prove that our model has a certain reliability and robustness.

[0150] In some embodiments, as Figure 4As shown, to evaluate the performance of the multi-modal dialogue for pest and disease control, the multi-modal model for pest and disease control can be evaluated according to the following scoring criteria:

[0151] (1) Accuracy: Evaluate the accuracy of the model for the multi-modal pest and disease control at the visual pixel level, including the recognition accuracy of pest and disease types.

[0152] (2) Professionalism: Evaluate whether the information provided by the model is professional and conforms to the best practices of agricultural science and pest and disease control.

[0153] (3) Relevance: Evaluate whether the information provided by the model is relevant to the pests and diseases queried by the user and whether it can provide targeted control suggestions.

[0154] (4) Interaction: Evaluate the interaction ability of the model during the dialogue process, including whether it can understand the user's intention, whether it can ask reasonable follow-up questions to obtain more information, and whether it can make appropriate adjustments according to the user's feedback.

[0155] (5) Recognition: Whether to use the multi-modal model for pest and disease control in the subsequent multi-modal control process of pest and disease control at the visual pixel level.

[0156] Specifically, 10 test samples are randomly selected from each of the twelve types of pests and diseases in the test set, for a total of 120 samples. Experts related to biological pests and diseases score the answers to each picture according to the specified criteria. Among the 120 pest and disease test cases, the recognition results of the multi-modal model for pest and disease control for a large part (86%) are evaluated as accurate by biological experts. At the same time, experts believe that the multi-modal model for pest and disease control has a high degree of relevance (82%) and professionalism (81.5%) in the provided control information, which is very meaningful for the automated implementation of pest and disease control. In addition, the multi-modal model for pest and disease control essentially aims to provide users with a one-stop interactive multi-modal control tool for pest and disease control at the visual pixel level. Experts also gave a score of 79% for the interactivity and expressed a willingness of 80% to use the multi-modal model for pest and disease control in subsequent pest and disease control. This shows that the multi-modal model for pest and disease control can provide accurate recognition results, give users reliable control opinions, and lower the threshold for users during the planting process, which is a step forward in the field of automated pest and disease identification and control.

[0157] In some embodiments, such as Figure 5As shown in the figure, in an actual farmland, the user uploads pest and disease images. The multi-modal pest and disease control model analyzes the image and text information through an image processing module and a natural language processing module, and generates the recognition results and control suggestions for pests and diseases. Therefore, the multi-modal method for pest and disease control based on visual pixel level provided in this embodiment can be integrated into a precision agriculture platform to provide farmers with real-time pest and disease monitoring and control suggestions, thereby reducing the use of pesticides, improving crop yield and quality, and enhancing the practicality and accuracy of the model.

[0158] In some embodiments, such as Figure 6 is a schematic structural diagram of the multi-modal pest and disease control model. When responding to the user's pest and disease image and the corresponding image text, the pest and disease image is input into the pre-trained multi-modal pest and disease control model. First, the trained multi-modal pest and disease control model processes the image data through the SAM module.

[0159] Specifically, first, the pest and disease image is segmented into multiple regions (for example, by a sliding window or a fixed-size block), and each region is encoded as a feature vector. These vectors together form an image embedding. Then, the SAM module uses a mask decoder to process the image embedding. Among them, the mask decoder can learn the dependencies between different regions and generate richer features, thereby achieving a more refined feature representation according to the context information.

[0160] Furthermore, based on the processed image embedding, the SAM module performs object detection to identify the key objects in the image and their positions. The results of object detection are usually given in the form of bounding boxes, and each bounding box corresponds to an object in the image.

[0161] Finally, the SAM module outputs the processed image embeddings, which contain an in-depth understanding of the image content, and passes the image embeddings to the MMA module for fusion with the text embeddings, so as to better support subsequent tasks.

[0162] Furthermore, the image text is processed by a large model embedding model to obtain the text features (text embeddings) extracted from the image text.

[0163] Furthermore, the image features and text features are fused and aligned through the MMA module to obtain the recognition results of the pest and disease image, and then a natural language answer is generated through the large model decoder.

[0164] Specifically, the MMA module is used to process multimodal data (images and texts). It aligns the image embeddings and text embeddings through a certain mechanism (such as the attention mechanism or linear transformation), enabling them to be compared and fused in the same semantic space. Then, the aligned image and text features are further fused to generate a comprehensive multimodal representation, which contains the joint information of the image and text. Finally, the MMA module outputs a fused multimodal representation for generating the final answer or prediction.

[0165] Among them, in the MMA module, the modality separation operation is performed by introducing the attention mechanism, such as Figure 6 where W represents a matrix for linear transformation to adjust the dimension of the feature vector or perform feature mapping. In multimodal alignment and fusion, and are the linear transformation matrices for calculating the key vectors, and are the linear transformation matrices for calculating the value vectors, and W Q is the linear transformation matrix for calculating the query vector.

[0166] Among them, norm 0 and norm l represent the normalization operation, which is used to ensure that the feature vectors have the same scale, thus avoiding some features dominating other features due to overly large numerical values.

[0167] Specifically, the normalization operation includes layer normalization and batch normalization. Layer normalization means normalizing the features of each sample so that the mean of each feature is 0 and the variance is 1; batch normalization means normalizing the data of the entire batch.

[0168] Furthermore, according to the matrix algorithm and scaling process, the linear transformation matrices and for calculating the key vectors and the linear transformation matrix W Q for calculating the query vector are computed, and the Softmax function is introduced to calculate the attention weights, which are used for subsequent matrix multiplication processing of the linear transformation matrices and for calculating the value vectors and the calculation results input to the Softmax function. Finally, the processing results are sequentially input into the linear layer and the normalization layer to obtain the multimodal representation output by the MMA module, effectively aligning and fusing the information from different modalities, thereby improving the performance of the model in multimodal tasks.

[0169] In this embodiment, personalized and accurate prevention and control suggestions are generated based on the recognition results, and based on the LLaMA-2-13b-chat model, through transfer learning and fine-tuning, it can quickly adapt to new text data and the needs of specific fields, improving the adaptability of the multi-modal model for pest and disease control to new information.

[0170] As Figure 7 shown, some embodiments of the present application provide a multi-modal system for pest and disease control based on the visual pixel level. The system includes an acquisition module 710 and an identification module 720. Specifically:

[0171] The acquisition module 710 is used to acquire pest and disease images and their corresponding image texts;

[0172] The identification module 720 is used to input the pest and disease images into a pre-trained multi-modal model for pest and disease control to obtain the recognition results output by the multi-modal model for pest and disease control;

[0173] In some embodiments, the identification module 720 may include: segmenting the pest and disease images into multiple initial images.

[0174] In some embodiments, the identification module 720 may include: inputting the first initial image into the encoding layer to obtain a first encoded image; the first initial image is any one of a part of the initial images randomly selected from the multiple initial images.

[0175] In some embodiments, the identification module 720 may include: occluding the second initial image and converting it into a unified embedding representation to obtain a second encoded image; the second initial image is any one of the multiple initial images other than the first initial image.

[0176] In some embodiments, the identification module 720 may include: splicing the first encoded image and the second encoded image according to the position embedding of the initial images to extract the image features of the pest and disease images from the spliced image.

[0177] In some embodiments, the identification module 720 may include: inputting the image text into the text encoding layer to obtain the text features of the image text.

[0178] In some embodiments, the identification module 720 may include: obtaining the recognition results of the pest and disease images through the image features and the text features.

[0179] In some embodiments, the identification module 720 may include: segmenting the pest and disease training images into multiple initial training images.

[0180] In some embodiments, the recognition module 720 may include: inputting a first initial training image into an encoding layer to obtain a first encoded training image; the first initial training image is any one of a part of the initial training images randomly selected from a plurality of initial training images.

[0181] In some embodiments, the recognition module 720 may include: occluding a second initial training image and converting it into a unified embedding representation to obtain a second encoded training image; the second initial training image is any one of the initial training images other than the first initial training image among the plurality of initial training images.

[0182] In some embodiments, the recognition module 720 may include: inputting the first encoded training image and the second encoded training image into a decoding layer in the order of the first initial training image to obtain a restored training image.

[0183] In some embodiments, the recognition module 720 may include: calculating a loss function according to the region corresponding to the second encoded training image in the second encoded training image and the restored training image to optimize the pest control multimodal model.

[0184] In some embodiments, the recognition module 720 may include: inputting the image text into the LLaMA-2-13b-chat large language model to obtain the text features of the image text output by the LLaMA-2-13b-chat large language model.

[0185] In some embodiments, the recognition module 720 may include: obtaining the modal features of the pest image according to the image features and the text features.

[0186] In some embodiments, the recognition module 720 may include: inputting the modal features into a linear layer to generate the recognition result of the pest image.

[0187] In some embodiments, the recognition module 720 may include: mapping the image features and the text features to a preset shared semantic space to obtain visual language features.

[0188] In some embodiments, the recognition module 720 may include: performing modal separation on the visual language features to obtain modal features.

[0189] In some embodiments, the recognition module 720 may include: inputting the visual language features into a modal adaptation layer to perform modal separation operations on the visual language features.

[0190] In some embodiments, the recognition module 720 may include:

[0191] φ(X,M,m)=X⊙| {M=m} ;

[0192]

[0193] n ∈ [1, n];

[0194] Wherein, X is the visual language feature, M is the modality indicator of the modality adaptation layer, m is the modality indication value, and ⊙ is the element-wise multiplication operation. is the query vector. is the hidden state of the previous layer given by the decoding layer, W l Q is the linear transformation matrix for calculating the query vector. is the key vector. is the feature with modality indication value 0. is the feature with modality indication value n. and is the linear transformation matrix for calculating the key vector. is the value vector. and is the linear transformation matrix for calculating the value vector, and n is the number of preset language encoder layers.

[0195] In some embodiments, the recognition module 720 may include: constructing a pest and disease resource database.

[0196] In some embodiments, the recognition module 720 may include: generating prevention and control suggestions corresponding to the recognition result of the pest and disease image according to the pest and disease resource database and the recognition result of the pest and disease image.

[0197] It should be noted that the multi-modal system for pest and disease control based on visual pixel level provided in this embodiment and the above-mentioned multi-modal method for pest and disease control based on visual pixel level are based on the same inventive concept. Therefore, the relevant content of the above-mentioned multi-modal method for pest and disease control based on visual pixel level is also applicable to the content of the multi-modal system for pest and disease control based on visual pixel level. Therefore, it will not be elaborated here.

[0198] To solve the technical problems of low accuracy and poor generalization ability in identifying pests and diseases in the prior art, the system divides the pest and disease image into multiple initial images; inputs the first initial image into the encoding layer to obtain the first encoded image; the first initial image is any one of a part of the initial images randomly selected from the multiple initial images; occludes the second initial image and converts it into a unified embedded representation to obtain the second encoded image; the second initial image is any one of the multiple initial images other than the first initial image; splices the first encoded image and the second encoded image according to the position embedding of the initial image to extract the image features of the pest and disease image from the spliced image; inputs the image text into the text encoding layer to obtain the text features of the image text; obtains the recognition result of the pest and disease image through the image features and the text features. In this way, it is possible to combine the text features and the image features to improve the multi-modal accuracy and generalization ability of pest and disease control based on visual pixel level, and then give prevention and control suggestions according to the recognition result to improve the user experience.

[0199] An embodiment of the present application further provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above multi-modal method for pest and disease control based on visual pixel level.

[0200] As Figure 8 , Figure 8 is a schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present application. The electronic device includes:

[0201] At least one battery;

[0202] At least one memory;

[0203] At least one processor;

[0204] At least one program;

[0205] The program is stored in the memory, and the processor executes at least one program to implement the above multi-modal method for pest and disease control based on visual pixel level disclosed in the present application.

[0206] The electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), an in-vehicle computer, etc.

[0207] The following provides a detailed introduction to the electronic device of the embodiment of the present application.

[0208] The processor 1600 can be implemented in the form of a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present disclosure;

[0209] The memory 1700 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1700 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1700 and are called by the processor 1600 to execute a multi-modal method for pest control based on visual pixel level provided by the embodiments of the present disclosure.

[0210] The input / output interface 1800 is used to implement information input and output;

[0211] The communication interface 1900 is used to implement communication interaction between this device and other devices, and can implement communication through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0212] The bus 2000 transmits information between various components of the device (such as the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900);

[0213] Among them, the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900 are communicatively connected to each other inside the device through the bus 2000.

[0214] The embodiments of the present disclosure also provide a storage medium, which is a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute the above-mentioned multi-modal method for pest control based on visual pixel level.

[0215] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include memories remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0216] The embodiments described in the embodiments of the present disclosure are for more clearly illustrating the technical solutions of the embodiments of the present disclosure, and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present disclosure are equally applicable to similar technical problems.

[0217] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present disclosure, and may include more or fewer steps than shown in the figures, or combine certain steps, or different steps.

[0218] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0219] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0220] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of this application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0221] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression refers to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0222] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0223] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0224] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0225] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing an electronic device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0226] The above has specifically described the preferred implementation of the embodiments of this application, but the embodiments of this application are not limited to the above implementation manners. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the embodiments of this application, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the embodiments of this application.

[0227] The above has described the embodiments of this application in detail with reference to the accompanying drawings, but this application is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art to which it pertains, various changes can also be made without departing from the purpose of this application.

Claims

1. A multimodal method for pest control based on visual pixel level, characterized in that: The method comprises: Obtain pest and disease images and their corresponding image texts; Inputting the pest image into a pre-trained pest control multimodal model to obtain a recognition result output by the pest control multimodal model; The multimodal model for pest control outputs recognition results including: Segmenting the pest image into a plurality of initial images; Inputting a first initial image into a coding layer to obtain a first coded image; the first initial image is any one of a portion of the initial images randomly selected from the plurality of initial images; Masking a second initial image and converting it into a unified embedding representation to obtain a second encoded image; the second initial image is any one of the multiple initial images except the first initial image; splicing the first coded image and the second coded image according to the position embedding of the initial image, so as to extract image features of the pest image from the spliced ​​image; Inputting the image text into a text encoding layer to obtain text features of the image text; The recognition result of the pest image is obtained through the image features and the text features.

2. The multimodal method for pest control based on visual pixel level according to claim 1, characterized in that: The training process of the multimodal model for pest control includes: Segment the pest and disease training image into multiple initial training images; Inputting a first initial training image into the encoding layer to obtain a first encoded training image; the first initial training image is any one of a portion of the initial training images randomly selected from the multiple initial training images; Masking a second initial training image and converting it into a unified embedding representation to obtain a second encoded training image; the second initial training image is any one of the multiple initial training images except the first initial training image; Inputting the first encoded training image and the second encoded training image into a decoding layer in the order of the first initial training image to obtain a restored training image; A loss function is calculated based on the second encoded training image and the area in the restored training image corresponding to the second encoded training image to optimize the multimodal model for disease and insect pest control.

3. The multimodal method for pest control based on visual pixel level according to claim 1, characterized in that: The text encoding layer is the LLaMA-2-13b-chat large language model; The step of inputting the image text into a text encoding layer to obtain text features of the image text includes: The image text is input into the LLaMA-2-13b-chat large language model to obtain text features of the image text output by the LLaMA-2-13b-chat large language model.

4. The multimodal method for pest control based on visual pixel level according to claim 1, characterized in that: The obtaining of the recognition result of the pest image by using the image feature and the text feature includes: Obtaining modal features of the pest image according to the image features and the text features; The modal features are input into the linear layer to generate the recognition result of the pest image.

5. The multimodal method for pest control based on visual pixel level according to claim 4, characterized in that: The obtaining of the modal features of the pest image according to the image features and the text features includes: Mapping the image features and the text features to a preset shared semantic space to obtain visual language features; The visual language features are modally separated to obtain the modal features.

6. The multimodal method for pest control based on visual pixel level according to claim 5, characterized in that: The modality separation of the visual language features includes: Inputting the visual language features into a modality adaptation layer to perform a modality separation operation on the visual language features; The modal separation operation includes: φ(X,M,m)=X⊙| {M=m} ; n∈[1,n]; Among them, X is the visual language feature, M is the modality indicator of the modality adaptation layer, m is the modality indicator value, ⊙ is the element-by-element multiplication operation, is the query vector, Given the hidden state of the previous layer to the decoding layer, To calculate the linear transformation matrix of the query vector, is the key vector, is a feature with a modal indicator value of 0, is the feature with modal indicator value n, and To calculate the linear transformation matrix of the key vector, is a value vector, and is the linear transformation matrix for calculating the value vector, and n is the number of preset language encoder layers.

7. The multimodal method for pest control based on visual pixel level according to claim 1, characterized in that: After obtaining the recognition result of the pest image through the image feature and the text feature, the method further includes: Build a pest and disease resource database; Based on the pest resource database and the recognition result of the pest image, a prevention and control suggestion corresponding to the recognition result of the pest image is generated.

8. A multimodal system for pest control based on visual pixel level, characterized in that: The system comprises: An acquisition module, used to acquire pest and disease images and their corresponding image texts; A recognition module, used for inputting the pest image into a pre-trained pest control multimodal model to obtain a recognition result output by the pest control multimodal model; The multimodal model for pest control outputs recognition results including: Segmenting the pest image into a plurality of initial images; Inputting a first initial image into a coding layer to obtain a first coded image; the first initial image is any one of a portion of the initial images randomly selected from the plurality of initial images; Masking a second initial image and converting it into a unified embedding representation to obtain a second encoded image; the second initial image is any one of the multiple initial images except the first initial image; splicing the first coded image and the second coded image according to the position embedding of the initial image, so as to extract image features of the pest image from the spliced ​​image; Inputting the image text into a text encoding layer to obtain text features of the image text; The recognition result of the pest image is obtained through the image features and the text features.

9. An electronic device, characterized in that: It includes at least one control processor and a memory for communicating with the at least one control processor; the memory stores instructions that can be executed by the at least one control processor, and the instructions are executed by the at least one control processor so that the at least one control processor can execute the multimodal method for pest control based on visual pixel level as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute a multimodal method for pest control based on visual pixel level as described in any one of claims 1 to 7.