A multi-modal method based on CLIP for realizing an intelligent recognition system for crops

Through the CLIP model, the image, text and environmental data of crops are fused, and the problem of low accuracy in pest identification in the prior art under complex environments is solved, and efficient and accurate pest identification and monitoring are achieved.

CN119580049BActive Publication Date: 2025-07-25ANHUI SAIDA TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510015098.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-07-25
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

The existing crop pest and disease identification methods have low accuracy in complex environments and insufficient fusion utilization of multimodal data, which limits the practicality and reliability of the system.

Method used

Using a multimodal method based on CLIP, we collect real-time images, text descriptions and environmental data of crops, use the CLIP model to extract images and text features, combine environmental data to generate multimodal feature vectors, and use a pre-trained classifier for identification.

Benefits of technology

It significantly improves the accuracy and robustness of crop pest identification, can maintain high recognition stability in complex environments, achieve real-time monitoring and rapid response to large areas of crops, and reduces the damage of pests and diseases to crops.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119580049B_ABST
    Figure CN119580049B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-modal method based on CLIP for implementing an intelligent crop recognition system, belonging to the technical field of crop recognition, specifically including: collecting real-time images, text descriptions, and environmental data of crops; performing image processing on the real-time images; using the CLIP model to extract features from the images and text descriptions, extracting visual features from the images and semantic features from the text, the text description including the types of crops and descriptions of the symptoms of pests and diseases, fusing the environmental data with the visual features to generate a multi-modal feature vector; using a pre-trained classifier to classify the multi-modal feature vector, identifying the specific types of crops and types of pests and diseases according to the extracted multi-modal features, displaying the recognition results to the user through a visualization interface, and feeding back the information to the monitoring center; the present invention significantly improves the accuracy of crop pest and disease recognition by fusing visual, semantic, and environmental features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of crop recognition, and specifically relates to a multi-modal method based on CLIP for realizing an intelligent crop recognition system. Background Art

[0002] With the development of modern agriculture, high yield and quality of crops have become important goals of agricultural production. However, the occurrence of pests and diseases poses a serious threat to crop yield and quality. Traditional pest and disease monitoring mainly relies on manual inspections, which are inefficient and easily affected by subjective factors.

[0003] In recent years, with the development of computer vision and artificial intelligence technologies, using high-point monitoring video images combined with image algorithms for crop pest and disease recognition has become a trend. However, when dealing with images in complex environments, the accuracy of existing methods still needs to be improved, and the fusion and utilization of multi-modal data are insufficient, which limits the practicability and reliability of the system. Summary of the Invention

[0004] The purpose of the present invention is to provide a multi-modal method based on CLIP for realizing an intelligent crop recognition system, and solve the following technical problems:

[0005] When dealing with images in complex environments, the accuracy of existing methods still needs to be improved, and the fusion and utilization of multi-modal data are insufficient, which limits the practicability and reliability of the system.

[0006] The purpose of the present invention can be achieved through the following technical solutions:

[0007] An intelligent crop recognition system implemented by a multi-modal method based on CLIP, comprising:

[0008] A data acquisition module for collecting real-time images, text descriptions, and environmental data of crops, where the text description includes the type of crops and the symptom description of pests and diseases;

[0009] An image processing module for performing noise reduction processing on the real-time image, adjusting the contrast and brightness of the real-time image to a set value, and distinguishing the crop area from the background area in the real-time image based on image segmentation technology;

[0010] A feature extraction module for extracting features from the real-time image and text description based on the CLIP model, extracting visual features from the real-time image, extracting semantic features from the text description, and fusing the environmental data with the visual features to generate a multi-modal feature vector;

[0011] A feature classification module for classifying the multi-modal feature vector using a pre-trained classifier, and identifying the type of crops and the corresponding type of pests and diseases according to the extracted multi-modal features;

[0012] The result output module is used to display the recognition result to the user through a visualization interface and feedback the information to the monitoring center.

[0013] As a further solution of the present invention: in the data acquisition module, the environmental data includes but is not limited to temperature, humidity, and soil nutrient concentration.

[0014] As a further solution of the present invention: in the image processing module, the process of noise reduction is as follows:

[0015] Convert the real-time image into a grayscale image, adjust the size of the grayscale image to a set specification, then perform wavelet transform on the grayscale, decompose the grayscale image into several scales, the grayscale image is decomposed into several wavelet coefficients, and the wavelet coefficients represent the local features of the image at different scales. Perform threshold processing on the wavelet coefficients of the high-frequency subbands, suppress the noise coefficients by setting a threshold, retain the signal coefficients, and after threshold processing, reconstruct the image based on the retained signal coefficients and wavelet coefficients. Through inverse wavelet transform, convert the processed coefficients back to the spatial domain to obtain the denoised grayscale image and the real-time image.

[0016] As a further solution of the present invention: in the feature extraction module, the specific process of the CLIP model for feature extraction is as follows:

[0017] The CLIP model includes an image encoder and a text encoder. The image encoder is based on the ResNet network. The image encoder converts the input real-time image into a visual feature vector and extracts the visual features in the visual feature vector. The text encoder is based on the Transformer structure. The text encoder is used to convert the input text description into a text feature vector and capture the semantic features in the text feature vector.

[0018] Mark the pre-paired visual feature vector and text feature vector as positive sample pairs, and the unpaired ones as negative sample pairs. The CLIP model embeds the visual feature vector and the text feature vector into a shared semantic space. In the semantic space, calculate the similarity between the visual feature vector and the text feature vector, and adjust the model parameters according to the calculation result so that the similarity of the positive sample pairs approaches 1 and the similarity of the negative sample pairs approaches 0. The CLIP model uses a symmetric loss function, and the symmetric loss function consists of two parts of cross-entropy, calculates the corresponding loss from the row direction and the column direction respectively, and finally takes the average of the losses in the two directions.

[0019] As a further solution of the present invention: for each real-time image, the model calculates the similarity between its visual feature vector and all text feature vectors, and converts all similarities into a probability distribution through the softmax function; for each text description, the model calculates the similarity between its text feature vector and all visual feature vectors, and minimizes the inner product with irrelevant features. During the training process, the parameters are continuously adjusted to minimize the value of the loss function until convergence.

[0020] As a further solution of the present invention: the data acquisition module is also used for unit conversion and normalization processing of all environmental data. The formula for normalization processing is x' = (x - min(x)) / (max(x) - min(x)), where x' represents the normalized data and x represents the original data.

[0021] As a further solution of the present invention: the data acquisition module is also used for duplicate removal processing of the text description. For any pair of text descriptions with different dates but the same number of words, they are respectively marked as the first text and the second text. The occurrence frequencies of the preset keywords of the text description are respectively collected, and the number of the top three keywords with the highest occurrence frequencies is collected. The number of occurrences of the first three keywords of the first text is marked as n1, n2, and n3, and the number of occurrences of the first three keywords of the second text is marked as m1, m2, and m3. The three values are respectively used as the upper base, lower base, and height of the trapezoid, and the first trapezoid and the second trapezoid are respectively generated. Calculate the area S1 of the first trapezoid = (n1 + n2)n3 / 2, calculate the area S2 of the first trapezoid = (m1 + m2)m3 / 2. The first trapezoid and the second trapezoid are superposed in the plane, and the superposition state with the largest superposition area is selected. The area S3 of the superposition area in the current superposition state is obtained. Calculate the value of S3 / (S1 + S2), and this value is the shape similarity. Two text descriptions with a similarity greater than 90% are marked as duplicate texts, and the text description with the later date is deleted.

[0022] As a further solution of the present invention: the process of the feature extraction module for fusing features is as follows:

[0023] Use an encoder to extract features from environmental data, and extract the numerical features of environmental data based on the attention mechanism; map them into vector form in the multi-modal space, and perform weighted fusion on the extracted visual features and numerical features, assign different weights to the visual features and numerical features, and then perform weighted summation to obtain the comprehensive features.

[0024] As a further solution of the present invention: the pre-training process of the classifier is as follows:

[0025] Label the marked image data, text data, and environmental data, clean, denoise, and standardize the collected data. Use a convolutional neural network as a classifier, and select a model trained on a large-scale dataset as the base model. Freeze the convolutional layers of the base model and only train the newly added classifier. After the classifier is trained, unfreeze the convolutional layers of the base model and continue to train the entire model based on the collected data. Train the model using the labeled dataset and evaluate the model through evaluation metrics, including accuracy, recall, and F1-score. During the training process, continuously adjust the parameters to minimize the value of the loss function until the model converges.

[0026] Advantages of the present invention:

[0027] (1) By fusing visual, semantic, and environmental features, the present invention significantly improves the accuracy of crop pest and disease identification. Compared with the prior art, the present invention introduces the CLIP model, which can process information described by both images and natural languages simultaneously. Through contrastive learning of visual features and text semantics, the CLIP model solves the limitation of traditional methods that rely only on single-modal data in complex environments. For example, traditional pest and disease identification technologies usually rely only on image data, which has poor identification effects in complex scenarios such as light changes, weather impacts, or crop occlusion. By combining image visual features with text semantic features, the present invention can rely on semantic-level supplementation when image information is insufficient, enhancing the model's understanding ability. At the same time, the introduction of environmental data enables the system to consider multiple factors when analyzing the occurrence of crop pests and diseases, thus making more accurate judgments. This innovation of multi-modal data fusion significantly improves the accuracy and robustness of identification, enabling the system to operate stably in various complex environments.

[0028] (2) Through the fusion of multi-modal features, the system can still maintain high recognition stability when facing different environmental conditions, such as light intensity, weather changes, and crop occlusion. Current crop pest and disease identification systems perform poorly in complex environments, especially the changes in light and weather conditions are likely to lead to a significant decrease in recognition accuracy. In the prior art, many systems often make misjudgments or missed judgments due to the degradation of image quality when facing these changes. However, by introducing multi-modal feature extraction and contrastive learning of the CLIP model, the present invention not only relies on image quality but also effectively overcomes the influence of these adverse factors by fusing environmental data and text descriptions of crops. The system can maintain high recognition accuracy in different light, weather, and crop growth cycles. Especially in low-light or adverse weather conditions where traditional methods are prone to errors, the system of the present invention can still maintain reliable performance. Therefore, the present invention has a significant improvement in system robustness and adaptability and can adapt to various complex farmland environments.

[0029] (3) By combining high - point monitoring devices with efficient algorithms, the present invention can achieve real - time monitoring and rapid response for large - area crops. Compared with traditional manual inspections or crop monitoring methods that rely on a single image recognition system, the system of the present invention can continuously obtain high - definition images of a large - scale farmland and perform real - time analysis and feedback in combination with environmental data. In the prior art, the response speed of real - time monitoring systems is often limited by computing resources and algorithm efficiency, and it is impossible to continuously and accurately monitor pests and diseases in large - scale crop areas. However, the present invention optimizes the data collection and processing process, efficiently extracts multi - modal features using the CLIP model, and quickly classifies the types of pests and diseases through a pre - trained classifier, greatly improving the real - time performance of the system. In addition, the system can also timely feedback the recognition results to the monitoring center, helping agricultural managers quickly take measures to reduce the damage of pests and diseases to crops. This real - time monitoring and response ability not only improves the efficiency of crop management but also reduces crop losses caused by delays. Brief Description of the Drawings

[0030] The present invention will be further described below in conjunction with the accompanying drawings.

[0031] Figure 1 It is a flow diagram of the present invention. Detailed Embodiments

[0032] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention.

[0033] Please refer to Figure 1 As shown, the present invention is a multi - modal method - based intelligent crop recognition system using CLIP, specifically including:

[0034] 1. Data collection

[0035] The present invention first obtains high - definition images of crops in real - time through high - point monitoring devices and synchronously collects relevant environmental data. The environmental data includes, but is not limited to, temperature, humidity, soil nutrient concentration, etc. These data provide important context information for subsequent pest and disease identification. Through the deployment of high - point monitoring devices, a large - area farmland can be covered to monitor the growth status of crops and the occurrence of pests and diseases in real - time. The collection of high - definition images provides high - quality raw data for subsequent image processing and feature extraction, while the environmental data provides a more comprehensive reference basis for pest and disease identification.

[0036] 2. Image pre - processing

[0037] Due to the complex farmland environment, the acquired images may be interfered by factors such as weather, light, and noise. Therefore, the preprocessing step is very necessary. The image preprocessing module performs operations such as denoising, enhancement, and segmentation on the acquired images. First, the denoising process can effectively reduce the random noise in the images, such as factors affecting uneven illumination and shadows. Second, the enhancement process improves the contrast and brightness of the images, making the characteristics of pests and diseases more prominent. Finally, the image segmentation technology can distinguish the crop regions from the background regions in the images, further improving the accuracy of feature extraction. This process greatly enhances the effectiveness and accuracy of subsequent feature extraction.

[0038] 3. Multi-modal Feature Extraction

[0039] After the image preprocessing is completed, the system enters the multi-modal feature extraction stage. The core innovation of this stage lies in using the CLIP model to extract features from images and text descriptions. The CLIP model is a pre-trained multi-modal neural network that can process image and text information simultaneously. Specifically, the CLIP model extracts visual features from images and semantic features from text. The text descriptions can include semantic information such as the types of crops and descriptions of the symptoms of pests and diseases. These text information form the basis for contrastive learning with the visual features in the images.

[0040] In addition to visual and semantic features, the present invention also fuses environmental data (such as temperature, humidity, soil nutrient concentration, etc.) with visual features to form a comprehensive multi-modal feature vector. This way of fusing multi-modal features greatly improves the accuracy of pest and disease identification. Environmental data is closely related to the occurrence of crop pests and diseases. With the help of this data, the system can better understand the occurrence conditions of pests and diseases and make more accurate identification and diagnosis.

[0041] 4. Feature Matching and Classification

[0042] After the multi-modal feature extraction is completed, the system uses a pre-trained classifier to classify the multi-modal feature vector. The classifiers in the present invention can support algorithms such as support vector machines (SVM), random forests, or deep neural networks. The role of the classifier is to identify the specific types of crops and their pest and disease types based on the extracted multi-modal features.

[0043] The present invention also introduces the method of contrastive learning to further optimize the feature representation. The core idea of contrastive learning is to optimize the feature extraction ability of the model by comparing different samples. In the technical solution of the present invention, contrastive learning trains on different images and descriptions of the same crop, enabling the model to better capture the correlation between images and texts, thereby improving the recognition accuracy. Especially in complex environments such as lighting changes and crop occlusion, feature optimization is particularly important.

[0044] 5. Result Output and Feedback

[0045] After the classification is completed, the system displays the recognition results to the user through a visualization interface. The visualization interface can intuitively display the types of crops, the types of pests and diseases, and the corresponding environmental data. Users can understand the real-time situation of the farmland through the interface and then make corresponding management decisions. In addition, the system also feeds back the recognition information to the monitoring center to achieve intelligent crop management. The monitoring center can take corresponding measures in a timely manner according to the recognition results, such as spraying pesticides and adjusting water content, so as to reduce the harm of pests and diseases to crops.

[0046] The core of the present invention lies in the adoption of the CLIP model for multi-modal feature extraction. Through contrastive learning of visual features and text semantic features, combined with additional information such as temperature, humidity, and soil nutrient concentration in the farmland environment, a comprehensive multi-modal feature vector is formed. This fusion method greatly improves the recognition ability of the system in complex and changing environments. Especially in the case of uneven lighting, many occlusions, or poor weather conditions, it can still accurately identify the types of pests and diseases of crops.

[0047] Application of the CLIP model: CLIP learns the correspondence between vision and language through joint training of images and texts. When applied to the identification of crop pests and diseases, the system can compare crop images with corresponding semantic descriptions to improve the understanding ability of objects in the images. Compared with traditional image classification methods, CLIP can not only identify visual features in images but also combine semantic information for deeper understanding, enabling the system to perform well in a wider range of scenarios.

[0048] Introduction of environmental data: The present invention particularly considers the introduction of agricultural environmental data such as temperature, humidity, and soil data. These data are combined with visual features to provide additional context information for the identification of crop pests and diseases. For example, the occurrence of certain pests and diseases is closely related to specific climate conditions. The introduction of environmental data can help the system better predict and identify these pests and diseases, further improving the practicality and accuracy of the system.

[0049] Contrastive learning optimization: To further improve the recognition accuracy, the present invention also adopts the method of contrastive learning. By comparing the feature representations in different modalities, the representation effect of the feature vector is optimized. This method can reduce error accumulation, especially when dealing with multi-modal data fusion, and further improve the accuracy and robustness of the classifier.

[0050] In another preferred embodiment of the present invention, in the data acquisition module, the environmental data includes but is not limited to temperature, humidity, and soil nutrient concentration.

[0051] In another preferred embodiment of the present invention, in the image processing module, the process of noise reduction is as follows:

[0052] 1. Conversion to grayscale image:

[0053] First, convert the color image obtained in real time into a grayscale image. This conversion simplifies subsequent processing steps and reduces the computational amount by removing color information and only retaining luminance information.

[0054] The grayscale processing can be achieved by various methods, such as taking the average value of the RGB three channels or weighted average.

[0055] 2. Adjust image size:

[0056] According to actual needs, adjust the size of the grayscale image to a preset specification. This step is usually to ensure that the image meets the input requirements of a specific algorithm or model, or to optimize storage and transmission efficiency.

[0057] The size adjustment can be achieved through techniques such as interpolation, cropping, or scaling.

[0058] 3. Wavelet transform:

[0059] Perform wavelet transform on the adjusted grayscale image. This is a multi-scale analysis tool that can decompose the image into components of different frequencies.

[0060] The result of wavelet transform is to represent the original image as a series of wavelet coefficients, which represent the local features of the image at different scales.

[0061] 4. Threshold processing:

[0062] Among the wavelet coefficients obtained by wavelet transform, the wavelet coefficients in the high-frequency subbands are often related to the details and noise of the image. Therefore, threshold processing of these high-frequency coefficients is a key step in denoising.

[0063] Set a suitable threshold to suppress the noise coefficients below the threshold and retain the signal coefficients above the threshold. This can effectively remove noise while retaining the important features of the image.

[0064] 5. Image Reconstruction:

[0065] Based on the signal coefficients and wavelet coefficients retained after threshold processing, image reconstruction is performed. This step is achieved through inverse wavelet transform, converting the processed coefficients back to the spatial domain to form a denoised grayscale image.

[0066] Compared with the original image, the noise in the reconstructed grayscale image is significantly suppressed, while more useful information is retained.

[0067] 6. Output Result:

[0068] The finally obtained denoised grayscale image is converted into a real-time image after noise reduction, which can be used for subsequent analysis, recognition, or other processing tasks.

[0069] In another preferred embodiment of the present invention, in the feature extraction module, the specific process of the CLIP model for extracting features is as follows:

[0070] The CLIP model includes an image encoder and a text encoder. The image encoder is based on the ResNet network. The image encoder converts the input real-time image into a visual feature vector and extracts visual features from the visual feature vector. The text encoder is based on the Transformer structure. The text encoder is used to convert the input text description into a text feature vector and capture semantic features in the text from the text feature vector.

[0071] The pre-paired visual feature vector and text feature vector are marked as positive sample pairs, and the unpaired ones are marked as negative sample pairs. The CLIP model embeds the visual feature vector and the text feature vector into a shared semantic space. In the semantic space, the similarity between the visual feature vector and the text feature vector is calculated. According to the calculation result, the model parameters are adjusted so that the similarity of the positive sample pairs approaches 1 and the similarity of the negative sample pairs approaches 0. The CLIP model uses a symmetric loss function, which consists of two parts of cross-entropy. The corresponding loss is calculated from the row direction and the column direction respectively, and finally the average value of the losses in the two directions is taken.

[0072] In another preferred embodiment of the present invention, for each real-time image, the model calculates the similarity between its visual feature vector and all text feature vectors, and converts all similarities into a probability distribution through the softmax function. For each text description, the model calculates the similarity between its text feature vector and all visual feature vectors, and minimizes the inner product with irrelevant features. During the training process, the parameters are continuously adjusted to minimize the value of the loss function until convergence.

[0073] In another preferred embodiment of the present invention, the data acquisition module is further configured to perform unit conversion and normalization processing on all environmental data. The formula for normalization processing is x' = (x - min(x)) / (max(x) - min(x)), where x' represents the normalized data and x represents the original data.

[0074] In another preferred embodiment of the present invention, the data acquisition module is further configured to perform duplicate removal processing on the text description. For any pair of text descriptions with different dates but the same number of words, they are respectively marked as the first text and the second text. The occurrence frequencies of the preset keywords in the text description are respectively collected, and the number of the top three keywords with the highest occurrence frequencies is collected. The numbers of the top three keywords in the first text are marked as n1, n2, and n3, and the numbers of the top three keywords in the second text are marked as m1, m2, and m3. The three values are respectively used as the upper base, lower base, and height of a trapezoid, and the first trapezoid and the second trapezoid are respectively generated. Calculate the area of the first trapezoid S1=(n1 + n2)n3 / 2, calculate the area of the first trapezoid S2=(m1 + m2)m3 / 2, perform plane superposition on the first trapezoid and the second trapezoid, select the superposition state when the superposition area is the largest, obtain the area S3 of the superposition region in the current superposition state, calculate the value of S3 / (S1 + S2), and this value is the shape similarity. Mark two text descriptions with a similarity greater than 90% as duplicate texts, and delete the text description with the later date.

[0075] In another preferred embodiment of the present invention, the process of the feature extraction module for fusing features is as follows:

[0076] Use an encoder to extract features from the environmental data, and extract the numerical features of the environmental data based on the attention mechanism; map them into vector form in the multi-modal space, perform weighted fusion on the extracted visual features and numerical features, assign different weights to the visual features and numerical features, and then perform weighted summation to obtain the comprehensive features.

[0077] In another preferred embodiment of the present invention, the pre-training process of the classifier is as follows:

[0078] 1. Data preparation and preprocessing:

[0079] First, collect a large amount of labeled image data, text data, and environmental data. These data are the basis for model training and need to cover various situations of crop pests and diseases, as well as relevant environmental factors such as temperature, humidity, and soil nutrients.

[0080] Perform cleaning, denoising, and standardization processing on the collected data. This step is to eliminate noise and outliers in the data, improve data quality, and at the same time convert the data into a format suitable for model processing.

[0081] 2. Select the base model:

[0082] Use a convolutional neural network (CNN) as the core architecture of the classifier. CNN has strong feature extraction capabilities in the field of image processing and is suitable for extracting features of crop pests and diseases from images.

[0083] Select a model trained on a large-scale dataset as the base model. Such a model has already learned rich feature representation capabilities and can be used as the initial model for new tasks, accelerating the training process and improving model performance.

[0084] 3. Freeze and fine-tune:

[0085] During the pre-training stage, freeze the convolutional layers of the base model and only train the newly added classifier part. The reason for this is that the convolutional layers of the base model have already learned general feature representations on a large-scale dataset, while the newly added classifier needs to be trained for specific crop pests and diseases.

[0086] After the classifier is trained, unfreeze the convolutional layers of the base model and continue to train the entire model based on the collected data. This step is to enable the entire model to adapt to new tasks and data distributions.

[0087] 4. Model training and evaluation:

[0088] Use the labeled dataset to train the model. During the training process, continuously adjust the model parameters through the backpropagation algorithm to minimize the value of the loss function.

[0089] Use evaluation metrics such as accuracy, recall, and F1-score to evaluate the model. These metrics can comprehensively reflect the performance of the model, including the accuracy, completeness, and comprehensive performance of recognition.

[0090] 5. Parameter adjustment and optimization:

[0091] During the training process, continuously adjust the parameter settings according to the performance of the model, such as the learning rate, batch size, number of iterations, etc. The adjustment of these parameters is crucial for the convergence speed and final performance of the model.

[0092] Through continuous training and evaluation until the model converges. Convergence means that the model has learned sufficient feature representations and can achieve stable and good performance on new data.

[0093] The above has described in detail an embodiment of the present invention, but the above content is only a preferred embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. Any equivalent changes and improvements made within the scope of the application of the present invention should still fall within the scope covered by the patent of the present invention.

Claims

1. A multi-modal method based on CLIP for implementing a crop intelligent recognition system, characterized in that, Including: A data acquisition module for acquiring real-time images, text descriptions, and environmental data of crops. The text description includes the type of crops and the symptom description of pests and diseases; An image processing module for performing noise reduction processing on the real-time image, adjusting the contrast and brightness of the real-time image to set values, and distinguishing the crop area and the background area in the real-time image based on image segmentation technology; A feature extraction module for extracting features from the real-time image and text description based on the CLIP model, extracting visual features from the real-time image, extracting semantic features from the text description, and fusing the environmental data with the visual features to generate a multi-modal feature vector; A feature classification module for classifying the multi-modal feature vector using a pre-trained classifier and identifying the type of crops and the corresponding type of pests and diseases according to the extracted multi-modal features; A result output module for displaying the recognition result to the user through a visualization interface and feeding back the information to the monitoring center; The data acquisition module is further configured to perform duplicate removal processing on the text description. For any pair of text descriptions with different dates but the same number of words, they are respectively marked as the first text and the second text. The occurrence frequencies of preset keywords in the text description are respectively collected, and the number of the top three keywords with the highest occurrence frequencies is collected. The number of the first three keywords in the first text is marked as n1, n2, and n3, and the number of the first three keywords in the second text is marked as m1, m2, and m3. The three values are respectively used as the upper base, lower base, and height of a trapezoid, and the first trapezoid and the second trapezoid are respectively generated. Calculate the area S1 of the first trapezoid = (n1 + n2)n3 / 2, calculate the area S2 of the first trapezoid = (m1 + m2)m3 / 2, perform planar superposition on the first trapezoid and the second trapezoid, select the superposition state with the largest superposition area, obtain the area S3 of the superposition area in the current superposition state, calculate the value of S3 / (S1 + S2), and this value is the shape similarity. Mark two text descriptions with a similarity greater than 90% as duplicate texts, and delete the text description with the later date among them.

2. The multi-modal method based on CLIP for implementing a crop intelligent recognition system according to claim 1, characterized in that, In the data acquisition module, the environmental data includes but is not limited to temperature, humidity, and soil nutrient concentration.

3. A multi-modal method based on CLIP for implementing a crop intelligent recognition system according to claim 1, characterized in that, In the image processing module, the process of noise reduction processing is as follows: Convert the real-time image into a grayscale image, adjust the size of the grayscale image to a set specification, then perform wavelet transform on the grayscale, decompose the grayscale image into several scales. The grayscale image is decomposed into several wavelet coefficients, and the wavelet coefficients represent the local features of the image at different scales. Perform threshold processing on the wavelet coefficients of the high-frequency subbands, suppress the noise coefficients through the set threshold, and retain the signal coefficients. After threshold processing, reconstruct the image based on the retained signal coefficients and wavelet coefficients. Through inverse wavelet transform, convert the processed coefficients back to the spatial domain to obtain the denoised grayscale image and the real-time image.

4. A multi-modal method based on CLIP for implementing a crop intelligent recognition system according to claim 1, characterized in that, In the feature extraction module, the specific process of the CLIP model for extracting features is as follows: The CLIP model includes an image encoder and a text encoder. The image encoder is based on the ResNet network. The image encoder converts the input real-time image into a visual feature vector and extracts visual features from the visual feature vector. The text encoder is based on the Transformer structure. The text encoder is used to convert the input text description into a text feature vector and capture semantic features in the text from the text feature vector. The pre-paired visual feature vectors and text feature vectors are marked as positive sample pairs, and the unpaired ones are marked as negative sample pairs. The CLIP model embeds the visual feature vectors and text feature vectors into a shared semantic space. In the semantic space, the similarity between the visual feature vector and the text feature vector is calculated. According to the calculation results, the model parameters are adjusted so that the similarity of the positive sample pairs approaches 1 and the similarity of the negative sample pairs approaches 0. The CLIP model uses a symmetric loss function, which consists of two parts of cross-entropy. The corresponding loss is calculated from the row direction and the column direction respectively, and finally the average value of the losses in the two directions is taken.

5. A crop intelligent recognition system implemented by a multimodal method based on CLIP according to claim 1, characterized in that, For each real-time image, the model calculates the similarity between its visual feature vector and all text feature vectors, and converts all similarities into a probability distribution through the softmax function. For each text description, the model calculates the similarity between its text feature vector and all visual feature vectors, and minimizes the inner product with irrelevant features. During the training process, the parameters are continuously adjusted to minimize the value of the loss function until convergence.

6. A multi-modal method based on CLIP for implementing an intelligent crop recognition system according to claim 1, characterized in that, The data acquisition module is also used to perform unit conversion and normalization processing on all environmental data. The formula for normalization processing is x' = (x - min(x)) / (max(x) - min(x)), where x' represents the normalized data and x represents the original data.

7. A multi-modal method based on CLIP for implementing a crop intelligent recognition system according to claim 1, characterized in that, For the feature extraction module, the process of fusing features is as follows: Use an encoder to extract features from environmental data, and extract numerical features of environmental data based on the attention mechanism. Map it into a vector form in the multi-modal space, and perform weighted fusion on the extracted visual features and numerical features. Different weights are assigned to the visual features and numerical features, and then the weighted sum is obtained to get the comprehensive features.

8. A multi-modal method based on CLIP for implementing a crop intelligent recognition system according to claim 1, characterized in that, The pre-training process of the classifier is as follows: Mark the labeled image data, text data, and environmental data, clean, denoise, and standardize the collected data. Use a convolutional neural network as the classifier, and select a model trained on a large-scale dataset as the base model. Freeze the convolutional layer of the base model and only train the newly added classifier. When the classifier training is completed, unfreeze the convolutional layer of the base model and continue to train the entire model based on the collected data. Use the labeled dataset to train the model, and evaluate the model through evaluation metrics, including accuracy, recall rate, and F1 score. During the training process, continuously adjust the parameters to minimize the value of the loss function until the model converges.