Cross-modal lycium barbarum pest retrieval model based on image and text hash feature learning

Through the cross-modal wolfberry pest retrieval model based on images and text, deeply integrating image and text information, the problem of difficult access to wolfberry pest recognition and control information in the existing technology is solved, accurate recognition and comprehensive understanding are achieved, and retrieval accuracy is improved.

CN120196774APending Publication Date: 2025-06-24ZHENGZHOU INST OF TECH +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510253876.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing technology is difficult to accurately and quickly identify wolfberry pest species and provide detailed control information, resulting in inefficient agricultural pest control and environmental pollution problems.

Method used

A cross-modal wolfberry pest retrieval model based on image and text hash feature learning is adopted. Through the image text encoder module, multimodal fusion module, label enhancement module and optimization module, the multimodal hash feature network model is built.

Benefits of technology

It realizes a comprehensive understanding of accurate identification and control information of wolfberry pests, improves retrieval accuracy, avoids the loss of modal information, and ensures that the model can be learned stably and efficiently under complex and diverse data conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196774A_ABST
    Figure CN120196774A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-modal lycium barbarum pest retrieval model based on image and text hash feature learning, and relates to the field of agricultural pest control, comprising an image text encoder module, a multi-modal fusion module, a label enhancement module and an optimization module; the image text encoder module is used for processing images and text data of the wolfberry pests in parallel; the multi-modal fusion module is used for carrying out deep fusion on image and text feature information; the label enhancement module is used for reserving label frequency information and obtaining high-quality supervision information; and the optimization module realizes optimal solution of variables in the model. According to the method, the multi-modal Hash feature learning network model is constructed, image-text cross-modal matching is realized, accurate and rich text description contents can be provided for Chinese wolfberry pest control, and the limitation of a conventional pest image recognition method is broken through, so that farmers and technicians can comprehensively know the pest condition, and the pest control efficiency is improved. And the capability of accurately mastering the characteristics of the Chinese wolfberry pests is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of agricultural pest control, and particularly relates to a cross-modal wolfberry pest retrieval model based on image and text hash feature learning. Background Art

[0002] As a crop with extremely high nutritional and medicinal value, wolfberry is widely used in the field of daily health care. However, wolfberry is vulnerable to pest attacks. Due to its poor pest resistance and being a multi-pest host, the rampant of numerous pests poses a severe challenge to wolfberry production and quality. At present, in order to block the spread of pests and improve the yield and quality of wolfberry, the abuse of pesticides is common, but inevitably causes serious environmental pollution problems. Therefore, accurately and quickly identifying the types of wolfberry pests and formulating detailed and targeted control measures are of great significance for promoting the green and healthy development of the wolfberry industry.

[0003] The traditional method of relying on agricultural experts to identify pests is not only time-consuming and laborious, but also extremely inefficient, making it difficult to achieve large-scale popularization and application. With the rapid development of artificial intelligence technology, the level of agricultural intelligence and precision has been significantly improved, and the agricultural pest identification technology based on artificial intelligence algorithms has attracted much attention. Many studies focus on using various algorithms to improve the pest identification accuracy. For example, combining the YOLOv8 backbone network and the Bi FPN network, developing improved vision Transformer methods, extracting local binary pattern features combined with support vector machines, and using convolutional neural network models, etc., have achieved certain results on different pest data sets. However, simply identifying the types of pests can no longer meet the needs of agricultural pest control. It is also necessary to master more pest information such as the source distribution, habitat, and prevention methods of pests to achieve precise prevention and control.

[0004] With the progress of modern agricultural sensing technology, various modal information has emerged. Cross-modal retrieval methods can construct the interaction and correlation between different modalities and obtain richer cross-modal insect information. Although there are many methods in the field of image-text retrieval, the image-text matching methods specifically for agricultural pests are extremely rare. There is an urgent need to deeply explore cross-modal image-text retrieval technology to meet the actual needs of the agricultural field. Summary of the Invention

[0005] In order to overcome the above-mentioned defects of the prior art, the embodiments of the present invention provide a cross-modal wolfberry pest retrieval model based on image and text hash feature learning to solve the problems raised in the above background art.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] A cross-modal wolfberry pest retrieval model based on image and text hash feature learning, characterized in that it includes an image-text encoder module, a multi-modal fusion module, a label enhancement module and an optimization module;

[0008] The image-text encoder module is used to process the image and text data of wolfberry pests in parallel;

[0009] The multi-modal fusion module is used to deeply fuse the image and text feature information;

[0010] The label enhancement module is used to retain the label frequency information and obtain high-quality supervision information;

[0011] The optimization module is used to optimize the model.

[0012] Preferably, the operation of the image-text encoder module specifically includes the following content:

[0013] Collect the image set and text set, and augment the original samples of the image data through random horizontal-vertical flipping, random brightness adjustment, random cropping and random shadow transformation;

[0014] The augmented text data is obtained through random insertion, synonym replacement, random deletion and random exchange.

[0015] Preferably, the operation of the multi-modal fusion module and the label enhancement module specifically includes the following content:

[0016] The paired images and texts are respectively represented by and respectively. The original image data and text data are respectively input into an image encoder and a text encoder with a Transformer structure, and the output image feature matrix is represented as and the text feature matrix is

[0017] The original label matrix is represented as L = [l1,…,l n ∈ R r×n , where r is the number of categories, and a smoothing function is adopted, defined as where w is the median of g, and the frequency of each category is calculated by The weighted label matrix is transformed into F = L ⊙ (f1 n ), and the normalized feature matrix of each modal sample is represented by so that the similarity of the same sample is equal to 1;

[0018] Rewrite the weighted label matrix F as Retain the modal similarity information and label frequency, and define the enhanced label G as:

[0019]

[0020] In the above formula γ m is the balance coefficient. Since the enhanced label G needs to add weight information, a constraint G = M⊙L is added to G and transformed into:

[0021]

[0022] After solving, the enhanced label matrix G is obtained, which promotes the generation of high-quality hash codes. The weighted fusion mode is adopted to obtain the unified feature matrix V0;

[0023]

[0024] where represents the mapping matrix that converts the m-th modality into the common space, and α m is the fusion coefficient of the m-th modality. Using the learned mapping matrix W m , the feature representation matrix V in the common subspace is obtained m = W m X m ;

[0025] There are differences between V1, V2 and the unified feature V0. Under the supervision of the enhanced label information, the representation ability of the final hash code is improved by minimizing the following equation;

[0026]

[0027] s.t. R T R = I;

[0028] where λ and φ are balance parameters, and β m is the weighted coefficient for the second-round fusion, P ∈ R c×r and P ∈ R c×c represent projection matrices;

[0029] Set joint feature learning, individual feature learning, and binary code learning. V0 is obtained through the joint feature learning part, and V1 and V2 are calculated according to the learned projection matrices W1 and W2 respectively;

[0030] At the same time, V0, V1, and V2 are used to optimize B, and the solved B is passed to the joint feature learning. The optimal B is obtained by iterating the above process. The objective function is summarized as follows:

[0031] Joint feature learning:

[0032]

[0033] Individual feature learning:

[0034] W1X1 → V1, W2X2 → V2;

[0035] Binary code learning:

[0036]

[0037] s.t. R T R = I.

[0038] Preferably, the operation of the optimization module specifically includes the following:

[0039] Fix other variables and rewrite the objective function of P m to obtain

[0040]

[0041] transformed into the maximization expression form with respect to ;

[0042] Let to obtain the solution of P m as follows:

[0043]

[0044] Fix other variables and update the label weight matrix M, and write the objective function of M as:

[0045]

[0046] The problem of G is reduced to L is used to constrain the structure of M, and M is solved as:

[0047]

[0048] To ensure that the element values of M are positive, when M ij < 0, set it to M ij = -M ij .

[0049] Preferably, the operation of the optimization module specifically further includes the following:

[0050] Fix other variables and rewrite the objective function of W m to obtain

[0051]

[0052] The derivative with respect to W m is zero, obtaining:

[0053]

[0054] Fix other variables and rewrite the objective function of V0 to obtain

[0055]

[0056] The derivative with respect to V0 is 0, obtaining

[0057]

[0058] Preferably in the present invention, the operation of the optimization module specifically further includes the following content:

[0059] Fix other variables and rewrite the objective function of α m to obtain

[0060]

[0061] wherein t is a smoothing parameter, obtaining the solution of α m ;

[0062]

[0063] Fix other variables and rewrite the objective function of R to obtain

[0064]

[0065] Simplify it to an expression in the following maximization form:

[0066]

[0067] The solution of R is R = UV;

[0068] wherein Preferably in the present invention, the operation of the optimization module specifically further includes the following content:

[0069] Fix other variables and rewrite the objective function of P to obtain

[0070]

[0071] The derivative with respect to P is zero, obtaining

[0072] P = (BB T + φI) -1 BL T ;

[0073] Fix other variables and rewrite the objective function of B to obtain

[0074]

[0075] wherein The bit-by-bit optimization strategy is adopted to obtain the closed-form solution of the $i$-th row of $\mathbf{B}$, that is

[0076] where $m$ i and $P$ i are the $i$-th rows of $\mathbf{M}$ and $\mathbf{P}$ respectively, is $\mathbf{P}$ excluding $P$ i , is $\mathbf{B}$ excluding $b$ i .

[0077] Preferably, the operation of the optimization module further includes the following:

[0078] Obtain the hash function that maps the original modal data to the Hamming space, and use the ridge regression method to learn the mapping matrix which is defined as follows:

[0079]

[0080] where $\theta$ is the penalty parameter, and its derivative with respect to $\mathbf{H}$ m is zero, so we can get

[0081]

[0082] The hash code of the $m$-th modality of the new query data is generated by the hash function:

[0083]

[0084] where $\text{sgn}$ represents the sign function, and $\mathbf{h}_m$ represents the output feature of the $m$-th modality after being processed by the pre-trained deep network.

[0085] The beneficial effects of the present invention are as follows:

[0086] 1. The CWPR model of the present invention constructs a multi-modal hash feature network model by deeply fusing image and text information, providing unprecedented rich text description content for the prevention and control of wolfberry pests, changing the limitation of simply relying on image recognition to identify pest species in the past, and enabling growers and technicians to comprehensively understand the pest situation.

[0087] 2. The innovative two-level fusion scheme fully considers the uniqueness between modalities, effectively avoids the loss of modal information, maximally integrates the advantages of images and texts, improves the model's accurate extraction ability of wolfberry pest characteristics, and thus improves the retrieval accuracy.

[0088] 3. Aiming at the thorny problem of uneven distribution of wolfberry pest and disease data, a label enhancement method is proposed. The enhanced labels provide high-quality supervision information for hash feature learning, ensuring that the model can still learn stably and efficiently and accurately identify the types and related information of pests when facing complex and diverse wolfberry pest data.

[0089] 4. It is proposed that the iterative alternating optimization algorithm can effectively solve the variables in the joint optimization problem. The evaluation results and visualization results in the actual wolfberry pest and disease data fully show that the proposed CWPR model can quickly and accurately provide pest and disease information content, inject strong impetus into the green and healthy development of the wolfberry industry, and help the process of agricultural modernization. Brief Description of the Drawings

[0090] Figure 1 is the species distribution map of the wolfberry pest dataset;

[0091] Figure 2 (a) is the PR curve of the CWPR model and other comparison methods in the I2I retrieval task;

[0092] Figure 2 (b) is the PR curve of the CWPR model and other comparison methods in the T2T retrieval task;

[0093] Figure 3 is the convergence curve of the CWPR model during the training stage on the wolfberry pest dataset;

[0094] Figure 4 (a) is the change curve of the mAP result with the change of μ;

[0095] Figure 4 (b) is the performance change curve of the CWPR model under different φ values;

[0096] Figure 4 (c) is the performance change curve of the CWPR model under different λ values;

[0097] Figure 4 (d) and Figure 4 (e) are the accuracy change curves of the CWPR model under different β0 values. Detailed Embodiment

[0098] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0099] The present invention proposes a cross-modal wolfberry pest retrieval model based on image and text hash feature learning, which is characterized by including an image-text encoder module, a multi-modal fusion module, a label enhancement module, and an optimization module;

[0100] The image-text encoder module is used to process the image and text data of wolfberry pests in parallel;

[0101] The multi-modal fusion module is used to deeply fuse the image and text feature information;

[0102] The label enhancement module is used to retain the label frequency information and obtain high-quality supervision information;

[0103] The optimization module is used to optimize the model.

[0104] The operation of the image-text encoder module specifically includes the following:

[0105] Collect the image set and text set, and augment the original samples of the image data through random horizontal-vertical flipping, random brightness adjustment, random cropping, and random shadow transformation;

[0106] The augmented text data is obtained through random insertion, synonym replacement, random deletion, and random swapping.

[0107] The operation of the multi-modal fusion module and the label enhancement module specifically includes the following:

[0108] The paired images and texts are respectively represented by and The original image data and text data are respectively input into the image encoder and text encoder with a Transformer structure, and the output image feature matrix is represented as and the text feature matrix is

[0109] The original label matrix is represented as L = [l1,…,l n ∈ R r×n , where r is the number of categories, and a smoothing function is adopted, defined as where w is the median of g, and the frequency of each category is calculated by The weighted label matrix is transformed into F = L ⊙ (f1 n ), and the normalized feature matrix of each modal sample is represented by such that the similarity of the same sample is equal to 1;

[0110] The weighted label matrix F is rewritten as Retaining the modal similarity information and label frequency, the enhanced label G is defined as:

[0111]

[0112] In the above formula γ m is the balance coefficient. Since adding weight information is required for the enhanced label G, a constraint G = M ⊙ L is added to G and transformed into:

[0113]

[0114] After solving, the enhanced label matrix G is obtained, which promotes the generation of high-quality hash codes. The unified feature matrix V0 is obtained using the weighted fusion mode;

[0115]

[0116] where represents the mapping matrix for converting the m-th modality into the common space, and α m is the fusion coefficient of the m-th modality. Using the learned mapping matrix W m , the feature representation matrix V in the common subspace is obtained m = W m X m ;

[0117] There are differences between V1, V2, and the unified feature V0. Under the supervision of the enhanced label information, the representational ability of the final hash code is improved by minimizing the following equation;

[0118]

[0119] s.t. R T R = I;

[0120] where λ and φ are balance parameters, and β m is the weighted coefficient for the second-round fusion. P ∈ R c×r and P ∈ R c×c represent projection matrices;

[0121] Set joint feature learning, individual feature learning, and binary code learning. V0 is obtained through the joint feature learning part, and V1 and V2 are calculated according to the learned projection matrices W1 and W2 respectively;

[0122] At the same time, V0, V1, and V2 are used to optimize B, and the solved B is passed to the joint feature learning. The optimal B is obtained by iterating the above process. The objective function is summarized as follows:

[0123] Joint feature learning:

[0124]

[0125] Individual feature learning:

[0126] W1X1 → V1, W2X2 → V2;

[0127] Binary code learning:

[0128]

[0129] s.t. R T R = I.

[0130] The optimization module operation specifically includes the following:

[0131] Fix other variables and rewrite the objective function of P m to obtain

[0132]

[0133] transformed into the form of a maximization expression with respect to ;

[0134] Let to obtain the solution of P m , as follows:

[0135]

[0136] Fix other variables and update the label weighting matrix M, and write the objective function of M as:

[0137]

[0138] The problem of G is reduced to L is used to constrain the structure of M, and M is solved as:

[0139]

[0140] To ensure that the element values of M are positive, when M ij < 0, set it to M ij = -M ij .

[0141] The optimization module operation specifically also includes the following:

[0142] Fix other variables and rewrite the objective function of W m to obtain

[0143]

[0144] The derivative with respect to W m is zero, obtaining:

[0145]

[0146] Fix other variables and rewrite the objective function of V0 to obtain

[0147]

[0148] The derivative with respect to V0 is 0, obtaining

[0149]

[0150] Optimizing the module operation specifically further includes the following:

[0151] Fix other variables and rewrite the objective function of α m to obtain

[0152]

[0153] where t is the smoothing parameter, obtaining the solution of α m ;

[0154]

[0155] Fix other variables and rewrite the objective function of R, obtaining

[0156]

[0157] Simplify it to an expression in the following maximization form:

[0158]

[0159] The solution of R is R = UV;

[0160] where

[0161] Optimizing the module operation specifically further includes the following:

[0162] Fix other variables and rewrite the objective function of P, obtaining

[0163]

[0164] The derivative with respect to P is zero, obtaining

[0165] P = (BB T + φI) -1 BL T ;

[0166] Fix other variables and rewrite the objective function of B, obtaining

[0167]

[0168] where Adopt a bit-by-bit optimization strategy to obtain the closed-form solution of the i-th row of B, that is

[0169] where m i and P i are the i-th rows of M and P respectively, is P excluding P i of P, is B excluding b i of B.

[0170] The specific operation of the optimization module further includes the following:

[0171] Obtain a hash function that maps the original modal data to the Hamming space, and use the ridge regression method to learn the mapping matrix is defined as follows:

[0172]

[0173] where θ is the penalty parameter, and the derivative with respect to H m is zero, and we can get

[0174]

[0175] The hash code of the m-th modality of the new query data is generated by the hash function:

[0176]

[0177] where sgn represents the sign function, represents the output feature of the m-th modality after being processed by the pre-trained deep network.

[0178] Select 10,598 samples of image data with text descriptions, including 8,363 samples of augmented data and 2,235 samples of original data. As Figure 1 shown, the text-image data is divided into 17 categories, and the number of samples in each category is different. Randomly select 5,000 sample subsets as the training set, and the remaining samples are used for testing.

[0179] The cross-modal wolfberry pest retrieval model based on image and text hash feature learning is labeled as CWPR. The ViT-Base model is introduced as the image encoder to process image data of size 224×224×3 to obtain deep image features. The text encoder adopts the structure of the bert-base-uncase model to extract text description features. The deep image features and text features are combined to solve the hash learning problem. In the training stage, the maximum number of iterations is set to 60 according to experience. Each hyperparameter is adjusted within a large range, and the optimal parameter values are found for the CWPR model. All experiments are conducted on a workstation equipped with Python 3.9.19 and MATLAB R2018b, running with 16GB of memory and an Intel(R) Core(TM) i7-10700 CPU @ 2.90 GHz. The commonly used mean average precision (mAP) is used to comprehensively evaluate the performance. The average precision (AP) for a given query q is defined as follows:

[0180]

[0181] where l q records the number of correct instances among the first R retrieved instances; P q (m) represents the accuracy rate of the first m instances; if the m-th position in the returned sequence is relevant to the query, δ q (m) = 1, otherwise, δ q (m) = 0. The average of the APs for all queries is the mAP. In the following experiments, the higher the mAP value, the better the performance of the model.

[0182] Table 1 Comparison of mAP for cross-modal retrieval tasks under different bit lengths

[0183]

[0184]

[0185] A comparative experiment was conducted between CWPR and hash models such as DCMH, DJRSH, DCHUC, AGCH, M IAN, I GNNCON, etc. The retrieval performance was verified using the wolfberry pest image-text dataset. All methods were trained on the training set and verified on the test set. Each method was run 10 times for experiments, and the average was taken to obtain the retrieval accuracy. Two typical retrieval cases were introduced, namely image query image and text query text, which are usually abbreviated as "I2I" and "T2T" respectively. The hash code lengths were set to 16 bits, 32 bits, 64 bits, and 128 bits respectively. Table 1 reports the retrieval performance of all methods under different code lengths, and these methods are arranged in chronological order of publication. In the I2T retrieval task and the T2I retrieval task, the average accuracy rates of CWPR under all hash length settings are 96.07% and 97.93% respectively. Compared with the mainstream deep learning-based cross-modal retrieval methods, the CWPR model has achieved the best performance in cross-modal wolfberry pest retrieval. In addition, precision and recall were adopted as additional evaluation metrics to evaluate the performance of CWPR. By changing the Hamming radius value to obtain the points falling within different radius ranges, a precision-recall (PR) curve was plotted, as shown in Figure 2 shown. Due to the complex and variable background environment of wolfberry pest images and the small semantic information differences between the text descriptions of different species of pests, cross-modal wolfberry pest retrieval is a challenging problem. In the wolfberry retrieval task, the performance of CWPR is the best among all methods. The reasons are as follows: on the one hand, since the ViT model and the Bert model are large language models pre-trained on large-scale natural data, they perform excellently in feature extraction. The CWPR model uses the state-of-the-art ViT model and Bert model to extract the shallow and deep features of images and texts respectively. On the other hand, the CWPR model effectively fuses the image modality features, text modality features, and their joint features, which can alleviate the heterogeneous differences between modalities and at the same time enable the fused features to retain the image and text modality information. In view of the problem of the imbalance in the number of species in the wolfberry insect samples, a label enhancement method was introduced in the CWPR model to retain the distribution information of the species, so as to provide high-quality supervision information for hash feature learning. The CWPR model comprehensively utilizes the above advantages to improve the performance in the two cross-modal wolfberry insect retrieval tasks. As the hash code length increases, the performance of the CWPR model improves slightly. Among all the preset possible lengths, the CWPR model reaches the best retrieval accuracy at 64-bit hash code.

[0186] Ablation experiment

[0187] The CWPR framework consists of three important parts: joint feature learning, independent modality feature learning, and label enhancement module. To study the impact of these three parts on the cross-modal retrieval performance of wolfberry pests, the following ablation experiments were conducted. The specific experimental settings are as follows: CWPR-I means ignoring independent modality feature learning in the CWPR framework. CWPR-II means omitting the joint feature learning part in the CWPR framework. CWPR-III means that the CWPR model directly uses the original labels to supervise hash feature learning without the label enhancement step. As shown in Table 2, CWPR is superior to CWPR-I, CWPR-II, and CWPR-III in terms of cross-modal retrieval accuracy. In the I2T task, the accuracy of CWPR is 1.15% higher than that of the best-performing incomplete model (CWPR-III). In addition, in the T2I retrieval task, the accuracy of CWPR is improved by at least 0.44% compared to other cases. The experimental results show that the multi-modal fusion learning module and the introduced label enhancement technology can help CWPR learn the fine hash features of wolfberry pest images and texts.

[0188] Table 2 Ablation experiment results of the CWPR model

[0189]

[0190] Convergence analysis

[0191] The CWPR algorithm obtains the optimal variable values by iteratively optimizing all variables. In the experiment, the convergence of the iterative optimization algorithm was observed by recording the values of the objective loss function after each iteration. As Figure 3 shown, the horizontal axis represents the number of iterations, and the vertical axis represents the objective value. It can be seen that the curve drops smoothly and approaches a stable state within 10 iterations.

[0192] Parameter sensitivity analysis

[0193] Experimental analysis was conducted on several hyperparameters in the CWPR model. In the experiment, by fixing other parameters and adjusting the parameters within the candidate range, we observed the impact of these changes on the CWPR retrieval performance. For ease of observation, the hash code length was uniformly set to 64 bits. The candidate ranges of the parameters μ, φ, λ, and θ were set to {1e -5 , 1e -3 , 1e -1 , 1e 3 , 1e 5} according to experience, and the value range of β0 is {0.1, 0.3, 0.5, 0.7, 0.9}. The parameter μ controls the quantitative loss term Figure 4 (a) shows the change process of the mAP results as μ changes. When μ > 1e -1When the mAP result first rises sharply and then drops rapidly. Therefore, within the candidate range, when μ is set to 1e -1 the CWPR model performs best. When the value of μ is too large, the performance deteriorates rapidly. The reason for this phenomenon is that if the value of μ is too large, this term will dominate and overly affect other terms in the objective function. The parameter φ controls the constraint term to avoid overfitting, and its value is usually not too large. Figure 4 (b) shows the performance change curve of the CWPR model under different φ values. When φ > 0.1, the performance of CWPR decreases. The parameter λ affects the relative importance of the classification term. As Figure 4 (c) shows, as the value of λ increases, the performance of the CWPR model improves among all possible values. This experimental phenomenon indicates that pest category information promotes hash feature learning. The parameter θ is the balance coefficient of the balance term in hash function learning. Figure 4 (d) shows the accuracy change curve of the CWPR model under different β0 values. As the value of β0 increases, the retrieval accuracy of the model gradually decreases. The parameter β0 is a fusion weight coefficient that determines the relative importance between source data information. As Figure 4 (e) shows, when the value of β0 is 0.5, the CWPR model achieves the best retrieval accuracy. In summary, CWPR can find the appropriate parameter settings to achieve optimal performance in cross-modal wolfberry pest retrieval.

[0194] Compared with the single-layer fusion mechanism, the CWPR model improves the average retrieval accuracy by at least 1.2%. In addition, a label enhancement technique considering label distribution information is introduced to guide the learning of hash features. Experimental verification shows that the CWPR model can improve the accuracy of cross-modal wolfberry pest retrieval by 0.8%. CWPR uses ViT and BERT as the backbone networks of image and text encoders respectively. Compared with the best comparison method, the CWPR model achieves an average performance improvement of 1.89% in the cross-modal wolfberry pest retrieval task. Therefore, the CWPR model shows great potential in content retrieval for wolfberry pest control and has broad application prospects as an information service technology for pest control and management in agriculture.

[0195] Although image-based pest classification methods have made rapid progress in the development of agricultural intelligence, such as the emergence of many successful classification models, these methods can only obtain the species information of pests and still require the participation or assistance of human experts to complete scientific pest control. The present invention directly realizes the matching between images and text descriptions, providing a new research perspective for the identification and management of agricultural pests. Experiments show that the cross-modal pest retrieval method can accurately match pest images with detailed text descriptions, expanding the application of traditional classification methods. The proposed CWPR model can be applied to the following scenarios: managers submit captured pest images to the pest management system, and the system automatically retrieves matching pest description information from the pest management database containing scientific literature, pest control guidelines, and expert knowledge sources. Therefore, the CWPR model proposed in this paper is an extension of the pest classification method, which can not only obtain the species information of pests but also provide accurate and comprehensive pest characteristic information.

[0196] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents;

[0197] What is described in the above specification is only the specific implementation manners of the present invention. Various examples do not constitute limitations on the essence of the present invention. Those of ordinary skill in the technical field can make modifications or variations to the previously described specific implementation manners after reading the specification without departing from the essence and scope of the invention.

Claims

1. A cross-modal wolfberry pest retrieval model based on image and text hash feature learning, characterized in that: It includes image-text encoder module, multimodal fusion module, label enhancement module and optimization module; The image-text encoder module is used to process the image and text data of wolfberry pests in parallel; The multimodal fusion module is used to deeply fuse image and text feature information; The label enhancement module is used to retain label frequency information and obtain high-quality supervision information; The optimization module is used to optimize the variables in the model.

2. The cross-modal wolfberry pest retrieval model based on image and text hash feature learning according to claim 1 is characterized in that: The image text encoder module operation specifically includes the following: Collect image sets and text sets, and expand the original samples of image data through random horizontal-vertical flipping, random brightness adjustment, random cropping, and random shadow transformation; The augmented text data is obtained through random insertion, synonym replacement, random deletion and random exchange.

3. The cross-modal wolfberry pest retrieval model based on image and text hash feature learning according to claim 2 is characterized in that: The operation of the multimodal fusion module and the label enhancement module specifically includes the following: The paired images and texts are and Indicates that the original image data and text data are input into the image encoder and text encoder with Transformer structure respectively, and the output image feature matrix is ​​expressed as And the text feature matrix is The original label matrix is ​​represented as L = [l1,…,l n ]∈R r×n , where r is the number of categories, using a smooth function, defined as Where w is the median of g, and the frequency of each category is given by Calculation, weighted label matrix is ​​transformed into F = L⊙(f1 n ), the normalized feature matrix of each modal sample is Indicates that the similarity of the same samples is equal to 1; Rewrite the weighted label matrix F as Retaining modality similarity information and label frequency, the enhanced label G is defined as: In the above formula γ m To balance the coefficient, since the enhanced label G needs to add weight information, a constraint G = M⊙L is added to G, which is transformed into: After solving, the enhanced label matrix G is obtained, which promotes the generation of high-quality hash codes, and the weighted fusion mode is used to obtain a unified feature matrix V0; in represents the mapping matrix that transforms the mth mode into the common space, α m is the fusion coefficient of the mth mode, using the learned mapping matrix W m , get the feature representation matrix V in the common subspace m =W m X m ; There are differences between V1, V2 and the unified feature V0. Under the supervision of the enhanced label information, the representation ability of the final hash code is improved by minimizing the following equation; Where λ and φ are equilibrium parameters, β m is the weighted coefficient of the second round of fusion, P∈R c×r and P∈R c×c represents the projection matrix; Set up joint feature learning, individual feature learning and binary code learning, obtain V0 through the joint feature learning part, and calculate V1 and V2 according to the learned projection matrices W1 and W2 respectively; Use V0, V1 and V2 to optimize B at the same time, pass the solved B to joint feature learning, and iterate the above process to obtain the optimal B. The objective function is summarized as follows: Joint Feature Learning: Individual characteristics study: W1X1→V1,W2X2→V2; Binary code learning:

4. The cross-modal wolfberry pest retrieval model based on image and text hash feature learning according to claim 3 is characterized in that: The optimization module operation specifically includes the following: Fix other variables and rewrite P m The objective function is Convert to about The maximization expression form of ; make Get P m The solution is as follows: Fix other variables and update the label weight matrix M. The objective function of M is written as: The problem with G boils down to L is the structure used to constrain M, and M is solved as: To ensure that the element values ​​of M are positive, when M ij <0, set to M ij =-M ij .

5. The cross-modal wolfberry pest retrieval model based on image and text hash feature learning according to claim 4 is characterized in that: The optimization module operation also includes the following: Fix other variables and rewrite W m The objective function is About W m The derivative of is zero, so we get: Fix other variables and rewrite the objective function of V0 to obtain The derivative about V0 is 0, so we get 6. The cross-modal wolfberry pest retrieval model based on image and text hash feature learning according to claim 5 is characterized in that: The optimization module operation also includes the following: Fix other variables and rewrite α m The objective function is in t is the smoothing parameter, and we get α m The solution; Fixing other variables and rewriting the objective function of R, we obtain This simplifies to the following expression in maximized form: The solution of R is R = UV; in 7. The cross-modal wolfberry pest retrieval model based on image and text hash feature learning according to claim 6 is characterized in that: The optimization module operation also includes the following contents: Fixing other variables and rewriting the objective function of P, we obtain The derivative with respect to P is zero, so we have P=(BB T +φI) -1 BL T ; Fixing other variables and rewriting the objective function of B, we get in A bitwise optimization strategy is used to obtain the closed-form solution for the i-th row of B, Right now Where m i and P i are the i-th row of M and P respectively, Does not include P i P, Does not include b i B.

8. The cross-modal wolfberry pest retrieval model based on image and text hash feature learning according to claim 7, characterized in that: The optimization module operation also includes the following contents: Obtain a hash function that maps the original modal data to the Hamming space and use the ridge regression method to learn the mapping matrix The definition is as follows: Where θ is the penalty parameter, m The derivative of is zero, so we have New query data The hash code of the mth modulus is generated by the hash function: where sgn represents the sign function, Represents the output features of the mth modality after processing by the pre-trained deep network.