Visual Transform-based silk image long tail distribution processing method

By adopting a network model based on visual Transformer in silk image processing, combining residual connection and self-distillation hashing scheme, the problem of poor generalization ability under long tail distribution is solved, and more efficient and accurate image retrieval is achieved.

CN120196779APending Publication Date: 2025-06-24DONGHUA UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510268964.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art has overfitting when processing the long-tail distribution of silk images, which seriously restricts the generalization ability of the model and lacks effective hash image retrieval methods.

Method used

A network model based on visual Transformer is adopted, and hash encoding is generated and searched using a coarse to fine-tune hierarchical search strategy through pre-training and fine-tuning, combining residual connection and self-distillation hash scheme.

Benefits of technology

It improves the retrieval performance of the model in long-tail data tasks, enhances the information extraction ability, reduces the data imbalance caused by quantization error and long-tail distribution, and improves the accuracy of image retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196779A_ABST
    Figure CN120196779A_ABST
Patent Text Reader

Abstract

The invention relates to a silk image long tail distribution processing method based on a visual Transform, and the method comprises the following steps: obtaining a visual Transform network, carrying out the pre-training of the visual Transform network on a pre-established image data set, and freezing the network parameters of the visual Transform network after the pre-training is carried out to an optimal effect; network parameters of the trained visual Transform network are loaded to a corresponding network layer of a pre-constructed network model, the visual Transform network serves as a main network of the network model, other network parameters of the network model are initialized randomly, and fine adjustment is conducted on the network model in a silk image data set obtained in advance; and obtaining a silk image to be retrieved, and performing retrieval in the image library through the fine-tuned network model to obtain an image retrieval result. Compared with the prior art, the method has the advantages that the information extraction capability is enhanced, and the retrieval accuracy is improved; and quantization errors are reduced, and data imbalance caused by long-tail distribution is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image retrieval, and in particular to a method for processing long-tail distribution of silk images based on visual Transformer. Background Art

[0002] Long-tail distribution is a data distribution in which a small number of categories (generally called head categories) occupy a large number of samples in the data set, while a large number of categories (generally called tail categories) occupy a small number of samples, showing great imbalance. The deep learning model applied in the long-tail distribution is called the long-tail learning model.

[0003] Data collected from the real world often have a long-tail distribution. For example, in the classification and recognition of wildlife images, the number of images containing rare animals is often far less than the number of images containing common animals such as tigers and elephants. Due to its practical significance, long-tail learning has received more and more attention in recent years.

[0004] Traditional long-tail learning methods are often implemented based on resampling or category-sensitive learning methods. The former modifies the data set to generate balanced data through undersampling, oversampling, etc.; the latter modifies the loss function by modifying weights, margins, etc. to change the weights of different categories in the loss function. However, traditional methods often show overfitting on the tail classes, which seriously restricts the generalization ability of the model.

[0005] In summary, there is currently a lack of a hash image retrieval method to solve or partially solve the aforementioned long-tail learning problem of silk images. Summary of the invention

[0006] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and provide a silk image long-tail distribution processing method based on visual Transformer to improve the retrieval performance of the model in long-tail data tasks.

[0007] The purpose of the present invention can be achieved by the following technical solutions:

[0008] A method for processing long-tail distribution of silk images based on visual Transformer, comprising the following steps:

[0009] Acquire and pre-train a visual Transformer network on a pre-established image dataset, and freeze network parameters of the visual Transformer network after pre-training to an optimal effect;

[0010] Load the network parameters of the trained Vision Transformer network into the corresponding network layers of the pre-constructed network model. The network model uses the Vision Transformer network as the backbone network, and the other network parameters of the network model are randomly initialized. Then, fine-tune the network model on the pre-acquired silk image dataset.

[0011] Obtain the silk image to be retrieved, and perform retrieval in the image library through the fine-tuned network model to obtain the image retrieval result.

[0012] Further, the fine-tuning process of the network model includes the following steps:

[0013] Obtain the input image from the silk image dataset, perform a convolution operation on the features of the input image, and output shallow features.

[0014] Perform a chunking operation on the shallow features to obtain multiple feature chunks, and convert the feature information of each feature chunk into a one-dimensional sequence respectively.

[0015] Through linear embedding operation, add position encoding information to each one-dimensional sequence, then perform feature extraction through multiple multi-head attention modules in sequence, and then splice the feature extraction results of each feature chunk to obtain the deep features of the input image.

[0016] Through the residual connection mechanism, fuse the shallow features and the deep features to obtain the fused features.

[0017] Generate a hash code for the obtained fused features through a hash function.

[0018] Based on the generated hash code and the fused features, adopt a coarse-to-fine hierarchical search strategy to perform retrieval in the silk image dataset, and calculate the loss function of the network model according to the retrieval results, so as to iteratively update the network parameters.

[0019] Further, in the process of generating the hash code, the function tanh(x) is used as the activation function.

[0020] Further, the calculation process of the loss function of the network model includes the following steps:

[0021] Calculate the cosine similarity of the hash codes corresponding to the same input image after different augmentations through the self-distillation scheme, so as to obtain the value of the self-distillation loss function.

[0022] Based on the input of the hash code into the classification layer of the network model, obtain the corresponding semantic label, and thus calculate the value of the hash proxy loss function.

[0023] Calculate the quantization loss function value by estimating the binary likelihood value of the hash code and comparing it with the corresponding binary likelihood label;

[0024] Construct the loss function of the network model based on the self-distillation loss function value, the hash proxy loss function value, and the quantization loss function value.

[0025] Further, the calculation expression of the self-distillation loss function value is:

[0026] L SdH (h T , h S ) = 1 - S(h T , h S ) In the formula, L SdH is the self-distillation loss function value, h T , h S represent the hash codes corresponding to the same sample after different augmentations, and S(h T , h S ) is the cosine similarity between the two hash codes.

[0027] Further, the calculation process of the hash proxy loss function value includes:

[0028] Use a set of trainable hash proxies P θ to implement representation learning based on the hash proxy P θ , and use the hash proxy P θ and the hash code h T of the input image to calculate the class-level classification prediction. The calculation expression of the class-level classification prediction is:

[0029] p T = [(p θ1 , h T ), S(p θ2 , h T ), …, S(pθN cls , h T )]

[0030] In the formula, p T is the class-level classification prediction result, p θi is the hash proxy value assigned to the i-th class, N cls is the number of classes to be distinguished, and S is the calculation of the cosine similarity;

[0031] Based on the class-level classification prediction result p T , learn the similarity of the class label corresponding to the input image by calculating the hash proxy loss, and obtain the hash proxy loss function value. The calculation expression of the hash proxy loss function value is:

[0032] L HP(y, p T , τ) = H(y, softmax(p T / τ))

[0033] where τ is the temperature scale hyperparameter, H(u, v) = -∑ k u k log v k is the cross entropy, and the softmax operation is applied along the dimension of p T , and y is the class label.

[0034] Furthermore, the calculation process of the quantization loss function value includes:

[0035] Using a Gaussian distribution estimator g(h) with a predefined mean of m and a standard deviation of σ to estimate the binary likelihood of the hash code h, and the corresponding calculation expression is:

[0036]

[0037] Based on the calculated binary likelihood value, calculate the quantization loss function value, and the corresponding calculation expression is:

[0038]

[0039] where L bce-Q (h T ) is the quantization loss function value, H b = -u log v + (1 - u) log(1 - v) is the binary cross entropy, is the estimated likelihood value of the k-th hash code element, is the binary likelihood label,

[0040] Furthermore, during the update process of the network parameters, the network parameters corresponding to the Vision Transformer network remain frozen.

[0041] Furthermore, the specific process of the coarse-to-fine hierarchical search strategy includes:

[0042] Coarse range retrieval step: Calculate the Hamming distance between the hash code corresponding to the input image and the hash codes corresponding to all images in the silk image dataset, and sort the images from small to large according to the calculation results, and add the first m images in the sorting results to the alternative image pool;

[0043] Precise range retrieval steps: Calculate the binary classification results between the feature vector of the input image and the feature vectors corresponding to all images in the alternative image pool, use the binary classification results as a measure of the similarity between images, sort all images in the alternative image pool according to the measure of similarity, and select the top k images in the sorting result as the final retrieval result.

[0044] Furthermore, the silk image dataset contains various types of silk images and covers different silk textures, patterns, colors, and texture features; the label distribution in the silk image dataset shows a long-tail effect, that is, the number of samples of specific patterns or rare silk types is small, and the number of samples of common categories is large;

[0045] The method also includes data augmentation on the obtained silk image dataset to increase the number and diversity of samples in the tail categories.

[0046] Compared with the prior art, the present invention has the following advantages:

[0047] (1) Enhance the ability of information extraction and improve retrieval accuracy: The present invention uses a network model based on ViT (Vision Transformer) as the backbone network and adds residual connections to solve the long-tail distribution problem. Different from traditional deep convolutional neural networks, the present invention uses Transformer as the backbone network and applies it to the field of silk image retrieval. This network can capture the shallow features and deep features of the input image at one time, reduce the loss of detailed information under the long-tail distribution, and at the same time introduce residual connections to fuse the shallow feature information and deep feature information to ensure the completeness of the finally extracted feature information, enhance the learning ability of the model for tail categories, and thus improve the performance of the model in long-tail data and the accuracy of image retrieval.

[0048] (2) Reduce quantization error: The hash layer of the present invention uses the tanh function as the quantization function, effectively improving the problem of gradient disappearance in backpropagation and improving the stability and convergence efficiency of model training. At the same time, a self-distillation hash scheme is adopted. By using the high-level semantic information of the model itself as the supervision signal, the representation ability of the hash code is further optimized, and the feature extraction and classification ability of the model for complex data is enhanced, thereby improving the overall performance.

[0049] (3) The network model of the present invention adopts a self-distillation scheme to generate hash codes consistent with the sample labels on the premise of maintaining pairwise similarity. Through the teacher-student architecture, the teacher model provides stable guidance information to help the student model generate more accurate hash codes, further improving the accuracy of hash coding.

[0050] (4)Reduce data imbalance caused by long-tailed distribution: The present invention uses data augmentation operations, such as rotation, flipping, cropping, etc., to expand the tail dataset. Description of the Drawings

[0051] Figure 1 It is a schematic flowchart of a method for processing long-tailed distribution of silk images based on Vision Transformer provided in an embodiment of the present invention;

[0052] Figure 2 It is a schematic flowchart during training on a silk image dataset provided in an embodiment of the present invention. Detailed Embodiments

[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Generally, the components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.

[0054] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0055] It should be noted that: Similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0056] Embodiment 1

[0057] The present invention provides a method for processing long-tailed distribution of silk images based on Vision Transformer, which uses a network model based on the ViT (Vision Transformer) network that has been pre-trained and fine-tuned, and performs retrieval based on the input dataset with long-tailed distribution, such as Figure 1 shown, including the following steps:

[0058] S1. Pre-train the ViT network on the ImageNet dataset. After training to the optimal effect, freeze the current parameters to improve the generalization ability of the model.

[0059] S2. Load the parameters of the ViT network trained in S1 into the corresponding network layers of the network model of this embodiment, and randomly initialize other parameters. Then fine-tune the network model in the silk image dataset.

[0060] Specifically, it includes the following steps:

[0061] Obtain and pre-train the Vision Transformer network on a pre-established image dataset. After pre-training to the optimal effect, freeze the network parameters of the Vision Transformer network;

[0062] Load the network parameters of the trained Vision Transformer network into the corresponding network layers of a pre-constructed network model. The network model uses the Vision Transformer network as the backbone network, and the other network parameters of the network model are randomly initialized. Fine-tune the network model in the pre-obtained silk image dataset;

[0063] Obtain the silk image to be retrieved, and retrieve it in the image library through the fine-tuned network model to obtain the image retrieval result.

[0064] Preferably, in step S2, the fine-tuning process of the network model includes the following steps:

[0065] Obtain an input image from the silk image dataset, perform a convolution operation on the features of the input image, and output shallow features;

[0066] Perform a chunking operation on the shallow features to obtain multiple feature chunks, and convert the feature information of each feature chunk into a one-dimensional sequence respectively;

[0067] Through a linear embedding operation, add position encoding information to each one-dimensional sequence, then perform feature extraction through multiple multi-head attention modules in sequence, and then splice the feature extraction results of each feature chunk to obtain the deep features of the input image;

[0068] Through a residual connection mechanism, fuse the shallow features and the deep features to obtain fused features;

[0069] Generate a hash code for the obtained fused features through a hash function;

[0070] Based on the generated hash code and the fused features, adopt a coarse-to-fine hierarchical search strategy to retrieve in the silk image dataset, and calculate the loss function of the network model according to the retrieval result, so as to iteratively update the network parameters.

[0071] Specifically, for the input image, use a linear projection to replace the traditional convolutional layer, divide the image into fixed-size image patches, and embed them into a high-dimensional feature space to extract initial shallow features;

[0072] Based on the embedded shallow features, after adding positional encoding to the image patches, they are input into the ViT model, and the multi-head self-attention module is used for feature extraction to obtain the deep features of the image;

[0073] Through the residual connection mechanism, the shallow features and the deep features are fused to generate fused features with richer context information;

[0074] The fused features are used to generate hash codes through a hash function to represent the compact features of the image;

[0075] Based on the generated hash codes and fused features, a hierarchical search strategy from coarse to fine is adopted to retrieve in the image library, and finally accurate retrieval results are obtained.

[0076] For the training of the network model, a self-distillation hashing scheme is proposed. The weight-sharing Siamese structure is used to simultaneously compare the hash codes of different views (enhanced results) of an image. Two independent enhancement groups are configured to generate weak transformation views and strong transformation views respectively to construct a training framework of a simple teacher and a difficult student. The difficulty is controlled in a random sampling manner: the same hyperparameters are used for all transformations in the group, and they occur less or more by scaling their own occurrence probabilities.

[0077] Based on the two different hash codes of the samples, calculate the value of the self-distillation loss function;

[0078] Based on the output obtained by inputting the hash code into the classification layer and the semantic label of the sample, calculate the value of the hash proxy loss function;

[0079] Based on the quantization loss of the hash code and binary cross-entropy, calculate the value of the quantization loss function;

[0080] Based on the values of the self-distillation loss function, the classification loss function, and the quantization loss function, fine-tune the parameters of the network model.

[0081] Specifically, the hierarchical search strategy from coarse to fine includes:

[0082] Calculate the Hamming distance between the hash code corresponding to the image feature and the hash codes corresponding to each image in the image library and sort them, and select the top m images and add them to the alternative image pool;

[0083] Calculate the fused feature corresponding to the image feature, and perform binary classification with the fused features of each image in the alternative image pool. Then, sort according to the classification results and select the top k images as the final retrieval results.

[0084] Specifically, in this embodiment, step S2 includes the following steps:

[0085] S21. Augment and preprocess the silk image dataset. During the collection of silk image data, it may be affected by various factors such as changes in lighting conditions, texture complexity, and pattern occlusion. Therefore, to enrich and expand the silk dataset and avoid overfitting of the model, various data augmentation operations are performed on the existing dataset, including random rotation and flipping, random scaling, color transformation, adding noise, and random cropping. These operations can effectively improve the generalization ability of the model for silk images and its adaptability to complex scenarios.

[0086] S22. For the feature F of an input image with size H×W×C in , perform a convolution operation on it, and the output feature information is F = Conv(Fin).

[0087] S23. Perform a block operation on the convolved feature F. Evenly divide the input feature F into 9 blocks, and the size of each feature block Fi is Convert the feature information of each block into a one-dimensional sequence with a length of .

[0088] S24. Through a linear embedding operation, perform position embedding on each one-dimensional feature sequence, that is, according to the original spatial position of each one-dimensional feature sequence in the feature map, perform the corresponding position encoding information splicing operation. At the same time, add a position embedding block representing global information to obtain the feature F E .

[0089] S25. After passing through L multi-head attention modules, extract effective feature information. Use F MHA (g) to represent the feature extracted by a multi-head attention module. The feature obtained after passing through the i-th multi-head attention module is:

[0090] F i = F MHAi (F i―1 ) = F MHAi (F MHAi―1 (…F MHA1 (F E )…)), i = 1, 2, …, L

[0091] S26. Through the concatation operation, splice the feature information of each input image block. Introduce a residual connection to fuse the shallow feature F and the finally extracted deep feature F L to obtain a fused feature, enabling the network to extract both the fine-grained key feature information of the input image and the coarse-grained information of the shallow feature, which helps to improve the feature expression ability of the network.

[0092] S27. Use the finally extracted fusion features and the coarse-grained feature information output by the shallow layer of the network model as the input of the hash layer to learn and generate corresponding hash codes.

[0093] To effectively avoid the problem of gradient disappearance during the quantization of hash codes using the activation function sign(x) in the hash layer of previous deep hash networks, the network model of the present invention uses the function tanh(x) as the activation function of the hash layer. This activation function is continuous and smooth, and is an approximation function of the function sign(x), which can effectively update the parameters through backpropagation and learn hash codes containing high-level semantic information. The relationship between the function sign(x) and tanh(x) is as follows:

[0094]

[0095] S28. Introduce a self-distillation scheme and compare the hash codes of different views (enhanced results) of an image at the same time. Two independent enhancement groups are configured to generate weakly transformed views and strongly transformed views respectively to construct a training framework of a simple teacher and a difficult student.

[0096] S29. The similarity between binary codes in the Hamming space can be represented by cosine similarity:

[0097]

[0098] Among them, S in the formula is the cosine similarity of the binary code. The greater the similarity, the smaller the distance between the two. Combining the cosine similarity and the two different hash codes generated by the self-distillation scheme, the self-distillation loss function can be obtained. Define the self-distillation loss function as L SdH , and the calculation formula is:

[0099] L Sd (h T , h S ) = 1 - S(h T , h S )

[0100] Among them, L SdH is the value of the self-distillation loss function, h T , h S represent the hash codes corresponding to the same sample after different augmentations, and S(h T , h S ) is the cosine similarity between the two hash codes.

[0101] S30. In addition to self-distillation hashing, the present invention uses the learned teacher hash code h S to calculate the loss in order to transfer the learned hash knowledge to the student's code. By using a set of trainable hash proxies Pθ , proxy-based representation learning is introduced in deep hashing, and P will be used first θ and h T to calculate the class-level classification prediction, as shown in the following formula:

[0102] p T = [(p θ1 , H T ), S(p θ2 , H T ), …, S(pθN cls , H T )]

[0103] Then, p T is used to learn the similarity with the class label by calculating the hash proxy loss. Define L HP as the hash proxy loss function, and the formula is:

[0104] L HP (y, p T , τ) = H(y, softmax(p T / τ))

[0105] where pT = [S(pθ1, HT), S(pθ2, HT), …, S(pθ Ncls , HT)], pθi is the hash proxy assigned to the i-th class, N cls represents the number of classes to be distinguished, τ is the temperature scale hyperparameter, H(u, v) = -∑k uk log vk is the cross entropy, and the softmax operation is applied along the dimension of pT.

[0106] S31. To enable consecutive hash code elements to work like binary bits, the deep hashing method aims to reduce the quantization error by minimizing the distance (e.g., Euclidean distance) between the hash code bits and their nearest binary targets (+1 or -1) in a regression manner. However, since the goal of hashing is to classify the sign of each bit, a more natural choice is to treat it as binary classification and use a Gaussian distribution estimator g(h) with a predefined mean of m and a standard deviation of σ to estimate the binary likelihood of the hash code element h, as shown in the following formula:

[0107]

[0108] The value of the quantization loss function is calculated using the following formula:

[0109]

[0110] In the formula, L bce-Q (h T) To quantify the loss function value, Hb = -ulogv + (1 - u)log(1 - v) is the binary cross-entropy, is the estimated likelihood value of the k-th hash coding element, is the binary likelihood label,

[0111] Therefore, the loss function of the network model in this embodiment is:

[0112] L = L bce-Q +L HP +L SdH

[0113] S3. Perform image retrieval on the input image. Adopt a hierarchical search strategy from coarse to fine, and jointly utilize the features extracted by the network model and the generated hash coding to perform fast and accurate image retrieval. First, retrieve a certain number of images with similar distances in the hash coding space from the image library through the generated hash coding as candidate images. In order to further exclude images with a certain similarity in appearance to the input image but not of the same class label from the candidate images, then perform similarity comparison and sorting based on the feature maps extracted by the deep feature extractor. The specific implementation details are as follows.

[0114] First, perform a rough range retrieval. Calculate the Hamming distance between the hash coding corresponding to the query image and the hash coding corresponding to all images in the image library and sort them from small to large, and add the first m images in the sorting result to the candidate image pool.

[0115] Then, perform an accurate range retrieval. First, calculate the binary classification result between the feature vector of the query image and the feature vectors corresponding to all images in the candidate image pool, and use the classification result as a measure of the similarity between images. The classification result can be marked as "similar" or "dissimilar", and sort the images according to this classification result. Then, take the first k images in the sorting result as the final retrieval result. During this process, the query library includes the fusion features and hash coding of multiple pictures, which are used to generate the binary classification result and perform efficient retrieval.

[0116] This embodiment provides a silk dataset with long-tailed distribution characteristics. The dataset contains various types of silk images, covering different silk textures, patterns, colors, and textures. The label distribution in the dataset shows a long-tailed effect, with fewer samples in a few categories (such as specific patterns or rare silk types) and more samples in common categories. To address this imbalance problem, the dataset is augmented through various data augmentation methods, including random rotation, flipping, color adjustment, and cropping, etc., so as to increase the number of samples and diversity of the tail categories. This dataset is applicable to computer vision tasks such as classification and retrieval of silk images, especially for the problem of data imbalance, and can effectively evaluate the robustness and performance of the model.

[0117] The present invention provides a basic method, and specific parameters, including the number of neural network layers, the number of blocks, and the training method, etc., can be adjusted as needed.

[0118] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention through logical analysis, reasoning, or limited experiments based on the concept of the present invention on the basis of the prior art should be within the protection scope determined by the claims.

Claims

1. A method for processing long-tail distribution of silk images based on visual Transformer, characterized in that: The following steps are involved: Acquire and pre-train a visual Transformer network on a pre-established image dataset, and freeze network parameters of the visual Transformer network after pre-training to an optimal effect; Loading the trained network parameters of the visual Transformer network into the corresponding network layers of a pre-built network model, wherein the network model uses the visual Transformer network as the backbone network, and the other network parameters of the network model are randomly initialized, and the network model is fine-tuned in a pre-acquired silk image dataset; Get the silk image to be retrieved, search it in the image library through the fine-tuned network model, and get the image retrieval result.

2. According to the method for processing long-tail distribution of silk images based on visual Transformer in claim 1, it is characterized in that: The fine-tuning process of the network model includes the following steps: Acquire an input image from the silk image dataset, perform a convolution operation on features of the input image, and output shallow features; Performing a block operation on the shallow features to obtain a plurality of feature blocks, and converting feature information of each feature block into a one-dimensional sequence; Through linear embedding operation, position encoding information is added to each one-dimensional sequence, and then feature extraction is performed through multiple multi-head attention modules in sequence. The feature extraction results of each feature block are then concatenated to obtain the deep features of the input image. The shallow features are fused with the deep features through a residual connection mechanism to obtain fused features; Generate hash codes using the obtained fusion features through hash functions; Based on the generated hash codes and fusion features, a hierarchical search strategy from coarse to fine is adopted to search the silk image dataset, and the loss function of the network model is calculated according to the search results, so as to iteratively update the network parameters.

3. According to claim 2, a method for processing long-tail distribution of silk images based on visual Transformer is characterized in that: In the process of generating the hash code, the function tanh(x) is used as the activation function.

4. According to claim 2, a method for processing long-tail distribution of silk images based on visual Transformer is characterized in that: The calculation process of the loss function of the network model includes the following steps: The cosine similarity of the hash codes corresponding to the same input image after different expansions is calculated through the self-distillation scheme to obtain the self-distillation loss function value; Based on the hash code, the hash code is input into the classification layer of the network model to obtain the corresponding semantic label, thereby calculating the hash proxy loss function value; By estimating the binary likelihood value of the hash code and comparing it with the corresponding binary likelihood label, the quantization loss function value is calculated; The loss function of the network model is constructed based on the self-distillation loss function value, the hash proxy loss function value and the quantization loss function value.

5. According to claim 4, a method for processing long-tail distribution of silk images based on visual Transformer is characterized in that: The calculation expression of the self-distillation loss function value is: L SdH (h T ,h S )=1-S(h T ,h S ) Where, L SdH is the self-distillation loss function value, h T ,h S It represents the hash code corresponding to the same sample after different expansions, S(h T ,h S ) is the cosine similarity between two hash codes.

6. The method for processing long-tail distribution of silk images based on visual Transformer according to claim 4, characterized in that: The calculation process of the hash proxy loss function value includes: Use a set of trainable hashing proxies P θ , based on hash proxy P θ Representation learning of , using hash proxy P θ and the hash code h of the input image T To calculate the class-level classification prediction, the calculation expression of the class-level classification prediction is: In the formula, p T is the class-level classification prediction result, p θi is the hash proxy value assigned to the i-th class, N cls is the number of classes to be distinguished, S is the calculation of cosine similarity; Based on the class-level classification prediction results p T , by calculating the hash proxy loss to learn the similarity of the class label corresponding to the input image, the hash proxy loss function value is obtained, and the calculation expression of the hash proxy loss function value is: L HP (y,p T ,τ)=H(y,softmax(p T / t)) Where τ is the temperature scale hyperparameter, H(u,v)=―∑ k u k logv k is the cross entropy, and the softmax operation is performed along p T The dimension of is applied, y is the class label.

7. The method for processing long-tail distribution of silk images based on visual Transformer according to claim 4, characterized in that: The calculation process of the quantization loss function value includes: Use the predefined Gaussian distribution estimator g(h) with mean m and standard deviation σ to estimate the binary likelihood of the hash code h. The corresponding calculation expression is: Based on the calculated binary likelihood value, the quantization loss function value is calculated, and the corresponding calculation expression is: Where, L bce-Q (h T ) is the quantized loss function value, H b =―ulogv+(1―u)log(1―v) is the binary cross entropy, is the estimated likelihood value of the kth hash code element, is the binary likelihood label, 8. The method for processing long-tail distribution of silk images based on visual Transformer according to claim 2, characterized in that: During the updating process of the network parameters, the network parameters corresponding to the visual Transformer network remain frozen.

9. The method for processing long-tail distribution of silk images based on visual Transformer according to claim 2, characterized in that: The specific process of the hierarchical search strategy from coarse to fine includes: Rough range search step: The hash code corresponding to the input image is calculated with the hash codes corresponding to all images in the silk image dataset for Hamming distance, and the images are sorted from small to large according to the calculation results, and the first m images in the sorting results are added to the candidate image pool; Precise range retrieval step: Calculate the binary classification result between the feature vector of the input image and the feature vectors corresponding to all images in the candidate image pool, use the binary classification result as a measure of similarity between images, sort all images in the candidate image pool according to the measure of similarity, and select the first k images in the sorting result as the final retrieval result.

10. The method for processing long-tail distribution of silk images based on visual Transformer according to claim 1, characterized in that: The silk image dataset contains multiple types of silk images and covers different silk textures, patterns, colors and texture characteristics; the label distribution in the silk image dataset shows a long-tail effect, that is, the number of samples of specific patterns or rare silk types is small, and the number of samples of common categories is large; The method also includes performing data augmentation on the acquired silk image dataset to increase the number and diversity of samples in the tail category.