Feature extraction and deep quantization learning method for unlabeled Huangmubi play image retrieval
Through a combination of negative data enhancement based on image blocks and deep self-supervised contrast learning, real negative samples are generated and local texture features are extracted, which solves the problem of false negative samples in label-free retrieval of Huangmei Opera images, and achieves efficient and accurate image retrieval effect.
Patent Information
- Application Number
- CN202510515625.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-04-23
AI Technical Summary
Among the massive Huangmei Opera image data, it is difficult for the prior art to efficiently and accurately search for label-free images, especially due to the complex content and diverse characteristics of Huangmei Opera images.
The negative data augmentation based on image blocks is used to generate real negative samples, combined with deep self-supervised comparison learning and accumulation quantization methods, local texture features are extracted through an adaptive high-pass filter generator, and accurate reconstruction vectors are obtained through the deep accumulation quantization layer, improving retrieval performance.
The accuracy and efficiency of label-free Huangmei Opera image retrieval is improved, the problem of false negative samples is solved, the ability to distinguish image features is enhanced, and the retrieval accuracy is improved.
Smart Images

Figure CN120429458A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image retrieval technology, and more specifically, relates to a feature extraction and deep quantitative learning method for unlabeled Huangmei Opera image retrieval. Technical Background
[0002] With the advancement of digital technology, a vast amount of imagery has been preserved and disseminated, encompassing classic Huangmei Opera repertoires, character modeling, and stage scenes. These images not only carry a wealth of historical and cultural information but also provide valuable reference material for artistic research, education and training, and cultural and creative design. However, due to the complex content, diverse costumes, and rich scenes of Huangmei Opera images, efficiently and accurately retrieving relevant images from this massive data set is a critical task.
[0003] Nearest neighbor search methods are fundamental and essential in many technical fields, such as image retrieval, pattern recognition, and computer vision. However, when processing high-dimensional data, such as image features, exact nearest neighbor search methods incur high computational costs and large storage overhead. To achieve a good balance between retrieval efficiency and accuracy, approximate nearest neighbor search methods have been proposed.
[0004] There are two mainstream approaches to approximate nearest neighbor search: hash-based and quantization-based. Hash-based methods use a hash function to convert feature vectors into compact binary codes while maintaining the similarity of the original feature vectors. For this type of method, the Hamming distance is used to measure the distance between binary codes. However, the limitation of this type of method is that the discriminability of the Hamming distance is limited by the length of the code used, and only a limited distance value can be used to represent the similarity relationship between feature vectors. It cannot effectively measure the complex distance relationship required for similarity between feature vectors. Quantization-based methods better solve this problem. They use an asymmetric distance formed by the Euclidean distance from the query feature vector to the codebook to measure the similarity between feature vectors. While ensuring retrieval accuracy, they solve the problem of high overhead when calculating the exact Euclidean distance. Among them, the Accumulative Quantization (AQ) method is one of the best quantization methods for quantizing feature vectors. It obtains the reconstructed vector corresponding to the feature vector by accumulating M subvectors quantized by M codebooks.
[0005] In recent years, quantization-based deep image retrieval methods have introduced differentiable quantization methods on continuous deep image feature vectors, enabling the direct learning of deep representations in real-valued space. However, achieving satisfactory retrieval accuracy for supervised image retrieval methods requires expensive image annotation. Therefore, quantization-based unsupervised image retrieval methods have been proposed. These methods, without the need for image annotation, explore similarities between images to generate discriminative binary codes. In the unsupervised domain, contrastive learning is often used to learn image feature vectors without image annotation to achieve superior retrieval performance.
[0006] Based on the above background, the present invention designs a feature extraction and deep quantization learning method for unlabeled Huangmei Opera image retrieval, which combines deep self-supervised contrastive learning with a cumulative quantization method. Specifically, first, negative samples of the input Huangmei Opera image are generated using image block-based negative data augmentation, which are dissimilar to all Huangmei Opera images, so that more discriminative image feature vectors can be learned. Then, the codebook and quantized image feature vectors are learned through a deep cumulative quantization method. In addition, an adaptive high-pass filter generator is used to capture the local texture features of the Huangmei Opera image, providing more refined image information for the contrastive learning-based quantization model to further improve the retrieval performance. Summary of the Invention
[0007] In view of this, the purpose of the present invention is to propose a feature extraction and deep quantization learning method for unlabeled Huangmei Opera image retrieval, which generates true negative samples of input Huangmei Opera images through negative data enhancement based on image blocks, thereby solving the problem of false negative samples generated by randomly selecting training images in a contrastive learning framework; a fusion filter is used to extract texture information reflecting the local features of the Huangmei Opera image to enhance the feature discrimination ability of the Huangmei Opera image; and a deep cumulative quantization method is proposed to obtain more accurate reconstruction vectors and improve retrieval accuracy.
[0008] To achieve the above objectives, the specific technical solutions implemented in the present invention include:
[0009] The system consists of five modules: random data enhancement, image feature encoder, local texture feature extraction, deep accumulation quantization layer, and negative view-based quantization contrast learning.
[0010] The random data augmentation includes generating three different views for the Huangmei Opera image using different data augmentations, and the specific steps include:
[0011] All Huangmei Opera images in the batch need to go through the following processing steps: n For example:
[0012] Step A1: Randomly select two positive data augmentations, such as random cropping, rotation, color jitter, Gaussian blur, etc., denoted as t1 and t2;
[0013] Step A2: Use positive data enhancement t1 and t2 to generate Huangmei Opera image I n Generate two front views and in And uniformly expressed as
[0014] Step A3: Using the image patch-based negative data augmentation t3, randomly shuffle the image patch positions to generate the true negative view of all Huangmei Opera images
[0015] The image feature encoder includes extracting global image features and extracting multiple scale features of the image, using the Huangmei Opera image I n For example, the specific steps include:
[0016] (1) Extracting global image features
[0017] Step B1: Adjust I n The image size is converted to 224×224×3;
[0018] Step B2: Extract image I through the Vision Transformer model pre-trained by the ImageNet dataset n The global eigenvector x of n ;
[0019] (2) Extracting multiple scale features of the image
[0020] Step C1: Extract image I through U-Net network n Multiple scale feature vectors of
[0021] Step C2: Bottom-up fusion to obtain the corresponding initial fusion multi-scale information feature vector
[0022] The texture feature extraction includes extracting local texture features of Huangmei Opera images and feature fusion, with Huangmei Opera images I n For example, the specific steps include:
[0023] (1) Extracting local texture features of Huangmei Opera images
[0024] Step D1: Given an image I n The corresponding initial fusion feature vector Input into a 3×3 convolution layer to obtain the initial convolution kernel
[0025] Step D2: Input into a softmax layer to convert the initial generated convolution into a probability distribution This ensures that the weights of the low-pass filter are non-negative and normalized;
[0026] Step D3: Apply the weights of the low-pass filter Subtracting from the unit kernel E to perform the inversion operation forms a high-pass filter;
[0027] Step D4: Feature Vector After high-pass filter processing, the Huangmei Opera image I is obtained. n The local texture feature vector
[0028] (2) Feature Fusion
[0029] Step E1: Transform the local texture feature vector into Projected into D-dimensional space, The dimension is converted to D dimension;
[0030] Step E2: Use vector concatenation or orthogonal fusion algorithm to transform the global feature vector x n and local texture feature vector Fusion, get the corresponding fusion feature vector
[0031] The depth accumulation quantization layer includes quantizing the feature vector using an accumulation quantization method, and the specific steps include:
[0032] Step F1: Given a feature vector, and codebook C={C1,...,C m ,...,C M}, and Divide into M sub-vectors in sequence
[0033] Step F2: normalize all codebooks and subvectors;
[0034] Step F3: Calculate the soft quantization probability of each vector and the corresponding codebook;
[0035] Step F4: quantize each sub-vector through a soft allocation mechanism to obtain the corresponding sub-reconstruction vector;
[0036] Step F5: Accumulate each sub-reconstruction vector to obtain The corresponding reconstruction vector z n ;
[0037] The quantitative contrast learning based on the negative view includes the quantitative contrast learning of the joint original image and the negative view and the quantitative contrast learning of the joint positive view and the negative view. n For example, the specific steps include:
[0038] (1) Quantized contrastive learning of the original image and the negative view
[0039] Step G1: Given an input Huangmei Opera image I n The fusion feature vector and the global feature vectors of the three views generated by it And through the depth accumulation quantization layer, the corresponding reconstructed vector z is obtained n ,
[0040] Step G2: Calculate z separately n and and similarity between
[0041] Step G3: Calculate z n The reconstructed vector corresponding to the forward view generated from other images and adopts the perceptual bias framework to reduce Zhong and z n Similar vectors have a negative impact on the model training process;
[0042] Step G4: Based on the similarity calculated above, obtain the quantized contrast loss L oib ;
[0043] (2) Quantized contrastive learning of joint positive and negative views
[0044] Step H1: Given an input Huangmei Opera image I n Generated global feature vectors of different views And through the depth accumulation quantization layer, the corresponding reconstructed vector is obtained
[0045] Step H2: Calculate separately and and The similarity between them, where y = 1, On the contrary, when y=2,
[0046] Step H3: Calculation The reconstructed vector corresponding to the forward view generated from other images and adopts the perceptual bias framework to reduce Zhongyu Similar vectors have a negative impact on the model training process;
[0047] Step H4: Based on the similarity calculated above, obtain the quantized contrast loss L pvb ; BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 This is a flow chart of the feature extraction and deep quantitative learning method for unlabeled Huangmei Opera image retrieval of the present invention.
[0049] Figure 2 Schematic diagram of the adaptive high-pass filter generator of the present invention.
[0050] Figure 3 It is a flowchart of the depth accumulation quantization method of the present invention.
[0051] Figure 4 This is a schematic diagram of the quantitative contrast learning of the combined original image and the negative view of the present invention.
[0052] Figure 5 3 is a schematic diagram of quantitative comparative learning of the combined positive view and negative view of the present invention. Specific implementation methods
[0053] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the technical solutions, drawings and embodiments.
[0054] This paper proposes a feature extraction and deep quantitative learning method for unlabeled Huangmei Opera image retrieval, which includes five modules: random data enhancement, image feature encoder, texture feature extraction, deep cumulative quantization layer and quantitative contrast learning based on negative view. The complete process is as follows Figure 1 As shown in the figure: First, two random positive data augmentations and image block-based negative data augmentation are used to generate two positive views and one negative view for the input Huangmei Opera image. Subsequently, an image feature encoder is used, which includes a Vision Transformer model and a U-Net to extract the global feature vector and multiple scale feature vectors of the image and the corresponding view. Secondly, the multiple scale feature vectors are fused from bottom to top to obtain the initial fused feature vector, which is input into the adaptive high-pass filter generator to obtain the local texture features corresponding to the Huangmei Opera image, and then fused with the global feature vector corresponding to the image so that the feature vector contains more image information. The feature vector then passes through a deep accumulation quantization layer to obtain the corresponding reconstructed vector. Finally, through quantization contrast learning based on negative views, the similarity between related reconstructed vectors is minimized and the similarity between unrelated reconstructed vectors is maximized.
[0055] More specifically, the following Figure 1 、2 , 3, 4, and 5 provide a detailed description of the feature extraction and deep quantitative learning method for the unlabeled Huangmei Opera image retrieval of the present invention.
[0056] (1) Adaptive high-pass filter generator
[0057] The specific structure of the adaptive high-pass filter generator is as follows Figure 2 As shown in , it consists of a 3×3 convolution layer, a softmax layer, and a filter inversion operation. For example:
[0058] Step I1: Pass through a 3×3 convolution layer to obtain the initial convolution kernel The calculation formula is expressed as:
[0059] Step I2: Input into a softmax layer to convert the initial generated convolution into a probability distribution This ensures that the weights of the low-pass filter are non-negative and normalized, and the calculation formula is:
[0060] Step I3: By adding the weights of the low-pass filter Subtracting from the unit kernel E to perform the inversion operation forms a high-pass filter where k = 3 and The calculation formula is:
[0061] Step I4: Eigenvectors After high-pass filter processing, the Huangmei Opera image I is obtained. n The local texture feature vector The calculation formula is:
[0062] Step I5: Through a fully connected layer, Converted into a D-dimensional feature vector, which is consistent with the image I n The global eigenvector x of n Splicing to perform feature fusion and obtain the fused feature vector
[0063] (2) Deep Accumulation Quantization Method
[0064] The specific process of the deep accumulation quantization method is as follows Figure 3 As shown, given a codebook containing M codebooks {C1,...,C m ,...,C M}accumulated quantization head Q, where the mth codebook C m Contains K code words {c m,1 ,...,c m,k,...,c m,K} and the kth codeword in the mth codebook Fusion feature vector For example:
[0065] Step J1: Divide into M sub-vectors in sequence The mth subvector And d = D / M, and the subvector is quantized using the Mth codebook;
[0066] Step J2: In the mth d-dimensional subspace, All subvectors and C m All code words in are normalized to and c m,k For example, the specific calculation formula is as follows:
[0067]
[0068] The function norm(.) represents the vector normalization operation;
[0069] Step J3: Calculate the soft quantization probability P m ={p m,1 ,...,p m,k ,...,p m,K}, which is calculated by α-softmax. The specific calculation process is as follows:
[0070]
[0071] Where α is a non-negative parameter used to scale the input to softmax. α-softmax is a differentiable alternative to argmax, relaxing the discrete optimization of hard-coded allocations into a continuously differentiable form.
[0072] Step J4: Each sub-vector Quantized into subvectors through soft allocation mechanism The soft allocation mechanism can be formalized as a function sa(.), and the specific calculation formula is as follows:
[0073]
[0074] Step J5: Accumulate each sub-vector z n,m , thus obtaining The reconstruction vector
[0075] (3) Schematic diagram of quantitative contrast learning of the original image and the negative view
[0076] The sample pairs constructed in the quantitative contrastive learning of the joint original image and the negative view are as follows Figure 4 As shown, the reconstructed vector z n For example, the specific process is as follows:
[0077] Step K1: Calculate z separately n and and The similarity between and Where s(.) calculates the cosine similarity, N B is the batch size;
[0078] Step K2: Calculate z n The reconstructed vector corresponding to the forward view generated from other images and adopts the perceptual bias framework to reduce Zhong and z n The negative impact of similar vectors on the model training process is calculated as follows:
[0079]
[0080] Where τ2 is a non-negative parameter, ρ + is a positive prior for bias correction;
[0081] Step K3: Based on the similarity pairs calculated above, obtain the quantitative contrast loss L oib , the calculation formula is as follows:
[0082]
[0083] (4) Schematic diagram of quantitative contrast learning of joint positive and negative views
[0084] The sample pairs constructed in the quantitative contrastive learning of joint positive and negative views are as follows Figure 5 As shown, the reconstructed vector For example, the specific process is as follows:
[0085] Step L1: Calculate separately and and The similarity between and Among them, when y=1, On the contrary, when y=2,
[0086] Step L2: Calculation The reconstructed vector corresponding to the forward view generated from other images and adopts the perceptual bias framework to reduce Zhongyu The negative impact of similar vectors on the model training process is calculated as follows:
[0087]
[0088] Step L3: Based on the similarity pairs calculated above, obtain the quantitative contrast loss L pvb , the calculation formula is as follows:
[0089]
[0090] The specific embodiments described above further describe the objectives and technical solutions of the present invention in detail. It should be understood that the above are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A feature extraction and deep quantitative learning method for unlabeled Huangmei Opera image retrieval, characterized in that: include: Random data augmentation, image feature encoders, texture feature extraction, deep cumulative quantization, and contrastive learning with negative views; The random data augmentation includes: using two random positive data augmentations to generate two positive samples for the unlabeled Huangmei Opera image; and using image block-based negative data augmentation to divide the input Huangmei Opera image into M×N non-overlapping image blocks and randomly shuffle and reconstruct the image to obtain negative samples of the Huangmei Opera image; The image feature encoder includes: using the Vision Transformer model as a global feature extractor for Huangmei Opera images. This model maps the image into a high-dimensional real-valued image feature vector through a hierarchical image block embedding mechanism, thereby obtaining the image's global feature vector. Simultaneously, a U-Net is used to extract multiple scale feature vectors of the input Huangmei Opera image. The texture feature extraction includes: fusing multi-scale feature vectors extracted by the U-Net network from bottom to top to obtain an initial fused feature vector, and then inputting the initial fused feature vector into an adaptive high-pass filter generator to obtain local texture features in the Huangmei Opera image; The deep accumulation quantization layer comprises: inputting the extracted image feature vector into a deep accumulation quantization module composed of L codebooks, quantizing the quantized output vectors respectively and accumulating them to obtain the corresponding reconstructed vector; The proposed negative-view-based quantitative contrastive learning method involves designing two types of negative-view-based quantitative contrastive learning: joint quantitative contrastive learning of the original image and the negative view, and joint quantitative contrastive learning of the positive view and the negative view. By constructing true negative samples of training images, the negative impact of false negative samples on the training process of the quantitative model based on contrastive learning is reduced. Furthermore, a perceptual debiasing framework is employed to further mitigate the adverse effects of false negative samples caused by randomly selecting unlabeled training images.
2. The texture feature extraction method according to claim 1, wherein: The process of extracting texture features is as follows: First, given a training set containing N unlabeled Huangmei Opera images Extract the nth Huangmei Opera image I through U-Net n Multiple scale feature vectors of Fusion from bottom to top layer by layer to obtain corresponding multiple scale feature vectors in F up (.) indicates upsampling, and Then, a 3×3 convolutional layer is input to the feature vector that fuses multiple scale information And output the feature vector Among them, k 2 is the number of kernel channels and high-pass filters, and k is the kernel size package of the filter; Secondly, the eigenvector After a softmax layer, the initially generated convolution is converted into a probability distribution to ensure that the filter weights are non-negative and normalized. Subsequently, the kernel weights of the low-pass filter are inverted through the filter inversion operation to convert it into a high-pass filter. Finally, the eigenvector After the above high-pass filter processing, the Huangmei Opera image I is obtained. n Local texture features 3. The depth accumulation quantization layer according to claim 1, characterized in that The depth accumulation quantization process is: First, give the feature vector to be quantized It is equally divided into several sub-vectors in the dimension direction and contains M codebooks {C1,...,C m ,...,C M }accumulated quantization head Q, where And d = D / M; Then, after normalizing each sub-vector and the codeword in the codebook, the mth sub-vector Through the soft quantization mechanism according to the mth codebook C m Quantized into sub-reconstructed vectors Finally, accumulate M sub-reconstruction vectors to obtain the feature vector The reconstruction vector 4. The method of claim 1, wherein: The quantitative contrast learning process of the joint original image and the negative view is: First, in a batch, N are randomly selected from the training set I. B Huangmei Opera image training, through two kinds of random positive data enhancement (such as random cropping, rotation, color jitter, Gaussian blur) and image block-based negative data enhancement, generates 2N B forward views and N B negative views; Then, the pre-trained Vision Transformer model is used to extract the Huangmei Opera image I n (n=1,2,...,N B ) and the global feature vectors of different views are represented as x n , where x n First, local texture feature vector Fusion obtains the corresponding fusion feature vector Then, the corresponding reconstructed vector z is obtained through the depth accumulation quantization layer n , Secondly, in the contrastive learning framework, for the reconstructed vector z n , and is similar to all the and the reconstructed vector corresponding to the forward view generated from other images are not similar. At the same time, the perceptual bias framework is used to reduce Zhong and z n Similar vectors have a negative impact on the model training process; Finally, by calculating z n and the cosine similarity between its similar vector and dissimilar vector, respectively. and Thus, the quantitative contrast loss L of the joint original image and the negative view is obtained oib .
5. The method of claim 1, wherein: The quantitative contrast learning process of the joint positive view and negative view is: First, two random positive data enhancements (such as random cropping, rotation, color jittering, Gaussian blur) and image block-based negative data enhancement are used to enhance the training image I n (n=1,2,...,N B ) generate 2N respectively B forward views and N B negative views; Then, extract the Huangmei Opera image I n The global feature vectors of the three different views are expressed as And through the depth accumulation quantization layer, the corresponding reconstructed vector is obtained Secondly, in the contrastive learning framework, for the reconstructed vector Its is similar to all the and the reconstructed vector corresponding to the forward view generated from other images are not similar, where y = 1, On the contrary, when y=2, At the same time, the perceptual bias framework is used to reduce Zhongyu Similar vectors have a negative impact on the model training process; Finally, by calculating and the cosine similarity between its similar vector and dissimilar vector, respectively. and Thus, the quantitative contrast loss L of the joint positive view and negative view is obtained pvb .
Citation Information
Patent Citations
Large-similarity image similarity retrieval method and system based on deep product quantization
CN107943938A
Tea cake anti-counterfeiting method based on tea cake image feature coding
CN113379720A
Sketch image-visible light image retrieval method based on CNN and Transform
CN115908855A
Image processing model processing method and device, equipment and storage medium
CN116450870A
Remote sensing image retrieval optimization method based on automatic weight distribution and comparison hash
CN119339114A