Pattern semantic enhanced traditional pattern image-text retrieval method

Through the pattern semantic enhancement module and CLIP pre-trained model, the image encoder is trained, which solves the problem of difficult representation of traditional pattern image features and improves the accuracy and efficiency of pattern and text retrieval.

CN120067377APending Publication Date: 2025-05-30BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411971722.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Traditional patterns have many types, large size differences and complex outlines, which lead to difficult representation of image features, which brings huge challenges to cross-modal retrieval.

Method used

The graphic and text search method with pattern semantic enhancement is adopted, and the image encoder is trained through CLIP pre-training model and pattern semantic enhancement module, and the image encoder is trained, using the fill-in-the-blank task and attention mechanism to improve the encoder's ability to mine and distinguish local complex patterns.

Benefits of technology

It improves the fine feature extraction and discrimination ability of pattern images, enhances the alignment ability between patterns and text features, thereby improving the accuracy and efficiency of traditional pattern graphic and text retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067377A_ABST
    Figure CN120067377A_ABST
Patent Text Reader

Abstract

The method comprises the following steps: S1, collecting image-text data, and screening a data set image text pair obtained after preprocessing; s2, the image branch and the text branch obtain parameters through a CLIP pre-training model image and a text encoder; s3, training the model by adopting a pattern semantic enhancement module; according to the method, on the basis of the strong migration capability of the pre-training model CLIP, a large amount of existing knowledge is applied to traditional pattern image-text data, an image encoder is trained through a pattern semantic enhancement module, and the model is guided to carry out a selection blank filling task by inputting a text lacking pattern information, so that the retrieval result is obtained. A matched pattern is selected from a pre-constructed pattern corpus to complement a text, the module improves the mining capability and discrimination capability of an encoder on local complex patterns in an image, and the model extracts fine features of the patterns on one hand and has certain discrimination capability on various patterns on the other hand, so that the patterns are aligned with text features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer image processing, and specifically to a method for retrieving traditional pattern images and texts with enhanced pattern semantics. Background Art

[0002] With the rapid development of deep learning, cross-modal intelligent applications in various fields have gradually matured. Compared with the retrieval of single-modal data by traditional search engines, cross-modal retrieval is more flexible and can better meet the retrieval requirements of the vast amount of multi-modal data in the current Internet era. In the field of cultural resources, cross-modal retrieval will contribute to the digitization of traditional Chinese patterns, promote the cross-integration of different disciplinary fields, and provide new ideas and methods for the protection and inheritance of cultural heritage.

[0003] Deficiencies of the Prior Art:

[0004] Currently, there are many types of traditional patterns, with large size differences and complex outlines, which pose great challenges to the feature representation of images. Even for the same pattern, there may be significant differences in its outline, background color, and material carrier. This diversity and complexity bring huge challenges to the production of data sets and the implementation of cross-modal retrieval algorithms. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for retrieving traditional pattern images and texts with enhanced pattern semantics to solve the problems raised in the above background art.

[0006] To achieve the above purpose, the present invention provides the following technical solution: A method for retrieving traditional pattern images and texts with enhanced pattern semantics, and this image and text detection method specifically includes the following steps:

[0007] S1. Collect image and text data, and screen the image-text pairs obtained after preprocessing;

[0008] S2. The image branch and the text branch obtain parameters through the image and text encoders of the CLIP pre-trained model;

[0009] S3. Train the model using the pattern semantic enhancement module;

[0010] S4. Calculate the retrieval results.

[0011] Preferably, the step S1 specifically includes the following steps:

[0012] a1. Divide the pattern data set into a training set and a test set according to a ratio of 8:2;

[0013] a2. Split the text on the original data set and randomly replace and combine each part, that is, replace some entries in a piece of text with entries that do not correspond to the image.

[0014] Preferably, the step S2 specifically includes the following steps:

[0015] b1. Initialize the two towers respectively using the existing image-text pre-trained model, and freeze the parameters on the image side;

[0016] b2. Unfreeze the parameters on the image side, and fine-tune the Chinese native image and text data through contrastive learning.

[0017] Preferably, the step S3 specifically includes the following steps:

[0018] c1. For the fill-in-the-blank task, the pattern name in the text dataset is blocked to construct the "question text", and the correct answer is the missing pattern name;

[0019] c2. The image and the question text extract features through their respective encoders, and select the correct answer from the pre-constructed pattern corpus.

[0020] Preferably, the step c2 includes:

[0021] d1. Obtain the feature vector T through the text encoder q , and the clothing image is input into the image encoder to obtain the feature vector Iv;

[0022] d2. Use the question text feature Tq as the query, the visual feature Iv as the key-value pair, and obtain the predicted pattern feature using the attention mechanism. The formula is as follows:

[0023] ω att = f(Attention(T q , I v , I v ))

[0024] Multiple attention features are iterated to obtain the final predicted feature representation ω predict , input the pattern phrase in the pattern corpus into the text encoder to obtain the feature representation ω option , if the pattern corresponding to the option is the correct answer, the training goal is to make the similarity of the two features large; if it is a wrong option, make the similarity small. The loss function is as follows:

[0025]

[0026] where τ is a hyperparameter, ω option+ represents the pattern name feature of the correct option, represents the feature representation of other patterns in the corpus, C is the total number of nouns in the pattern corpus; in the m-th layer of the module, receive the cross-modal attention representation and calculated from the cross-modal attention representation The feature combination with the (m - 1)-th layer is continuously input into the next layer, and a complete predicted feature representation is output:

[0027]

[0028] Preferably, in step S4: after extracting features through a text encoder, a similarity calculation is performed with image features, and the TopK with the highest similarity score is taken as the retrieval result. Assume that the text input by the user is processed by the text encoder to obtain a feature vector T query , and the cosine distance is used to calculate the similarity between the query feature and the image library features. The calculation formula is as follows:

[0029]

[0030] where I feature is the feature of a certain image in the image library.

[0031] Preferably, in step d2, the traditional pattern corpus is constructed based on the cataloging system of fine classification of cultural relics and used as prior knowledge to guide model training.

[0032] Preferably, in step d2, generally τ = 0.07.

[0033] Preferably, in step S4, the cosine similarity is used to calculate the similarity between the query feature and the feature library.

[0034] Compared with the prior art, the beneficial effects of the present invention are:

[0035] For the method for retrieving traditional pattern images and texts with enhanced pattern semantics, based on the powerful transfer ability of the pre-trained model CLIP, a large amount of existing knowledge in it is applied to traditional pattern image and text data, and the image encoder is trained through a pattern semantics enhancement module. By inputting text lacking pattern information, the model is guided to perform a fill-in-the-blank task, and matching patterns are selected from a pre-constructed pattern corpus to complete the text. This module improves the ability of the encoder to mine and discriminate local complex patterns in images. On the one hand, the model extracts fine features of patterns, and on the other hand, it has a certain discrimination ability for various patterns, so as to align with text features. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 is the overall method flow chart of the present invention;

[0037] Figure 2 is the overall method effect diagram of the present invention;

[0038] Figure 3 is the overall method effect diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0040] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation to the present invention.

[0041] In the description of this patent, it should be noted that unless otherwise clearly specified and defined, the terms "installation", "connection", "setting" should be understood in a broad sense. For example, it can be fixedly connected and set, or detachably connected and set, or integrally connected and set. For those of ordinary skill in the art, the specific meanings of the above terms in this patent can be understood according to specific situations.

[0042] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, the meaning of "several" is two or more unless otherwise clearly and specifically defined.

[0043] Embodiment

[0044] Please refer to Figures 1-3 As shown, a technical solution of a traditional pattern image-text retrieval method with enhanced pattern semantics provided by the present invention:

[0045] S1. Collect image-text data. Collect cultural relic image-text data from museum and website channels. The dataset obtained after screening and preprocessing contains 25,760 pairs of image-text pairs. The data types include traditional costumes, embroidery, plum vases, celadon, etc. Under traditional costumes, there are detailed classifications, such as dragon robes, mandarin jackets, patterned fabrics, waistcoats, and sachets.

[0046] a1. Divide the pattern dataset into a training set and a test set according to a ratio of 8:2.

[0047] a2. Split the text on the original dataset and randomly replace and combine each part, that is, replace some entries in a piece of text with entries that do not correspond to the image. For example, replace "Ming yellow kesi wisteria-patterned cotton cloak" with "Great red kesi hundred-flower-patterned cotton cloak". Four negative sample texts are generated for each picture, so that each picture has five corresponding texts;

[0048] S2. The image branch and the text branch obtain parameters through the image and text encoders of the CLIP pre-trained model. There are two types of CLIP image encoders, the ResNet network based on CNN and the Visual Transformer network based on Transformer. This method uses the ViT B / 16 network. The pre-trained ViT model is an efficient feature extractor and can complete other downstream tasks. This method tries to keep the overall parameters of the pre-trained model unchanged and only connects a pattern perception module after the output layer, which not only utilizes the feature extraction ability of the existing model but also optimizes for the scenario of this article;

[0049] b1. Initialize the two towers using the existing image-text pre-trained model and freeze the parameters on the image side;

[0050] b2. Unfreeze the parameters on the image side and fine-tune the Chinese native image and text data through contrastive learning;

[0051] S3. Train the model using the pattern semantic enhancement module;

[0052] c1. The fill-in-the-blank task will obscure the pattern name in the text dataset to construct the "question text", and the correct answer is the missing pattern name;

[0053] c2. The image and the question text extract features through their respective encoders and select the correct answer from the pre-constructed pattern corpus;

[0054] d1. Obtain the feature vector Tq through the text encoder, and input the clothing image into the image encoder to obtain the feature vector Iv;

[0055] d2. Use the question text feature Tq as the query, the visual feature Iv as the key-value pair, and use the attention mechanism to obtain the predicted pattern feature. The formula is as follows:

[0056] ω att =f(Attention(T q ,I v ,I v ))

[0057] Multiple attention features are iteratively used to obtain the final predicted feature representation ω predict , and input the pattern phrases in the pattern corpus into the text encoder to obtain the feature representation ω option, if the pattern corresponding to the option is the correct answer, the training goal is to maximize the similarity between the two features; if it is a wrong option, minimize the similarity. The loss function is as follows:

[0058]

[0059] where τ is a hyperparameter, typically τ = 0.07, ω option+ represents the feature of the pattern name of the correct option, represents the feature representation of other patterns in the corpus. C is the total number of nouns in the pattern corpus, which prompts the image encoder to mine the semantic information of local patterns. The pattern semantic enhancement module receives multiple text-image modality feature representations and attempts to restore the text features based on the visual features of local patterns. The first layer uses the pre-trained encoder introduced above, and each subsequent layer is a stacked Transformer encoder, which makes full use of the attention mechanism to mine the pattern data features. In the m-th layer of the module, it receives the text-image features from the same layer and to calculate the cross-modal attention representation which is combined with the features of the (m - 1)-th layer and then input into the next layer to output the complete predicted feature representation:

[0060]

[0061] S4. Calculate the retrieval result;

[0062] After extracting features through the text encoder, calculate the similarity with the image features, and take the TopK with the highest similarity score as the retrieval result. Assume that the text input by the user is processed by the text encoder to obtain the feature vector T query , and use the cosine distance to calculate the similarity between the query feature and the image library features. The calculation formula is as follows:

[0063]

[0064] where I feature is the feature of an image in the image library.

[0065] The working principle of the present invention is as follows:

[0066] When this embodiment is in use, the method includes: S1, collecting graphic and text data, and screening the dataset image-text pairs obtained after preprocessing; S2, obtaining parameters for the image branch and the text branch through the image and text encoders of the CLIP pre-trained model; S3, training the model using the pattern semantic enhancement module; S4, calculating the retrieval results. Based on the powerful transfer ability of the pre-trained model CLIP, a large amount of existing knowledge in it is applied to traditional pattern graphic and text data, and the image encoder is trained through the pattern semantic enhancement module. By inputting the text lacking pattern information, the model is guided to perform a fill-in-the-blank task, and matching patterns are selected from the pre-constructed pattern corpus to complete the text. This module improves the ability of the encoder to mine and discriminate local complex patterns in the image. On the one hand, the model extracts fine features of the patterns, and on the other hand, it has a certain discrimination ability for various patterns, so as to align with the text features.

[0067] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and all these changes and improvements fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.

Claims

1. A traditional pattern image and text retrieval method with pattern semantic enhancement, characterized by: The image and text detection method specifically includes the following steps: S1, collect image and text data, and filter the image and text pairs obtained after preprocessing; S2, the image branch and the text branch obtain parameters through the CLIP pre-trained model image and text encoder; S3, using pattern semantic enhancement module to train the model; S4. Calculate the search results.

2. According to the method for traditional pattern image and text retrieval with pattern semantic enhancement in claim 1, it is characterized by: The step S1 specifically includes the following steps: a1. Divide the pattern dataset into training set and test set in a ratio of 8:2; a2. Split the text in the original data set and randomly replace and combine each part, that is, replace some entries in a text with entries that do not correspond to the image.

3. According to the method of claim 1, the method is characterized by: The step S2 specifically includes the following steps: b1. Use the existing image and text pre-training model to initialize the two towers respectively and freeze the image side parameters; b2. Unfreeze the image side parameters and fine-tune the native Chinese image and text data through comparative learning.

4. According to the method for traditional pattern image and text retrieval with pattern semantic enhancement as claimed in claim 1, it is characterized by: The step S3 specifically comprises the following steps: c1. Selecting the fill-in-the-blank task will obscure the pattern names in the text dataset to construct the "question text", and the correct answer is the missing pattern name; c2. The image and question text are passed through their respective encoders to extract features and select the correct answer from a pre-built pattern corpus.

5. The traditional pattern image and text retrieval method with pattern semantic enhancement according to claim 4 is characterized by: The step c2 comprises: d1. Get the feature vector T through the text encoder q , the clothing image is input into the image encoder to obtain the feature vector Iv; d2. Take the question text feature Tq as the query and the visual feature Iv as the key-value pair, and use the attention mechanism to obtain the predicted pattern features. The formula is as follows: ω att =f(Attention(T q ,I v ,I v )) Multiple attention features are iterated to obtain the final prediction feature representation ω predict , input the pattern phrases in the pattern library into the text encoder to obtain the feature representation ω option , if the pattern corresponding to the option is the correct answer, the training goal is to make the similarity of the two features large; if it is an incorrect option, the similarity is small, and the loss function is as follows: Among them, τ is a hyperparameter, ω option+ The pattern name features that indicate the correct option, represents the feature representation of other patterns in the corpus, C is the total number of nouns in the pattern corpus; the mth layer of the module receives the graphic features from the same layer and Calculate the cross-modal attention representation Combined with the features of the m-1th layer, it continues to be input into the next layer and outputs a complete prediction feature representation:

6. The traditional pattern image and text retrieval method with pattern semantic enhancement according to claim 1 is characterized by: In step S4: after extracting features through the text encoder, similarity calculation is performed with the image features, and the TopK with the highest similarity score is taken as the retrieval result. Assume that the text input by the user is processed by the text encoder to obtain the feature vector T query , use the cosine distance to calculate the similarity between the query feature and the image library feature, and the calculation formula is as follows: Among them, I feature is the feature of an image in the image library.

7. The traditional pattern image and text retrieval method with pattern semantic enhancement according to claim 5 is characterized by: In step d2, the traditional pattern corpus is constructed based on the cataloging system of cultural relics subclassification, and serves as prior knowledge to guide model training.

8. The traditional pattern image and text retrieval method with pattern semantic enhancement according to claim 5 is characterized by: In the step d2, generally τ=0.

07.

9. The traditional pattern image and text retrieval method with pattern semantic enhancement according to claim 1 is characterized by: In step S4, cosine similarity is used to calculate the similarity between the query feature and the feature library.