Trademark retrieval library generation method and device based on multi-modal representation generation model
Through multimodal characterization generation model training and fine-tuning of trademark rejection data, the problems of manual dependence and accuracy in trademark search are solved, more accurate trademark search results are achieved, and retrieval accuracy and rule adaptability are improved.
Patent Information
- Application Number
- CN202510646315.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-05-20
AI Technical Summary
Existing trademark search technology relies on manual operations and accuracy depends on staff experience. The multimodal model is not effective in trademark search, and it is impossible to effectively understand the graphic and text elements in trademark images, and the trademark similarity determination rules are difficult to embed into the model.
A multimodal characterization generation model is adopted, and by obtaining trademark image-text pairs, training the visual coding module and the cross-modal generation module, fine-tuning is carried out in combination with trademark rejection data, trademark image feature vectors are generated, and initial and final trademark search libraries are constructed.
It improves the accuracy of trademark search, improves the search effect by 8%, and makes the search results more in line with trademark review rules, alleviating the cross-modal semantic gap and difficulty in embedding rules.
Smart Images

Figure CN120256664A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of trademark retrieval, and particularly to a method and device for generating a trademark retrieval library based on a multi-modal representation generation model. Background Art
[0002] In the field of trademarks, most of the picture similarity comparisons rely on manual operations. Specifically, when a new trademark is submitted, its picture needs to be compared with a large number of pictures of registered trademarks for similarity. The existing technologies usually require manual separation of the graphic elements and text elements in the submitted trademark picture, and then separate comparisons with the pictures and texts of the registered trademarks. This process requires a large amount of manual operations, and the accuracy of the comparison results often depends on the experience and carefulness of the comparison staff.
[0003] Although some existing technologies have started to adopt multi-modal models (combining images and texts for similarity detection), however, these methods generally only perform simple fine-tuning on trademark pictures and lack in-depth analysis of the characteristics of trademark pictures. Traditional deep learning methods (such as CNN, MAE, MoCo, etc.) are also used as auxiliary tools in some existing technologies for the similarity comparison of trademark pictures, but the effects are often not ideal. The reasons are as follows:
[0004] 1. Single-modal image encoding models: For example, image encoders such as ResNet, MAE, MoCo, etc. based on CNN cannot understand the specific graphic and text elements contained in trademark images. These models pay more attention to the overall structure of the image and lack semantic understanding, so their performance in trademark retrieval is poor.
[0005] 2. Multi-modal models: For example, CLIP (combining an image encoder based on ViT and a BERT text encoder), although it has made some progress in the alignment of image-text feature spaces, there is still no deep fusion. Especially due to the limitation of the BERT text understanding ability, the model cannot effectively extract the content in the image when facing abstract and complex images. Therefore, the effects of models such as CLIP in trademark retrieval are also limited.
[0006] 3. Artificial rules for trademark retrieval: The determination of trademark similarity is often affected by artificial rules such as the "Trademark Examination and Adjudication Guidelines", and it is very difficult to effectively embed these rules into the model. This results in a possible deviation between the results retrieved by the model and the results expected by humans. For example, a trademark picture may contain multiple graphic elements, but some of these graphic elements may not be important and do not need special attention during retrieval. If these unimportant elements are incorporated into the feature vector, it may lead to deviation in the retrieval results. Summary of the Invention
[0007] To this end, the present application provides a method and device for generating a trademark retrieval library based on a multi-modal representation generation model to solve the problem of inaccurate trademark retrieval results in the prior art.
[0008] To achieve the above object, the present application provides the following technical solutions:
[0009] In a first aspect, a method for generating a trademark retrieval library based on a multi-modal representation generation model includes:
[0010] Step 1: Obtain a trademark image, and label and classify the trademark image according to the Vienna Graphic Elements Classification Standard to obtain a trademark image-text pair; the trademark image includes a graphic trademark image, a text trademark image, and a combined trademark image;
[0011] Step 2: Train a multi-modal representation generation model according to the trademark image-text pair to obtain a trained multi-modal representation generation model; the multi-modal representation generation model includes a visual encoding module and a cross-modal generation module; the training process specifically includes: using the visual encoding module to extract a high-dimensional visual feature vector of the trademark image, and converting the guiding text instruction into a text vector sequence through a tokenizer; according to the trademark image-text pair, aligning the dimensions of the high-dimensional visual feature vector and the text vector sequence through a linear projection layer and then splicing and fusing them to obtain a joint feature; inputting the joint feature into the cross-modal generation module to generate a complete description of the graphic elements and text elements corresponding to the trademark image;
[0012] Step 3: Use the trained multi-modal representation generation model to generate a trademark image feature vector and input it into a vector database to obtain an initial trademark retrieval library;
[0013] Step 4: Obtain trademark rejection data, select trademarks similar to the trademark rejection data from the initial trademark retrieval library as negative samples, and select trademarks corresponding to the trademark rejection data from the rejection data as positive samples, and use the adapter technology and the triple contrast learning method to fine-tune the trained multi-modal representation generation model to obtain an optimized multi-modal representation generation model;
[0014] Step 5: Use the optimized multi-modal representation generation model to regenerate a trademark image feature vector and input it into a vector database to obtain a final trademark retrieval library.
[0015] Preferably, in step 2, the visual encoding module uses a Vision Transformer model pre-trained by CLIP.
[0016] Preferably, in step 2, the cross-modal generation module uses a pre-trained large language model based on a decoder architecture.
[0017] Preferably, in step 2, when training the multi-modal representation generation model according to the trademark image-text pair, the loss function is:
[0018]
[0019] Wherein,
[0020]
[0021]
[0022] , represent hyperparameters, L ce represents the difference between the predicted distribution and the true distribution of the model, N represents the batch size, C represents the vocabulary size, and y ic represents the true label of the i-th sample belonging to category C, represents the probability predicted by the model, represents the contrastive learning loss, represents the positive sample corresponding to x, represents the negative sample corresponding to x.
[0023] Preferably, in step 3, the vector database adopts a milvus vector database.
[0024] Preferably, step 4 specifically includes:
[0025] Step 401: Obtain trademark rejection data, select trademarks similar to the trademark rejection data from the initial trademark retrieval library as negative samples, and select the trademark corresponding to the trademark rejection data from the rejection data as the positive sample;
[0026] Step 402: Input the trademark rejection data, the positive sample, and the negative sample into the trained multi-modal representation generation model respectively to obtain multiple multi-modal representations;
[0027] Step 403: Input the multiple multi-modal representations into the adapter module respectively to obtain a positive sample feature vector, a rejected trademark feature vector, and a negative sample feature vector;
[0028] Step 404: Use the triplet contrastive learning method to perform contrastive learning on the positive sample feature vector, the rejected trademark feature vector, and the negative sample feature vector, and update the parameters of the adapter module.
[0029] Preferably, in step 404, during the contrastive learning, the loss function is:
[0030]
[0031] Among them, A represents the positive sample feature vector, P represents the trademark rejection feature vector, and N represents the negative sample feature vector. and respectively represent the distances between the anchor sample and the positive sample, and between the anchor sample and the negative sample in the embedding space. Margin represents a preset threshold for controlling the difference between the positive sample and the negative sample.
[0032] Preferably, the adapter module consists of two feed-forward sub-layers.
[0033] In a second aspect, a trademark retrieval library generation device based on a multi-modal representation generation model includes:
[0034] A trademark image-text pair construction module for obtaining trademark images and classifying and annotating the trademark images according to the Vienna Graphic Elements Classification Standard to obtain trademark image-text pairs; the trademark images include graphic trademark images, text trademark images, and combined trademark images;
[0035] A multi-modal representation generation model training module for training a multi-modal representation generation model according to the trademark image-text pairs to obtain a trained multi-modal representation generation model; the multi-modal representation generation model includes a visual encoding module and a cross-modal generation module; the training process specifically includes: using the visual encoding module to extract high-dimensional visual feature vectors of trademark images, and converting the guiding text instructions into a text vector sequence through a tokenizer; according to the trademark image-text pairs, aligning the dimensions of the high-dimensional visual feature vectors and the text vector sequence through a linear projection layer and then splicing and fusing them to obtain a joint feature; inputting the joint feature into the cross-modal generation module to generate a complete description of the graphic elements and text elements corresponding to the trademark image;
[0036] An initial trademark retrieval library construction module for using the trained multi-modal representation generation model to generate trademark image feature vectors and inputting them into a vector database to obtain an initial trademark retrieval library;
[0037] A multi-modal representation generation model optimization module for obtaining trademark rejection data, selecting trademarks similar to the trademark rejection data from the initial trademark retrieval library as negative samples, and selecting trademarks corresponding to the trademark rejection data from the rejection data as positive samples, and using the adapter technology and the triplet contrast learning method to fine-tune the trained multi-modal representation generation model to obtain an optimized multi-modal representation generation model;
[0038] A trademark retrieval library optimization module for using the optimized multi-modal representation generation model to regenerate trademark image feature vectors and inputting them into a vector database to obtain a final trademark retrieval library.
[0039] In a third aspect, a computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of a method for generating a trademark retrieval library based on a multi-modal representation generation model are implemented.
[0040] Compared with the prior art, the present application has at least the following beneficial effects:
[0041] The present application provides a method and apparatus for generating a trademark retrieval library based on a multi-modal representation generation model. By obtaining trademark images and performing annotation and classification according to the Vienna Graphic Elements Classification Standard, trademark image-text pairs are obtained; a multi-modal representation generation model is trained based on the trademark image-text pairs; trademark image feature vectors are generated using the trained multi-modal representation generation model and entered into a vector database to obtain an initial trademark retrieval library; trademark rejection data is obtained, trademarks similar to the trademark rejection data are selected from the initial trademark retrieval library as negative samples, and trademarks corresponding to the trademark rejection data are selected from the rejection data as positive samples, and the trained multi-modal representation generation model is fine-tuned using the adapter technology and the triplet contrast learning method to obtain an optimized multi-modal representation generation model; trademark image feature vectors are regenerated using the optimized multi-modal representation generation model and entered into the vector database to obtain a final trademark retrieval library. The present application performs two fine-tunings on the multi-modal representation generation model using trademark image-text pairs and trademark rejection data, which improves the overall retrieval effect of the trademark retrieval library constructed using the multi-modal representation generation model, thereby making the trademark retrieval results more accurate. Description of the Drawings
[0042] To more intuitively illustrate the prior art and the present application, exemplary drawings are given below. It should be understood that the specific shapes and structures shown in the drawings generally should not be regarded as limiting conditions when implementing the present application; for example, those skilled in the art are capable of making conventional adjustments or further optimizations to the addition / deletion / attribution division of certain units (components), specific shapes, positional relationships, connection methods, dimensional proportional relationships, etc. based on the technical concept disclosed in the present application and the exemplary drawings.
[0043] Figure 1 It is a basic flowchart of a method for generating a trademark retrieval library based on a multi-modal representation generation model provided in Embodiment 1 of the present application;
[0044] Figure 2 It is a detailed flowchart of a method for generating a trademark retrieval library based on a multi-modal representation generation model provided in Embodiment 1 of the present application;
[0045] Figure 3 It is a flowchart of training a multi-modal representation generation model using trademark image-text pairs provided in Embodiment 1 of the present application;
[0046] Figure 4 This is a flowchart for fine-tuning a multi-modal representation generation model using trademark rejection data provided in the first embodiment of this application. Detailed implementation manners
[0047] The following further details this application through specific embodiments in conjunction with the accompanying drawings.
[0048] In the description of this application: Unless otherwise specified, "a plurality of" means two or more. Terms such as "first", "second", "third", etc. in this application are intended to distinguish the objects being referred to, and do not have special significance in terms of technical connotations (for example, they should not be understood as emphasizing importance or order, etc.). Expressions such as "including", "comprising", "having", etc. also mean "not limited to" (certain units, components, materials, steps, etc.).
[0049] Terms such as "upper", "lower", "left", "right", "middle", etc. cited in this application are generally indications of the general relative position relationship for the convenience of intuitively understanding with reference to the accompanying drawings, and are not absolute limitations on the position relationship in the actual product.
[0050] First embodiment
[0051] Please refer to Figure 1 and Figure 2 , this embodiment provides a method for generating a trademark retrieval library based on a multi-modal representation generation model, which provides a systematic solution to problems such as cross-modal semantic matching, feature space alignment, and trademark rejection data adaptation optimization in the trademark examination scenario, and is applicable to scenarios such as trademark examination and brand rights protection, including:
[0052] S1: Obtain trademark images, and label and classify the trademark images according to the Vienna Classification of Graphic Elements to obtain trademark image-text pairs; the trademark images include graphic trademark images, text trademark images, and combined trademark images;
[0053] Specifically, in this step, a trademark image-text pair training set is constructed by collecting trademark images. The collected trademark images need to cover graphic trademark images (dominant types such as geometric figures and figurative patterns), text trademark images (dominant types such as pure text and font designs), and combined trademark images (composite types with mixed graphic and text arrangements).
[0054] After collecting the trademark images, label and classify the trademark images according to the Vienna Classification of Graphic Elements. The classification is accurate to the group level (the Vienna Classification includes three levels of classification: section, class, and group). The text descriptions constructed by this method are more in line with trademark examination rules and can simplify the model learning difficulty.
[0055] In the trademark image-text pairs constructed in this step, the text is composed of the graphic element names in the Vienna Classification Table and the text information contained in the figure. This method is closer to the similarity determination rules in the "Trademark Examination and Adjudication Guidelines", and the unified naming rules can make it easier for the model to learn.
[0056] S2: Train a multi-modal representation generation model based on the trademark image-text pairs to obtain a trained multi-modal representation generation model;
[0057] Specifically, in this embodiment, the multi-modal representation generation model includes a visual encoding module and a cross-modal generation module. Among them, the visual encoding module uses the Vision Transformer (ViT) model pre-trained by CLIP as an image feature extractor. Its advantage is that through contrastive learning, it can achieve a unified representation of the image-text semantic space, ensure the consistency of the hidden space between visual features and text descriptions, and lay a foundation for subsequent cross-modal interaction; the cross-modal generation module uses a pre-trained large language model based on the decoder architecture (for example: LLaMA3 8B), and this model uses the autoregressive mechanism in the Transformer architecture to perform text generation tasks.
[0058] Please refer to Figure 3 , in this step, when training the multi-modal representation generation model based on the trademark image-text pairs, a graphic-text pair supervised learning strategy is adopted, including:
[0059] S201: Use the visual encoding module to extract the high-dimensional visual feature vector of the trademark image, and convert the guiding text instruction into a text vector sequence through a tokenizer;
[0060] In this step, the trademark image is input into the ViT encoder to extract the high-dimensional visual feature vector, and at the same time, the guiding text instruction (such as "Describe this picture") is converted into a text vector sequence through a tokenizer.
[0061] S202: Align the dimensions of the high-dimensional visual feature vector and the text vector sequence through a linear projection layer according to the trademark image-text pair, and then splice and fuse them to obtain a joint feature;
[0062] To achieve the feature interaction between the visual and language modalities, in this step, the two types of feature vectors (high-dimensional visual feature vectors and text vector sequences) are dimensionally aligned through a linear projection layer and then spliced and fused to obtain a joint feature.
[0063] S203: Input the joint feature into the cross-modal generation module to generate a complete description of the graphic elements and text elements corresponding to the trademark image.
[0064] When training the multi-modal representation generation model based on the trademark image-text pair in this step, the training objective is to use the complete descriptions of the graphic elements and text elements corresponding to the trademark image as the supervision signal, and optimize the model parameters through an end-to-end progressive fine-tuning strategy.
[0065] When training the multi-modal representation generation model based on the trademark image-text pair in this step, the goal of multi-task learning is achieved by jointly optimizing the text generation loss (cross-entropy) and the feature contrast loss (InfoNCE), and its loss function is:
[0066]
[0067] Among them,
[0068]
[0069]
[0070] , denote hyperparameters. In this embodiment, , , L ce represents the difference between the predicted distribution and the true distribution of the model. N represents the batch size, C represents the vocabulary size, and y ic represents the true label of the i-th sample belonging to category C, represents the probability predicted by the model, represents the contrastive learning loss, represents the positive sample corresponding to x, represents the negative sample corresponding to x.
[0071] Finally, in this step, for the trained model, the last valid token of the last hidden state is taken as the multi-modal representation. This fusion mechanism inherits the dual-modal understanding ability of the pre-trained model, and at the same time enhances the generation effect of the trademark image features through parameter updates.
[0072] It should be noted that the training of the multi-modal representation generation model adopts a progressive training strategy. Step S2 is the first stage of the training of the multi-modal representation generation model. In this stage, the ViT parameters are frozen and only the projection layer and the decoder (the pre-trained large language model of the decoder architecture) are fine-tuned.
[0073] The reason for using the trademark graphic data to fine-tune the multi-modal representation generation model instead of using the original model with pre-trained parameters is that the original model with pre-trained parameters is trained with general data. After fine-tuning with trademark data, the model can better understand the characteristics of trademark images.
[0074] In this step, the text encoder in CLIP is replaced with a pre-trained large language model, which can improve the model's understanding of the semantics in the graph, and the generated feature vectors can better fuse the graphic elements and text information in the graph.
[0075] S3: Use the trained multi-modal representation generation model to generate trademark image feature vectors and record them in the vector database to obtain an initial trademark retrieval library;
[0076] Specifically, in this step, the fine-tuned multi-modal representation generation model is used to encode trademark images and record them in the milvus vector database to obtain an initial trademark retrieval library, which is used by the subsequent secondary fine-tuning model.
[0077] S4: Obtain trademark rejection data, select trademarks similar to the trademark rejection data from the initial trademark retrieval library as negative samples, and select trademarks corresponding to the trademark rejection data from the rejection data as positive samples, and use the adapter technology and triplet contrast learning method to fine-tune the trained multi-modal representation generation model to obtain an optimized multi-modal representation generation model;
[0078] Specifically, in this step, aiming at the characteristics of trademark retrieval, trademark rejection data is collected, and the adapter enhancement technology is used. When freezing the parameters of the above multi-modal representation generation model, by introducing a small number of additional parameters (adapter module parameters), artificial rule experience is implicitly learned from the trademark rejection data. Without changing the ability of trademark text and image fusion basically, the multi-modal representation generation model learns the potential trademark similarity judgment rules.
[0079] Please refer to Figure 4 , step S4 specifically includes:
[0080] S401: Obtain trademark rejection data, select trademarks similar to the trademark rejection data from the initial trademark retrieval library as negative samples, and select trademarks corresponding to the trademark rejection data from the rejection data as positive samples;
[0081] S402: Input the trademark rejection data (i.e., the rejected trademark image), positive samples, and negative samples into the trained multi-modal representation generation model respectively to generate multiple feature vectors (multi-modal representations);
[0082] S403: Input the generated multiple feature vectors into the adapter module respectively to obtain multiple new feature vectors, namely positive sample feature vectors, rejected trademark feature vectors, and negative sample feature vectors;
[0083] S404: Use the triplet contrast learning method to perform contrast learning on the positive sample feature vectors, rejected trademark feature vectors, and negative sample feature vectors, and update the adapter module parameters.
[0084] In this step, in the triple (Anchor, Positive, Negative), Anchor represents a randomly selected trademark image, Positive represents the corresponding rejected trademark, and Negative represents the trademarks that are relatively similar to Anchor in the initial trademark retrieval database.
[0085] The triple contrast learning strategy is as follows: for a rejected trademark image, find its corresponding trademark image in the rejection data as its positive sample, and find the top 20 trademarks that are most similar to the rejected trademark image in the initial trademark retrieval database as its negative sample set. Then, according to the training process, starting from the easiest to the most difficult, select the corresponding negative sample with the appropriate difficulty from the negative sample set.
[0086] For example: in the first two rounds, select the 20th-ranked trademark as the negative sample, and in the 3rd - 4th rounds, select the 19th-ranked trademark as the negative sample. Its training loss function is:
[0087]
[0088] Among them, A represents the feature vector of Anchor after passing through the adapter, that is, the positive sample feature vector, P represents the feature vector of Positive after passing through the adapter, that is, the rejected trademark feature vector, and N represents the feature vector of Negative after passing through the adapter, that is, the negative sample feature vector. and respectively represent the distances between the anchor sample and the positive sample, and between the anchor sample and the negative sample in the embedding space, and are specifically defined as , where margin represents a preset threshold used to control the difference between the positive sample and the negative sample. In this embodiment, margin = 0.2, and it is expected that the distance between the anchor sample and the negative sample is greater than the distance between the anchor sample and the positive sample. The purpose of doing this is to optimize the model performance by gradually increasing the difficulty of the contrast learning task.
[0089] In this embodiment, the adapter module generally consists of two feedforward sub-layers. The feedforward sub-layer can be abstracted into the following mathematical expression:
[0090]
[0091] Among them, W1, b1, W2, and b2 represent the parameters learned by the feedforward sub-layer. Its principle is to project the high-dimensional features of the original input to low-dimensional features for training and then restore them to the original feature dimension, that is, the first feedforward sub-layer projects the input dimension n (high-dimensional features) to m (low-dimensional features), where m < n. Then, in the output stage, the input dimension is restored through the second feedforward sub-layer, and m (low-dimensional features) is remapped back to n (the original high-dimensional features).
[0092] It should be noted that step S4 is the second stage of the training of the multi-modal representation generation model. In this stage, full-parameter fine-tuning is required, which can improve the retrieval effect. In this embodiment, a two-level technical architecture can be used to achieve a stepped improvement in the retrieval effect.
[0093] In this step, trademark rejection data is used, and the triple contrast learning method is used to further fine-tune the multi-modal representation generation model, enabling the model to implicitly learn trademark similarity rules and making trademark retrieval more accurate.
[0094] S5: Use the optimized multi-modal representation generation model to regenerate the trademark image feature vectors and input them into the vector database to obtain the final trademark retrieval library.
[0095] In this embodiment, the finally optimized multi-modal representation generation model is used to re-encode the trademark data and input it into the milvus vector database to obtain the final trademark retrieval library for subsequent trademark retrieval.
[0096] A method for generating a trademark retrieval library based on a multi-modal representation generation model provided by this embodiment has the following advantages:
[0097] (1) Using trademark data to fine-tune the multi-modal representation generation model enables the model to better understand trademark graphic elements and text information. Compared with the trademark retrieval library constructed using the original model, the retrieval effect is improved by 8%.
[0098] (2) Using trademark rejection data to fine-tune the model for the second time enables the model to implicitly learn trademark retrieval rules, making it more in line with the retrieval expectations of the real scenario, and further improving the overall retrieval effect of the constructed trademark retrieval library by 5%.
[0099] (3) It can effectively alleviate the core pain points such as cross-modal semantic gap and difficulty in rule embedding in the existing trademark retrieval system, and provides a reusable technical paradigm for the in-depth application of artificial intelligence in the field of intellectual property.
[0100] Embodiment 2
[0101] This embodiment provides a device for generating a trademark retrieval library based on a multi-modal representation generation model, including:
[0102] A trademark image-text pair construction module, configured to obtain a trademark image, and label and classify the trademark image according to the Vienna Graphic Elements Classification Standard to obtain a trademark image-text pair; the trademark image includes a graphic trademark image, a text trademark image, and a combined trademark image;
[0103] The multi-modal representation generation model training module is used to train a multi-modal representation generation model based on the trademark image-text pair to obtain a trained multi-modal representation generation model; the multi-modal representation generation model includes a visual encoding module and a cross-modal generation module; the training process specifically includes: using the visual encoding module to extract the high-dimensional visual feature vector of the trademark image, and converting the guiding text instruction into a text vector sequence through a tokenizer; aligning the dimensions of the high-dimensional visual feature vector and the text vector sequence through a linear projection layer according to the trademark image-text pair and then splicing and fusing them to obtain a joint feature; inputting the joint feature into the cross-modal generation module to generate a complete description of the graphic elements and text elements corresponding to the trademark image;
[0104] The initial trademark retrieval library construction module is used to generate a trademark image feature vector by using the trained multi-modal representation generation model and input it into the vector database to obtain an initial trademark retrieval library;
[0105] The multi-modal representation generation model optimization module is used to obtain trademark rejection data, select trademarks similar to the trademark rejection data from the initial trademark retrieval library as negative samples, and select the trademark corresponding to the trademark rejection data from the rejection data as positive samples, and use the adapter technology and the triple contrast learning method to fine-tune the trained multi-modal representation generation model to obtain an optimized multi-modal representation generation model;
[0106] The trademark retrieval library optimization module is used to regenerate the trademark image feature vector by using the optimized multi-modal representation generation model and input it into the vector database to obtain the final trademark retrieval library.
[0107] For the specific implementation content of each module in a trademark retrieval library generation device based on a multi-modal representation generation model, reference can be made to the definition of a trademark retrieval library generation method based on a multi-modal representation generation model in the above text, which will not be elaborated here.
[0108] Embodiment III
[0109] This embodiment provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of a trademark retrieval library generation method based on a multi-modal representation generation model are implemented.
[0110] The technical features of the above embodiments can be combined arbitrarily (as long as there is no contradiction in the combination of these technical features). For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described; these embodiments not explicitly written out should also be considered to be within the scope described in this specification.
Claims
1. A method for generating a trademark retrieval library based on a multi-modal representation generation model, characterized in that, Including: Step 1: Obtain a trademark image, and label and classify the trademark image according to the Vienna Classification of Figurative Elements standard to obtain a trademark image-text pair; the trademark image includes a graphic trademark image, a word trademark image, and a combined trademark image; Step 2: Train a multi-modal representation generation model based on the trademark image-text pair to obtain a trained multi-modal representation generation model; the multi-modal representation generation model includes a visual encoding module and a cross-modal generation module; The training process specifically includes: using the visual encoding module to extract a high-dimensional visual feature vector of the trademark image, and converting the guiding text instruction into a text vector sequence through a tokenizer; according to the trademark image-text pair, align the dimensions of the high-dimensional visual feature vector and the text vector sequence through a linear projection layer and then splice and fuse them to obtain a joint feature; input the joint feature into the cross-modal generation module to generate a complete description of the graphic elements and text elements corresponding to the trademark image; Step 4: Use the trained multi-modal representation generation model to generate trademark image feature vectors and enter them into a vector database to obtain an initial trademark retrieval library; Step 5: Obtain trademark rejection data, select trademarks similar to the trademark rejection data from the initial trademark retrieval library as negative samples, and select trademarks corresponding to the trademark rejection data from the rejection data as positive samples, and use the adapter technology and the triplet contrast learning method to fine-tune the trained multi-modal representation generation model to obtain an optimized multi-modal representation generation model; Step 6: Use the optimized multi-modal representation generation model to regenerate trademark image feature vectors and enter them into a vector database to obtain a final trademark retrieval library.
2. The method for generating a trademark retrieval library based on a multi-modal representation generation model according to claim 1, wherein In step 2, the visual encoding module uses a Vision Transformer model pre-trained by CLIP.
3. The method for generating a trademark retrieval library based on a multi-modal representation generation model according to claim 1, wherein In step 2, the cross-modal generation module uses a pre-trained large language model based on a decoder architecture.
4. The method for generating a trademark retrieval library based on a multi-modal representation generation model according to claim 1, wherein In step 2, when training the multi-modal representation generation model according to the trademark image-text pair, the loss function is: ; Where ; ; , represents a hyperparameter, L ce represents the difference between the predicted distribution of the model and the true distribution, N represents the batch size, C represents the vocabulary size, and y ic represents the true label that the i-th sample belongs to class C, represents the probability predicted by the model, represents the contrastive learning loss, represents the positive sample corresponding to x, represents the negative sample corresponding to x.
5. The method for generating a trademark retrieval library based on a multi-modal representation generation model according to claim 1, wherein In step 3, the vector database uses a milvus vector database.
6. The method for generating a trademark retrieval library based on a multi-modal representation generation model according to claim 1, wherein Step 4 specifically includes: Step 401: Obtain trademark rejection data, select trademarks similar to the trademark rejection data from the initial trademark retrieval library as negative samples, and select trademarks corresponding to the trademark rejection data from the rejection data as positive samples; Step 402: Input the trademark rejection data, the positive sample, and the negative sample into the trained multi-modal representation generation model respectively to obtain multiple multi-modal representations; Step 403: Input the multiple multi-modal representations into the adapter module respectively to obtain a positive sample feature vector, a rejected trademark feature vector, and a negative sample feature vector; Step 404: Use the triplet contrast learning method to perform contrast learning on the positive sample feature vector, the rejected trademark feature vector, and the negative sample feature vector, and update the parameters of the adapter module.
7. The method for generating a trademark retrieval library based on a multi-modal representation generation model according to claim 6, wherein In step 404, when performing the contrast learning, the loss function is: ; Wherein, A represents the positive sample feature vector, P represents the trademark rejection feature vector, and N represents the negative sample feature vector. and respectively represent the distances between the anchor sample and the positive sample, and between the anchor sample and the negative sample in the embedding space. Margin represents a preset threshold for controlling the difference between the positive sample and the negative sample.
8. The method for generating a trademark retrieval library based on a multi-modal representation generation model according to claim 6, wherein The adapter module consists of two feed-forward sub-layers.
9. A trademark retrieval library generation device based on a multi-modal representation generation model, characterized in that It includes: A trademark image-text pair construction module, which is used to obtain trademark images and label and classify the trademark images according to the Vienna Classification of the Figurative Elements, so as to obtain trademark image-text pairs; the trademark images include graphic trademark images, word trademark images, and combined trademark images; A multi-modal representation generation model training module, which is used to train a multi-modal representation generation model according to the trademark image-text pairs to obtain a trained multi-modal representation generation model; The multi-modal representation generation model includes a visual encoding module and a cross-modal generation module; The training process specifically includes: using the visual encoding module to extract the high-dimensional visual feature vectors of trademark images, and converting the guiding text instructions into a text vector sequence through a tokenizer; aligning the dimensions of the high-dimensional visual feature vectors and the text vector sequence through a linear projection layer according to the trademark image-text pairs and then splicing and fusing them to obtain joint features; inputting the joint features into the cross-modal generation module to generate a complete description of the graphic elements and text elements corresponding to the trademark images; An initial trademark retrieval library construction module, which is used to generate trademark image feature vectors by using the trained multi-modal representation generation model and input them into a vector database to obtain an initial trademark retrieval library; A multi-modal representation generation model optimization module, which is used to obtain trademark rejection data, select trademarks similar to the trademark rejection data from the initial trademark retrieval library as negative samples, and select trademarks corresponding to the trademark rejection data from the rejection data as positive samples, and use the adapter technology and the triple contrast learning method to fine-tune the trained multi-modal representation generation model to obtain an optimized multi-modal representation generation model; A trademark retrieval library optimization module, which is used to regenerate trademark image feature vectors by using the optimized multi-modal representation generation model and input them into a vector database to obtain a final trademark retrieval library.
10. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 8.
Citation Information
Patent Citations
Combined commodity retrieval method and system based on multi-modal pre-training model
CN114445201A
Combined commodity retrieval method and system based on multi-modal pre-training model
CN114840705A
ViT-based trademark retrieval method and system
CN117453938A
Image-text retrieval method and device in remote sensing field based on prompt learning
CN118690034A
Trademark multi-modal retrieval method and device based on image and category similarity and storage medium
CN119513416A
Cited By
Trademark examination method and device, electronic equipment and storage medium
CN121833919A
Image based trademark search method and apparatus
KR102966390B1