Fine-grained costume image retrieval method and device based on large language model common knowledge injection
By introducing the attributes generated by large language models into the clothing image retrieval system, and combining fine-grained visual features, the problem of unknown or missing attributes in open-world scenarios is solved, and fine-grained clothing image retrieval with high accuracy and robustness is achieved.
Patent Information
- Application Number
- CN202510671212.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-23
AI Technical Summary
Existing clothing image retrieval methods cannot effectively deal with unknown or missing attributes in open-world scenarios, resulting in a degradation in search of clothing products and the inability to achieve fine-grained clothing product search.
Using a method based on the injection of common sense knowledge of large language model (LLM), fine-grained clothing image retrieval is achieved through image feature representation, attribute and common sense knowledge representation, robust fusion of attribute knowledge under modal loss, and attribute-guided cross-modal reasoning.
It improves the accuracy and robustness of the clothing image retrieval system in open scenarios, can effectively deal with unknown or missing attributes, and improves the ability to adapt to incomplete information input in multiple modes.
Smart Images

Figure CN120196777A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and information retrieval, and particularly to an open-scene fine-grained clothing image retrieval method and device based on the injection of common sense knowledge of large language models (LLMs). Background Art
[0002] Clothing image retrieval has a wide range of applications in various e-commerce platforms, including clothing product recommendations, fashion trend prediction, and detection of copied clothing products. Traditional clothing image retrieval methods usually retrieve visually similar products by measuring the global similarity between images and adopt a shared embedding space. Although this method performs well in overall visual matching, it cannot effectively capture the fine-grained visual features that are crucial for practical applications. Especially when performing specific design element retrieval or copied product detection, traditional methods often cannot provide a sufficiently fine match, resulting in inaccurate search results or not meeting user needs.
[0003] To address this issue, in recent years, the Attribute-Specific Fashion Retrieval (ASFR) task has been proposed, aiming to achieve more fine-grained clothing product search by retrieving based on specific attribute values. Different from traditional retrieval methods, ASFR focuses on retrieving products based on given attributes (such as skirt length, clothing color, etc.), which enables users to retrieve clothing products that exactly match the query attributes. However, existing ASFR methods usually assume that all attribute values are predefined and known in the training dataset, and this assumption has limitations in the actual open-world scenario. In the open-world scenario, many attributes do not appear in the training set, resulting in a significant reduction in the retrieval performance of existing methods when encountering unknown or missing attributes. Therefore, ASFR methods exhibit low generalization ability when facing cross-domain and open-set application scenarios.
[0004] Therefore, a new method is needed to handle unknown attributes, missing attributes, and complex multi-modal information in the open-world ASFR scenario, so as to improve the robustness and accuracy of the clothing image retrieval system in practical applications. Summary of the Invention
[0005] To address the limitations of the existing technology, the present invention proposes an open-scene fine-grained clothing image retrieval method and device based on the injection of common sense knowledge of large language models (LLMs). For the Attribute-Specific Fashion Retrieval (ASFR) task proposed in the background art, this method is no longer limited to global visual similarity, but achieves precise retrieval through fine-grained attribute matching. Specifically, given an input image and a specific fashion attribute (such as skirt length), the system can retrieve a list of images that best match this attribute.
[0006] The object of the present invention is achieved by the following technical solutions: In the first aspect, the present invention provides a fine-grained clothing image retrieval method for injecting common sense knowledge of large language models, and the method includes the following steps:
[0007] (1) Image feature representation: Obtain clothing images, and extract fine-grained visual features based on the image encoder of CLIP;
[0008] (2) Attribute and common sense knowledge representation: Map the given attributes to attribute embedding vectors, and generate commonsense descriptions using large language models;
[0009] (3) Robust fusion of attribute knowledge under modality absence: Based on the modality configuration vector, judge the availability of the current modality, including complete modality, attribute-only modality or context-only modality; when a certain modality is absent, supplement the missing modality information through an interpolation mechanism with default values or proxy embeddings;
[0010] (4) Attribute-guided cross-modal reasoning: Construct an attribute-guided query vector, align the fine-grained visual features with the attribute information, and generate fine-grained image retrieval results according to the similarity.
[0011] Further, in step (1), the image encoder is optimized through a low-rank adapter LoRA, and the image is split into multiple non-overlapping patches to generate the embedding representation of each patch.
[0012] Further, in step (2), enhanced text embeddings are obtained through a large language model to enhance the semantic representation of the attributes, and the attribute embeddings and the enhanced context jointly form a conditional query vector to guide the retrieval process.
[0013] Further, in step (3), a prompt vector is defined to encode the current modality configuration. When a modality is unavailable, an interpolation mechanism is introduced to replace the missing embedding with a trainable default value or proxy text.
[0014] Further, in step (4), the attribute embeddings and the attribute-enhanced context embeddings are combined to construct an attribute-guided query vector to guide the retrieval process and accurately capture the visual content related to the query attributes.
[0015] Further, in step (4), cross-modal alignment is performed between the attribute-guided query vector and the patch features extracted by the image encoder, the images are weighted by calculating the similarity between the query and the image patches, and finally image features matching the query are generated.
[0016] Furthermore, a triplet loss function is used to optimize the model, reducing the matching distance between the query vector and the positive sample image while increasing the distance from the negative samples. Through this optimization strategy, the query vector is aligned with the target attribute image, and it is ensured that the most relevant image to the query attribute can be accurately selected during retrieval.
[0017] In a second aspect, the present invention also provides a fine-grained clothing image retrieval device for injecting commonsense knowledge of a large language model, including a memory and one or more processors. Executable code is stored in the memory, and when the processor executes the executable code, the described fine-grained clothing image retrieval method for injecting commonsense knowledge of a large language model is implemented.
[0018] In a third aspect, the present invention also provides a computer-readable storage medium with a program stored thereon. When the program is executed by a processor, the described fine-grained clothing image retrieval method for injecting commonsense knowledge of a large language model is implemented.
[0019] In a fourth aspect, the present invention also provides a computer program product including a computer program. When the computer program is executed by a processor, the described fine-grained clothing image retrieval method for injecting commonsense knowledge of a large language model is implemented.
[0020] Advantages of the present invention:
[0021] 1. The present invention innovatively combines the attribute-enhanced context knowledge generated by the LLM with fine-grained visual features, effectively addressing the problem of unknown or missing attributes in open-world scenarios and enhancing the accuracy and robustness of the retrieval system in open scenarios.
[0022] 2. The present invention does not rely on traditional attribute predefinitions, greatly improving the adaptability to unknown attributes and being able to handle incomplete information in multimodal inputs, having broad practical application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is a schematic diagram of the fine-grained clothing image retrieval for the open scenario of the present invention.
[0024] Figure 2 It is a flowchart of the fine-grained clothing image retrieval based on the injection of commonsense knowledge of a large language model of the present invention.
[0025] Figure 3 It is a performance comparison diagram of the open-world ASFR task of the present invention.
[0026] Figure 4 It is a performance comparison diagram of the cross-domain open-world ASFR task of the present invention.
[0027] Figure 5Structural diagram of a fine-grained clothing image retrieval device for injecting common sense knowledge into large language models provided by the present invention. Detailed implementation manners
[0028] The present invention will be described in detail below with reference to the accompanying drawings and specific implementation manners.
[0029] To address the limitations of traditional clothing image retrieval methods in the face of unknown and missing attributes in open-world scenarios, the present invention proposes an open-scene fine-grained clothing image retrieval method based on injecting common sense knowledge of large language models (LLMs). By fusing the attribute-enhanced context knowledge generated by LLMs and fine-grained visual features, the present invention improves the robustness and accuracy of the retrieval system in cross-modal, multi-modal, and open-world scenarios. As Figure 1 and Figure 2 shown, the specific steps are as follows:
[0030] (1) Image feature representation
[0031] First, a CLIP-based image encoder is used to process the input image to extract fine-grained visual features. The image encoder is optimized through a low-rank adapter (LoRA). The image is split into multiple non-overlapping patches, and an embedded representation of each patch is generated, enhancing the patch-level representation ability of the image. The output of image encoding is:
[0032]
[0033] Among them, represents the visual encoding function, K represents the number of non-overlapping patches in the image, and d represents the embedded dimension of each patch. This patch-level representation is crucial for identifying and aligning specific visual regions related to attributes, laying the foundation for subsequent cross-modal interaction and fine-grained retrieval.
[0034] (2) Attribute and common sense knowledge representation
[0035] To enable the model to handle attributes in the open world, including those unseen in the training dataset, the present invention enriches information by extracting the original attribute representations and their commonsense knowledge from a pre-trained large language model (LLM). Specifically, the present invention uses the GPT-4o model API interface provided by OpenAI as a means to enhance the commonsense of attributes. The system first inputs the given attributes (such as "sleeve length", "lapel design", etc.) into GPT-4o in the form of natural language prompts. For example, the input format is: "As a fashion design and styling expert, your task is to create a detailed description of a specific fashion attribute to support fine-grained fashion retrieval. The attribute is: sleeve length. Please clearly explain how this attribute affects the appearance, functionality, and overall style of the clothing. Highlight its aesthetic significance, practical implications, and interactions with other design elements or contexts. Ensure that your explanation is well-structured, insightful, easy to understand, and provides meticulous details to enhance the semantic representation of this attribute." The natural language text returned by GPT-4o contains a commonsense description of the attribute, such as: "Sleeve length determines the degree of arm coverage and affects the overall balance between exposure and coverage of the clothing." More samples can be referred to Table 1 in the specification. Subsequently, the specified attribute and its corresponding commonsense text description are jointly input into the attribute-aware context encoder module to embed these heterogeneous inputs into a shared latent space for multimodal complementary information integration. It should be noted that the GPT-4o model used in this method is only invoked for inference through the API interface, without involving model retraining, and has good implementability and convenience for engineering deployment.
[0036] (2-1) Original attribute representation
[0037] The original attribute is mapped to a d-dimensional embedding vector through a fully connected layer and a tanh activation function, denoted as:
[0038]
[0039] where represents the attribute embedding function. This process bridges the gap between categorical attributes and high-dimensional latent representations, mapping attributes from simple categorical labels to more expressive high-dimensional semantic embedding vectors, which can effectively integrate visual and text modalities.
[0040] (2-2)Attribute-aware context knowledge representation
[0041] To enrich attribute-related information in the unseen situation, the present invention associates each attribute Expand to a comprehensive text description. Table 1 shows the text prompts designed in the present invention, which can capture the aesthetic and functional dimensions of a given attribute.
[0042] Taking "sleeve length" as an example, common sense knowledge is generated by providing the text description "a photo of fashion with focus on sleeve length" to the text encoder of CLIP, enriching the semantic representation of the attribute. The generated attribute-enhanced context and the attribute embedding together form a conditional query vector to guide the retrieval process. This process helps improve the adaptability to unknown or missing attributes in the open world. The generated attribute-enhanced context expression is:
[0043]
[0044] where t is the text embedding enhanced by the attribute, represents the text encoding function.
[0045] Table Descriptions of several representative attributes generated by LLM
[0046]
[0047] (3) Robust fusion of attribute knowledge under modality absence
[0048] Although the construction of query q assumes that both attribute and text common sense knowledge are available, real-world scenarios usually involve flexible modalities. That is, one or more attribute information may be missing. To address these situations, the present invention introduces the following mechanism, using the vector p to indicate the available state of the current modality, including three modes: complete, attribute-only, and context-only. When a certain modality is missing, on the basis of maintaining the p prompt indication, the missing modality imputation mechanism is called to fill in the missing part using the proxy embedding a or t obtained during training, ensuring effective reasoning ability even when modalities are missing. The specific implementation process is as follows:
[0049] (3-1) Modality switching prompt mechanism
[0050] By defining a modality configuration vector, the availability of the current modality is clarified, such as the complete modality, the attribute-only modality, or the context-only modality. The modality configuration is switched according to the actual situation to ensure efficient query processing under different modality conditions. Specifically: By defining a prompt vector to clarify the encoding of the current modality configuration.
[0051]
[0052] where, is the complete modality, is the attribute-only available modality, Is a context - only available modality.
[0053] (3 - 2) Missing modality imputation mechanism
[0054] When a modality is unavailable, an imputation mechanism is introduced to replace the missing embedding with a trainable default value. For example, when text information is missing, the system uses a surrogate text embedding learned during the training process to replace the real - text description. The corresponding imputation formula is:
[0055]
[0056] where and are surrogate embedding vectors. The introduction of this mechanism ensures that the system can still maintain efficient retrieval performance in the case of missing modalities.
[0057] (4) Attribute - guided cross - modal reasoning
[0058] After obtaining the representations of attributes, attribute - aware commonsense context, and image input, the present invention designs an attribute - guided cross - modal attention mechanism for reasoning about the correlations between them for fine - grained, image - patch - based content matching. First, an attribute - aware query is constructed to provide matching guidance, and then this query is used to retrieve relevant content in the image input.
[0059] (4 - 1) Attribute - guided query construction
[0060] To retrieve an image under specific attribute - text conditions, an attribute - aware conditional query q is defined, which provides fine - grained guidance by fusing the attribute and its context embedding as follows:
[0061]
[0062] where a and t are from steps (2 - 1) and (2 - 2) respectively, and is a learnable transformation that combines these embeddings into a single vector . This query encodes the user's fine - grained requirements (i.e., attribute a and text description t) and can effectively guide the retrieval process to accurately capture the visual content related to the query attributes.
[0063] (4 - 2) Cross - modal attention interaction
[0064] After obtaining the fine - grained patch features from the image encoder (step 1), the attribute - guided query Perform cross-modal attention alignment. By calculating the similarity between the query and the image patches, the image is weighted, and finally image features matching the query are generated. Specifically, q serves as the query (Q), and the image patch feature matrix X is projected into keys (K) and values (V). For each attention head h, we have:
[0065]
[0066] where is the dimension of each head, then we have:
[0067]
[0068] where are the learnable projection parameters. The outputs of all attention heads are concatenated and linearly transformed, and finally a multi-head output is obtained. Flatten O into a vector , giving:
[0069]
[0070] where o contains the most relevant patch-level information aligned with the specified attributes and attribute-aware context. By precisely highlighting the patches related to the fine-grained query, this cross-modal attention mechanism enables condition-specific attribute retrieval with high fine-grainedness.
[0071] (5) Model Training
[0072] After constructing the query q (possibly with modal cues or imputation vectors) and calculating the condition-specific visual output o, the present invention trains the model so that the correct image matches are ranked higher than the mismatches. To make the model preferentially match the correct images over the incorrect ones, a triplet loss is defined and the property-text alignment and retrieval accuracy are balanced.
[0073]
[0074] where, is the distance function (e.g., Euclidean distance or cosine similarity), is the margin, controlling the minimum distance difference between positive and negative samples, represent the i-th query vector, positive sample feature, and negative sample feature respectively. Minimizing pushes the query embedding vector closer to the matching result of the positive sample.
[0075] The final loss function balances the property-text alignment and retrieval accuracy:
[0076]
[0077] Among them, is the alignment loss, control its relative contribution.
[0078] (6) Model Inference
[0079] During the inference process, the missing modalities (attributes or texts) are first identified. If necessary, the system applies the corresponding modality prompts and imputation strategies to generate a consistent and complete representation. Then, a query vector q is constructed, and the candidates are ranked by calculating the similarity with the candidate images. Specifically, the system calculates the similarity between the output o(q) of the query image and o(c) of the candidate images, and finally ranks the candidates by the similarity to generate fine-grained retrieval results.
[0080]
[0081] Among them, is the attention mechanism function, represents the feature of the positive sample query image, represents the feature of the candidate set images. To verify the effectiveness of the method of the present invention, the present invention is evaluated not only in the common domain ASFR tasks, but also in challenging cross-domain scenarios where the image distributions are significantly different between the training set and the test set.
[0082] Comparison within the domain: The method proposed by the present invention achieves state-of-the-art performance on three traditional ASFR benchmarks (FashionAI, DARN, and DeepFashion) (Tables 2, 3, and 4). It is worth noting that different from two-stage methods such as ASEN++ and RPF, which are vulnerable to error accumulation and computational costs, the present method adopts a unified end-to-end single-stage pipeline, which simplifies the inference process while maintaining high accuracy. In addition, using the same backbone network, such as ResNet-50, the present method always performs better among its opponents (such as ASEN, ISLN, and AttnFashion).
[0083] Table Performance comparison on the FashionAI dataset
[0084]
[0085] Table Performance comparison on the DARN dataset
[0086]
[0087] Table Performance comparison on the DeepFashion dataset
[0088]
[0089] Cross - domain comparison: The traditional ASFR model was also evaluated in cross - domain settings where the image distributions differ between datasets. Table 5 compares the method of the present invention with baselines in two transfer scenarios: FashionAI → DARN and DARN → FashionAI. Although the two datasets have similar attributes (e.g., comparing the length of clothes to the length of coats), the cross - domain shift challenges the generalization ability of traditional methods. This method, by leveraging its strong attribute - enhanced representation, outperforms the baselines and is able to achieve accurate retrieval in cross - domain retrieval.
[0090] Table Cross - domain evaluation on FashionAI → DARN and DARN → FashionAI
[0091]
[0092] Generalization to unknown attributes: The open - world ASFR task requires the model to be able to handle unknown attributes across different domains. To evaluate the effectiveness of the model of the present invention, several experiments were conducted, focusing on the generalization ability and the generalization of unknown attributes across domains.
[0093] Figure 3 Demonstrated the model's ability to handle unknown attributes, focusing on eight fashion attributes: skirt length, sleeve length, coat length, pants length, lapel design, neckline design, neckline design, and neck design, and experiments were conducted in three scenarios with gradually increasing difficulty. In scenario Figure 3 (a), sleeve length and lapel design were excluded from training. In scenario Figure 3 (b), the difficulty was increased by excluding four attributes - skirt length, sleeve length, lapel design, and neckline design. Scenario Figure 3 (c) represents the most challenging condition, where only pants length and neck design were included in the training data.
[0094] Since traditional ASFR methods rely on fixed attribute distributions, they are unable to handle unknown attributes and thus perform equivalently to a random baseline, i.e., randomly sorting all candidate images. In contrast, this method integrates the attribute - enhanced context generated by the LLM, enabling the model to learn semantically rich representations. As Figure 3 shown, this method consistently outperforms the random baseline, demonstrating its ability to effectively capture and align the semantics of unknown attributes, even in the most challenging scenario (c).
[0095] Cross - domain generalization of unknown attributes: Cross - domain generalization tests the model's ability to adapt to unknown attributes in a new domain, combining the challenges of domain transfer and unknown attribute configuration. Figure 4 Evaluated the performance of training on the DARN dataset and testing on the FashionAI dataset, where unknown attributes (e.g., skirt length, pants length, and lapel design) were introduced.
[0096] In this setting, the random baseline performs poorly because it cannot adapt to the domain - specific attributes missing in training. In contrast, our method effectively bridges the gap between domains by enhancing the context with attributes generated by the LLM, enriching the embeddings and endowing them with semantic knowledge. As Figure 4 shown, our method has achieved a significant improvement in performance, demonstrating a strong ability to adapt across domains and cross - attribute configurations. These results highlight its robustness in real - world open - world ASFR tasks.
[0097] Corresponding to the foregoing embodiment of the method for fine - grained clothing image retrieval with large language model common sense knowledge injection, the present invention also provides an embodiment of a device for fine - grained clothing image retrieval with large language model common sense knowledge injection.
[0098] See Figure 5 , an embodiment of a device for fine - grained clothing image retrieval with large language model common sense knowledge injection provided by the embodiment of the present invention includes a memory and one or more processors. Executable code is stored in the memory. When the processor executes the executable code, it is used to implement the method for fine - grained clothing image retrieval with large language model common sense knowledge injection in the above - mentioned embodiment.
[0099] An embodiment of the device for fine - grained clothing image retrieval with large language model common sense knowledge injection provided by the present invention can be applied to any device with data - processing capabilities. The any device with data - processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as a logically - defined device, it is formed by the processor of any device with data - processing capabilities reading the corresponding computer program instructions in the non - volatile memory into the memory for operation. In terms of the hardware level, as Figure 5 shown, it is a hardware structure diagram of any device with data - processing capabilities where the device for fine - grained clothing image retrieval with large language model common sense knowledge injection provided by the present invention is located. In addition to Figure 5 the shown processor, memory, network interface, and non - volatile memory, the any device with data - processing capabilities where the device in the embodiment is located usually also includes other hardware according to the actual functions of the any device with data - processing capabilities, which will not be elaborated here.
[0100] For the implementation process of the functions and roles of each unit in the above device, please refer to the implementation process of the corresponding steps in the above method for details, which will not be elaborated here.
[0101] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative work.
[0102] The embodiment of the present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements a fine-grained clothing image retrieval method with common sense knowledge injection of a large language model in the above embodiment.
[0103] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device of any device with data processing capabilities, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and can also be used to temporarily store the data that has been output or will be output.
[0104] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the fine-grained clothing image retrieval method with common sense knowledge injection of a large language model described above.
[0105] The above embodiments are used to explain the present invention, rather than limiting the present invention. Any modifications and changes made within the spirit and scope of the claims of the present invention fall within the protection scope of the present invention.
Claims
1. A fine-grained clothing image retrieval method with common sense knowledge injection from a large language model, characterized in that: The method includes the following steps: (1) Image feature representation: Obtain a clothing image and extract fine-grained visual features based on the CLIP image encoder; (2) Attribute and commonsense knowledge representation: Map the given attributes to attribute embedding vectors and generate commonsense descriptions using a large language model; (3) Robust fusion of attribute knowledge under modality absence: Based on the modality configuration vector, judge the availability of the current modality, including the complete modality, the attribute-only modality, or the context-only modality; When a modality is absent, use an interpolation mechanism to supplement the missing modality information with default values or proxy embeddings; (4) Attribute-guided cross-modal reasoning: Construct an attribute-guided query vector, align the fine-grained visual features with the attribute information, and generate fine-grained image retrieval results based on the similarity.
2. According to claim 1, the fine-grained clothing image retrieval method with common sense knowledge injection of a large language model is characterized by: In step (1), the image encoder is optimized through a low-rank adapter LoRA, and the image is split into multiple non-overlapping patches to generate the embedding representation of each patch.
3. According to claim 1, the fine-grained clothing image retrieval method with common sense knowledge injection of a large language model is characterized by: In step (2), enhanced text embeddings are obtained through the large language model to enhance the semantic representation of the attributes. The attribute embeddings and the enhanced context jointly form a conditional query vector to guide the retrieval process.
4. A fine-grained clothing image retrieval method for commonsense knowledge injection of large language models according to claim 1, characterized in that, In step (3), a prompt vector is defined to encode the current modality configuration. When a modality is unavailable, an interpolation mechanism is introduced to replace the missing embedding with a trainable default value or proxy text.
5. According to claim 1, the fine-grained clothing image retrieval method with common sense knowledge injection of a large language model is characterized by: In step (4), the attribute embeddings and the attribute-enhanced context embeddings are combined to construct an attribute-guided query vector to guide the retrieval process and accurately capture the visual content related to the query attributes.
6. According to claim 1, the fine-grained clothing image retrieval method with common sense knowledge injection of a large language model is characterized by: In step (4), the attribute-guided query vector is used for cross-modal alignment with the patch features extracted by the image encoder. By calculating the similarity between the query and the image patches, the images are weighted, and finally, image features matching the query are generated.
7. A fine-grained clothing image retrieval method for commonsense knowledge injection of large language models according to claim 1, characterized in that The triplet loss function is used to optimize the model, reducing the matching distance between the query vector and the positive sample images while increasing the distance to the negative samples. Through this optimization strategy, the query vector is aligned with the target attribute images, and it is ensured that the most relevant images to the query attributes can be accurately selected during retrieval.
8. A fine-grained clothing image retrieval device for common sense knowledge injection of large language models, comprising a memory and one or more processors, wherein executable code is stored in the memory, characterized in that, When the processor executes the executable code, it implements a fine-grained clothing image retrieval method with commonsense knowledge injection of a large language model as described in any one of claims 1-7.
9. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements a fine-grained clothing image retrieval method with commonsense knowledge injection of a large language model as described in any one of claims 1-7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements a fine-grained clothing image retrieval method with commonsense knowledge injection of a large language model as described in any one of claims 1-7.
Citation Information
Patent Citations
Cross-modal retrieval method and device based on low-rank learning
CN115186143A
Semantic knowledge guided vehicle re-identification method
CN118230321A
Modal missing RGBT tracking method and system based on missing perception prompt
CN118887592A
Cross-modal image-text retrieval method and device
CN119322861A
Cross-modal pedestrian search key semantic complete alignment method based on large model knowledge
CN119474438A
Cited By
Fabric matching method and system based on multi-modal information matching
CN120976585A
Zero sample anomaly detection method and system based on triple perception learning enhanced visual language model
CN121121767A
A zero-shot anomaly detection method and system based on triple perception learning enhanced visual language model
CN121121767B