Plug-and-play conditional representation learning method and system

By generating descriptive text related to the specified criteria and encoding it as text basis, and projecting image representation into conditional feature space in combination with multimodal models, the adaptability problem of characterization learning methods in customized tasks in vertical fields is solved, and better task performance and interpretability are achieved.

CN120296194APending Publication Date: 2025-07-11SICHUAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510434921.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When facing customized tasks in vertical fields, existing representation learning methods are difficult to dynamically adapt to the needs of specific scenarios, resulting in a lack of effective expression of key information in the feature space, resulting in a bottleneck in task performance.

Method used

Descriptive text related to the specified criteria is generated through a large language model and encoded it into a standardized text basis. Combined with a multimodal model, the input image is encoded into an image representation, and projected into the conditional feature space under the specified criteria formed by the text basis, and the conditional representation used for downstream customization tasks is learned.

Benefits of technology

It realizes capturing rich semantic information under specified criteria, has good interpretability, and can be directly applied to existing representation learning methods, improving the performance of downstream customized tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296194A_ABST
    Figure CN120296194A_ABST
Patent Text Reader

Abstract

The invention discloses a plug-and-play conditional representation learning method and system. The method comprises the following steps: giving an input image and a specified criterion; generating a descriptive text related to a specified criterion through a large language model, and encoding the descriptive text into a standardized text base; encoding the input image into a standardized image representation; and projecting the image representation to a conditional feature space under a specified criterion spanned by the text base, and learning to obtain a conditional representation for a downstream customization task. Condition representation obtained through learning is more expressive under a specified criterion, has plug-and-play performance, universality and excellent interpretability, and can be suitable for various customized downstream tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of representation learning, and in particular relates to a plug-and-play conditional representation learning method and system. Background Art

[0002] We are in an era of information explosion. Processing large amounts of data quickly and efficiently has become an urgent need for industry and academia. In this context, representation learning that can autonomously extract semantic information from data has a very broad research prospect. At present, representation learning is mainly based on self-supervised learning methods. By designing various proxy tasks, it automatically mines the intrinsic structure of data without manual annotation, and then extracts discriminative representations. Specifically, typical representation learning includes: 1) Contrastive learning: by constructing positive and negative sample pairs, the distance between similar data (such as different enhanced views of the same image) in the feature space is shortened, while different types of data are pushed away, thereby learning discriminative features; 2) Mask prediction: by masking or destroying the input data (such as masking image blocks or text paragraphs), the model is trained to reconstruct the original information, forcing the model to understand the contextual dependencies and semantic structure of the data.

[0003] Although current representation learning methods have achieved excellent performance in general tasks such as classification and retrieval, their learning objectives are often limited to explicit semantic information such as object categories and geometric shapes, and are essentially still general representation paradigms for common needs. When faced with customized tasks in vertical fields (for example, animal habitat analysis needs to focus on vegetation types and landform features, and medical imaging diagnosis needs to emphasize the texture heterogeneity of pathological tissues, etc.), the limitations of general representations become increasingly obvious. It is difficult to dynamically adapt to the needs of specific scenarios, resulting in a lack of effective expression of key information in the feature space, which ultimately leads to task performance bottlenecks. This reveals the shortcomings of the current representation learning framework in terms of adaptation and interpretability to customized tasks, and highlights the urgency of evolving from "general feature extraction" to "conditional feature extraction." Especially in professional scenarios with high labeling costs and rich domain prior knowledge, how to bridge domain knowledge and data distribution through representation learning has become a key challenge for practical applications. Summary of the invention

[0004] In view of the above-mentioned deficiencies in the prior art, the plug-and-play conditional representation learning method and system provided by the present invention solves the problem that the general representation obtained by the current representation learning method mainly captures a single prominent criterion (shape or type), resulting in poor performance of customized downstream tasks that rely on other criteria.

[0005] In order to achieve the above-mentioned invention object, the technical solution adopted by the present invention is: a plug-and-play conditional representation learning method, comprising the following steps:

[0006] Given an input image and a specified criterion;

[0007] Generate descriptive text related to the specified criteria through a large language model and encode it into a standardized text base;

[0008] Encode the input image into a standardized image representation;

[0009] Project the image representation into the conditional feature space under the specified criteria spanned by the text base, and learn the conditional representation for downstream customization tasks.

[0010] Furthermore, the descriptive text W is expressed as:

[0011] W = LLM(P1, C)

[0012] In the formula, P1 represents the prompt template of the large language model, C represents the specified criteria, and LLM(·) represents the large language model.

[0013] Furthermore, the prompt template P1 of the large language model is a common expression for generating descriptions of the specified criteria C.

[0014] Furthermore, the prompt template P1 also includes the format requirements for the generated descriptive text, including the positional arrangement of the generated text elements and the uniqueness of the generated text elements.

[0015] Furthermore, the encoded text base T is expressed as:

[0016] T = VLM text (P2, C, W)

[0017] In the formula, P2 represents the prompt template of the multimodal model for encoding the descriptive text, C represents the specified criteria, W represents the descriptive text, and VLM text (·) represents the text encoder of the multimodal model for encoding the descriptive text; the prompt template P2 includes all prompt sentences determined based on the descriptive text type under the specified criteria.

[0018] Furthermore, the image representation I is expressed as:

[0019] I = VLM image (X)

[0020] In the formula, VLM image (·) represents the image encoder of the multimodal model for encoding the input image, and X represents the input image.

[0021] Furthermore, the learned conditional representation R is expressed as:

[0022] R = IT T

[0023] Wherein, I represents the image representation, T represents the text basis, and the superscript T represents the matrix transpose operation.

[0024] A conditional representation learning system, comprising:

[0025] A text basis generation module: used to generate a standardized text basis for a given input image under a specified criterion;

[0026] An image representation generation module: used to perform multimodal encoding on the input image through a multimodal model to generate an image representation;

[0027] A basis projection module: used to project the image representation into a conditional feature space under a specified criterion spanned by the text basis, and learn a conditional representation for downstream customized tasks.

[0028] Further, the text basis generation module includes:

[0029] A large language model: used to generate descriptive text related to a specified criterion for a given input image;

[0030] A multimodal model, used to perform multimodal encoding on the descriptive text to generate a standardized text basis.

[0031] Further, the prompt template when the large language model generates descriptive text related to a specified criterion is to generate common expressions describing the specified criterion, and limit the positional arrangement and uniqueness of the generated text elements.

[0032] The beneficial effects of the present invention are:

[0033] 1. Different from traditional representation learning methods that only focus on a single dominant semantics, the present invention proposes a conditional representation method that can capture semantic information under any specified criterion; for example, for an image of an animal, traditional methods only focus on its species, while the present invention can focus on more abundant information such as its quantity and habitat environment.

[0034] 2. The present invention only needs to construct a text basis under a specified criterion to complete the directional conversion of the picture representation, without the need to master prior knowledge of the specified criterion in advance, and has good interpretability; for example, under the color criterion, each dimension of the representation of the picture after conversion can be interpreted as the similarity between the picture and the color corresponding to that dimension.

[0035] 3. The present invention is plug-and-play, has excellent compatibility, and can be conveniently and efficiently applied to existing representation learning methods; for example, it can be directly combined with existing clustering or retrieval methods for personalized clustering or personalized retrieval tasks. Description of the Drawings

[0036] Figure 1Flowchart of the plug-and-play conditional representation learning method provided by the present invention.

[0037] Figure 2 Framework diagram of the conditional representation learning process provided by the present invention.

[0038] Figure 3 Schematic diagram of the personalized classification result provided by the present invention.

[0039] Figure 4 Schematic diagram of the personalized retrieval result provided by the present invention. Detailed implementation manners

[0040] The following describes the detailed implementation manners of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the detailed implementation manners. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.

[0041] An embodiment of the present invention provides a plug-and-play conditional representation learning method, as Figure 1 shown, including the following steps:

[0042] Given an input image and a specified criterion;

[0043] Generate descriptive text related to the specified criterion through a large language model and encode it into a standardized text basis;

[0044] Encode the input image into a standardized image representation;

[0045] Project the image representation into the conditional feature space under the specified criterion spanned by the text basis, and learn the conditional representation for downstream customization tasks.

[0046] In the process of implementing conditional feature learning in the embodiment of the present invention, considering the high annotation cost required for supervised learning, the method of the present invention is transformed on the basis of the general representation, so as to efficiently obtain the conditional representation applicable to customized downstream tasks. Specifically, as Figure 2In the overall framework diagram shown, given an input image and a specified criterion of the user (taking "color" as an example), first, by querying a large language model, it generates descriptive text related to the specified criterion (such as the three primary colors "red", "green", and "blue" of "color"). Then, through the text encoder and the image encoder in the multimodal model, the generated descriptive text and the input image are encoded respectively to obtain a standardized text basis and an image representation. Next, the general image representation is projected into the conditional feature space ("color" space) spanned by the text basis to obtain the transformed conditional feature; the transformed conditional feature will be more expressive under the specified criterion, suitable for customized downstream tasks, and has excellent interpretability.

[0047] In this embodiment, a basis refers to a set of linearly independent vectors that span the entire space; in a specific example, in a three-dimensional Cartesian coordinate system, the vectors (1, 0, 0), (0, 1, 0), and (0, 0, 1) representing the x, y, and z axes constitute a set of bases because any vector in this space can be represented as a linear combination of these three vectors. Similarly, in the three-color space, "red", "green", and "blue" constitute a set of bases because they can form all possible colors. From a broader perspective, a set of descriptive texts related to the user-specified criterion spans the conditional feature space based on this criterion and essentially acts as a basis; therefore, in this embodiment, the general image representation can be mapped to the required conditional feature through the text basis.

[0048] In this embodiment, the descriptive text W related to the specified criterion is generated by the large language model and is expressed as:

[0049] W = LLM(P1, C)

[0050] In the formula, P1 represents the prompt template of the large language model, C represents the specified criterion, and LLM(·) represents the large language model.

[0051] In an example of this embodiment, the large language model can be GPT-4 and DeepSeek, or other large language models that can generate descriptive text based on the prompt template, and specific limitations are not made here.

[0052] In an example of this embodiment, the specified criterion can be a criterion determined based on the input image according to the requirements of downstream customized tasks, such as "texture", "color", and "shape", etc. In this embodiment and subsequent descriptions, "color" is taken as an example of the specified criterion for illustration.

[0053] In this embodiment, the prompt template P1 of the large language model is a common expression for generating a description of the specified criterion C.

[0054] Furthermore, since large language models usually generate repetitive expressions; therefore, to avoid this situation, in this embodiment, the prompt template P1 is defined. For specific specified criteria, the prompt template P1 further includes format requirements for the generated descriptive text, including the positional arrangement of generated text elements and the uniqueness of generated text elements.

[0055] In a specific example of this embodiment, taking the specified criterion as "color" as an example, the following complete prompt template is given:

[0056] "Please generate as many common expressions describing colors as possible, with the format requirement: ["…", "…", "…"];

[0057] Ensure that all elements are written on one line and each element is unique;

[0058] Multiple expressions can be used to describe the same color, such as " Red ", " Deep Red " or " Scarlet ";

[0059] Only generate this list and do not generate extra information."

[0060] Among them, the bold part is the specified criterion, and the underlined part is an example of this specified criterion.

[0061] Based on the above method, after obtaining descriptive text that is highly semantically relevant to the specified criterion, it is encoded through the text encoder of the multimodal model to obtain a standardized text base T represented as:

[0062] T = VLM text (P2, C, W)

[0063] In the formula, P2 represents the prompt template of the multimodal model for encoding the descriptive text, C represents the specified criterion, W represents the descriptive text, and VLM text (·) represents the text encoder of the multimodal model for encoding the descriptive text; the prompt template P2 includes all prompt word sentences determined based on the descriptive text type under the specified criterion.

[0064] In an example of this embodiment, the multimodal model is a CLIP or CLIP-like model, which consists of a text encoder and an image encoder, and can encode the descriptive text and the input image respectively to obtain the corresponding text base and image representation.

[0065] In this embodiment, for the prompt template P2, taking the specified criterion as "color" and the generated descriptive text as "red", "green" and "blue" as an example, then the complete prompt template P2 input into the multimodal model is:

[0066] The color of the object is red.

[0067] The color of the object is green.

[0068] The color of the object is blue.

[0069] After encoding the above three prompt sentences, they will become three dimensions of the text basis T, and then span the color space.

[0070] It should be noted that when the prior knowledge of the dataset can be obtained, the "object" in the above sentence can be replaced with a more specific expression, such as "car", "backpack", etc.

[0071] In this embodiment, in order to achieve the representation conversion, the input image X is input into the image encoder of the multi-modal model for image encoding to obtain a standardized image representation I, which is expressed as:

[0072] I = VLM image (X)

[0073] In the formula, VLM image (·) represents the image encoder of the multi-modal model that encodes the input image, and X represents the input image.

[0074] In this embodiment, the image representation is projected into the conditional feature space under the specified criterion spanned by the text basis, and the learned conditional representation R is expressed as:

[0075] R = IT T

[0076] In the formula, I represents the image representation, T represents the text basis, and the superscript T represents the matrix transpose operation.

[0077] In this embodiment, the conditional representation learned by the above method can replace the original image representation I and be directly applied to downstream customized tasks. Since the conditional representation R is the projection of the input image in the corresponding criterion space, it has better discriminability than the image representation I under this criterion and has better performance in downstream customized tasks.

[0078] In the embodiment of the present invention, a conditional representation learning system based on the above conditional representation learning method is also provided, including:

[0079] Text basis generation module: used to generate a standardized text basis under the specified criterion corresponding to the given input image;

[0080] Image representation generation module: used to perform multi-modal encoding on the input image through a multi-modal model to generate an image representation;

[0081] Base Projection Module: It is used to project the image representation into the conditional feature space under the specified criterion spanned by the text base, and learn the conditional representation for downstream customization tasks.

[0082] In this embodiment, the text base generation module includes:

[0083] Large Language Model: It is used to generate descriptive text related to the specified criterion for the given input image;

[0084] Multimodal Model, which is used to perform multimodal encoding on the descriptive text to generate a standardized text base.

[0085] In this embodiment, the working process of the conditional representation learning system is as follows:

[0086] Given an input image and the user's specified criterion, in the text base generation module, first, by asking the large language model, it generates descriptive text related to the specified criterion, and then encodes it through the text encoder in the multimodal model to generate a standardized text base; at the same time, in the image representation generation module, the input image is encoded through the image encoder of the multimodal model to obtain a standardized image representation; finally, in the base projection module, the general image representation is projected into the conditional feature space spanned by the text base to obtain the transformed conditional feature. The conditional feature transformed through the above process will be more expressive under the specified criterion, suitable for customized downstream tasks, and has excellent interpretability.

[0087] In this embodiment, during the above working process, the prompt template for the large language model to generate descriptive text related to the specified criterion is to generate common expressions describing the specified criterion, and limit the position arrangement and uniqueness of the generated text elements.

[0088] In a specific example of this embodiment, in the text base generation module, taking the "color" of the input image as the specified criterion, the following complete prompt template P1 is given:

[0089] "Please generate as many common expressions describing colors as possible, in the format: ["…", "…", "…"];

[0090] Make sure to write all elements on one line and each element is unique;

[0091] You can use multiple expressions to describe the same color, such as "red", "dark red" or "scarlet";

[0092] Only generate this list and do not generate any extra information."

[0093] Among them, the bold part is the specified criterion, and the underlined part is an example of the specified criterion.

[0094] Based on the above prompt template P1, the descriptive text W related to the specified criterion generated by the large language model is expressed as:

[0095] W = LLM(P1, C)

[0096] Wherein, P1 represents the prompt template of the large language model, C represents the specified criterion, and LLM(·) represents the large language model.

[0097] Encoding the above descriptive text T through the text encoder of the multimodal model to obtain the standardized text base T, which is expressed as:

[0098] T = VLM text (P2, C, W)

[0099] Wherein, P2 represents the prompt template of the multimodal model for encoding the descriptive text, C represents the specified criterion, W represents the descriptive text, and VLM text (·) represents the text encoder of the multimodal model for encoding the descriptive text; the prompt template P2 includes all prompt sentences determined based on the descriptive text type under the specified criterion.

[0100] In a specific example of this embodiment, for the prompt template P2, taking the specified criterion as "color" and the generated descriptive texts as "red", "green", and "blue" as examples, then the complete prompt template P2 input into the multimodal model is:

[0101] The color of the object is red.

[0102] The color of the object is green.

[0103] The color of the object is blue.

[0104] After encoding the above three prompt sentences, they will become three dimensions of the text base T, and then span the color space. When the prior knowledge of the dataset can be obtained, the "object" in the above sentences can be replaced with a more specific expression, such as "car", "backpack", etc.

[0105] In a specific example of this embodiment, in the image representation generation module, input the input image X into the image encoder of the multimodal model for image encoding to obtain the standardized image representation I, which is expressed as:

[0106] I = VLM image (X)

[0107] Wherein, VLM image (·) represents the image encoder of the multimodal model for encoding the input image, and X represents the input image.

[0108] In a specific example of this embodiment, in the base projection module, the image representation is projected into the conditional feature space under the specified criterion spanned by the text basis, and the learned conditional representation R is expressed as:

[0109] R = IT T

[0110] In the formula, I represents the image representation, T represents the text basis, and the superscript T represents the matrix transpose operation.

[0111] In this embodiment, the conditional representation learned by the above method can replace the original image representation I and be directly applied to downstream customization tasks. Since the conditional representation R is the projection of the input image in the corresponding criterion space, it has better discriminability than the image representation I under this criterion and has more excellent performance in downstream customization tasks.

[0112] In the embodiments of the present invention, examples of applying the above conditional representation learning method to personalized classification and personalized retrieval scenarios are provided to demonstrate the plug-and-play and universality of the conditional representation learning method proposed by the present invention; for other downstream customization tasks, the method proposed by the present invention can also be applied in a similar manner.

[0113] Specifically, for the personalized classification scenario:

[0114] Personalized classification means dividing samples into different clusters according to the semantics of the samples under the specified criterion; as Figure 3 shown, the playing card data can be classified according to the two criteria of "suit" and "rank".

[0115] For the personalized classification task, based on the method of the present invention, the conditional representation can be directly obtained according to the picture data and the criterion, and then it can be used to replace the original picture representation for classification; the experimental results show that the present invention can obtain more accurate classification results.

[0116] For the personalized retrieval scenario:

[0117] Personalized retrieval means retrieving candidate pictures that have the same attributes as the query picture under the criterion given the query picture and the criterion, as Figure 4 shown.

[0118] For the personalized retrieval task, the method of the present invention can also obtain the conditional representation according to the picture data and the criterion, and then use it to replace the original picture representation for training (the training method is the same as the original retrieval method); the experimental results show that whether trained or not, using the conditional representation for retrieval can achieve more excellent retrieval performance.

[0119] In the present invention, specific embodiments are used to illustrate the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

[0120] Those of ordinary skill in the art will realize that the embodiments described herein are for helping the reader understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention according to these technical revelations disclosed by the present invention, and these deformations and combinations are still within the protection scope of the present invention.

Claims

1. A plug-and-play conditional representation learning method, characterized in that, It includes the following steps: Given an input image and a specified criterion; Generate descriptive text related to the specified criterion through a large language model and encode it into a standardized text basis; Encode the input image into a standardized image representation; Project the image representation into the conditional feature space under the specified criterion spanned by the text basis to learn the conditional representation for downstream customized tasks.

2. The method according to claim 1, wherein The descriptive text W is expressed as: W = LLM(P1, C) wherein, P1 represents the prompt template of the large language model, C represents the specified criterion, and LLM(·) represents the large language model.

3. The method according to claim 2, wherein The prompt template P1 of the large language model is a common expression for generating descriptions of the specified criterion C.

4. The method according to claim 2, wherein The prompt template P1 further includes format requirements for the generated descriptive text, including the position arrangement of the generated text elements and the uniqueness of the generated text elements.

5. The method according to claim 1, characterized in that, The encoded text basis T is expressed as: T = VLM text (P2, C, W) Wherein, P2 represents the prompt template of the multimodal model for encoding descriptive text, C represents the specified criterion, W represents the descriptive text, and VLM text (·) represents the text encoder of the multimodal model for encoding descriptive text; the prompt template P2 includes all prompt word sentences determined based on the descriptive text type under the specified criterion.

6. The method according to claim 1, characterized in that The image representation I is expressed as: i = VLM image (X) where, VLM image (·) represents the image encoder of the multimodal model that encodes the input image, and X represents the input image.

7. The method according to claim 1, characterized in that, The learned conditional representation R is expressed as: where I represents the image representation, T represents the text basis, and the superscript represents the matrix transpose operation.

8. A conditional representation learning system implemented based on the method according to any one of claims 1 to 7, characterized in that It includes: A text basis generation module: used to generate a standardized text basis for a given input image under a specified criterion; An image representation generation module: used to perform multimodal encoding on the input image through a multimodal model to generate an image representation; A basis projection module: used to project the image representation into the conditional feature space under the specified criterion spanned by the text basis to learn the conditional representation for downstream customized tasks.

9. The system according to claim 8, characterized in that The text basis generation module includes: A large language model: used to generate descriptive text related to a specified criterion for a given input image; A multimodal model, used to perform multimodal encoding on the descriptive text to generate a standardized text basis.

10. The system according to claim 9, wherein, When the large language model generates descriptive text related to a specified criterion, the prompt template is a common expression for generating descriptions of the specified criterion, and the position arrangement of the generated text elements and the uniqueness of the generated text elements are limited.