Zero-shot Text Classification Method and Device Based on Cross-modal Information Completion
The method addresses semantic alignment issues in zero-shot text classification by generating domain-specific keywords and using dual-path enhancement strategies, resulting in improved classification accuracy and reduced manual intervention.
Patent Information
- Application Number
- CN202510621202.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-05-14
AI Technical Summary
The existing zero-sample text classification methods have semantic bias and information missing problems in image selection, label embedding and multimodal fusion, resulting in limited performance of model cross-modal information utilization.
By designing a new image tag mapping mechanism, using a large-scale pre-trained language model to generate keyword phrases, combining cross-modal information-completion image construction, using a dual-path enhancement strategy for multimodal collaborative reasoning, and optimizing text semantic perception capabilities.
It significantly improves the zero-sample reasoning ability of the CLIP model, improves semantic accuracy and domain adaptability, and achieves higher text classification accuracy.
Smart Images

Figure CN120144764B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of zero-shot text classification, and particularly to a zero-shot text classification method and device based on cross-modal information completion. Background Art
[0002] Text classification is a fundamental task in the field of natural language processing (NLP), and plays an important role in many application scenarios such as news recommendation, sentiment analysis, and medical diagnosis. Its core goal is to map text to a predefined category space through semantic understanding. However, due to the difficulty in obtaining category annotation data, especially in specific professional fields where domain experts are usually required to participate in the annotation, the cost of manual annotation is high, making it difficult to effectively construct traditional supervised learning methods in the case of scarce data.
[0003] To alleviate the problem of insufficient category annotation data, few-shot and zero-shot text classification methods have emerged. Among them, few-shot text classification methods usually adopt methods such as transfer learning, meta-learning, or data augmentation to achieve classification of new categories with only a small amount of labeled data. Although few-shot text classification methods reduce the dependence on a large amount of labeled data, they still require a certain number of labeled samples, and their performance is difficult to guarantee in the case of rapid category changes or complete lack of annotation. With the increasing requirements for model generalization ability in practical applications and the exacerbation of the data scarcity problem, zero-shot text classification has gradually become a research hotspot. Different from few-shot text classification that relies on a small number of labeled samples, zero-shot text classification requires the model to achieve effective classification through cross-modal knowledge transfer or external knowledge injection in the case of completely lacking labeled data in the target domain, which poses higher requirements for the model's semantic understanding and knowledge transfer ability. The mainstream methods are mainly based on the text-text semantic matching paradigm, that is, by converting the category label into descriptive text and calculating its semantic similarity with the input text for classification. Although this method can effectively alleviate the data scarcity problem, its performance is limited by the semantic integrity of the label description and fails to fully utilize the synergistic effect of cross-modal information.
[0004] To address the above limitations, some researchers have started to explore the enhancing effect of visual modality information on text classification tasks. The CLIPTEXT framework proposed by Qin et al. was the first to apply the cross-modal ability of the CLIP model to zero-shot text classification, using images to enhance category information and verifying the effectiveness of the text-image matching method in zero-shot text classification. However, this method only focuses on visual image information and ignores the semantic knowledge embedded in text labels. Therefore, Zhang et al. further proposed the LABCLIP method, verifying the effectiveness of embedding label semantic information in image visual information in improving the performance of zero-shot text classification. In addition, to further improve the model performance, Wang et al. proposed the CLIPMulti framework, which innovatively adopted the text-image & text matching paradigm, aiming to achieve deeper multi-modal fusion by combining the image features and text features of labels. At the same time, to avoid the time cost of manually selecting texts and reduce the uncertainty brought by manually writing texts, the automated label generation strategy they proposed effectively reduces the dependence on manually designed label templates.
[0005] Although these methods have verified the effectiveness of combining visual image information for zero-shot text classification tasks to a certain extent, there are still challenges in aspects such as image source, the way of embedding label semantics in images, and multi-modal fusion:
[0006] (1) When selecting images, the above methods mainly calculate the similarity between the image candidate set and the corresponding label name, and verify through the validation set to select the image with the best performance as the mapped image of the label. However, label names are usually short and abstract, making it difficult to cover the diverse semantics in the actual data. When relying solely on label names for image selection, there are text semantic deviations between some mapped images and the real texts, resulting in the model being unable to clearly distinguish certain categories during the inference process, thus leading to uncertain or incorrect classification results. For example, when mapping the "Business" label to general business scene pictures (such as offices, stock charts), the model may not be able to distinguish the semantics of other subclasses;
[0007] (2) When embedding label semantics in images, Zhang et al. combined semantic knowledge by embedding fixed label texts at fixed positions in the images. Although it improved the classification performance to a certain extent, it relied on manual design of label positions and contents, and the short label texts could not provide sufficient semantic knowledge to achieve the effect of reducing semantic deviations. In addition, different embedding label positions would affect the classification results of the model. For the entire dataset, the best effect was not achieved only at the center of the image;
[0008] (3) In multi-modal fusion methods (such as matrix weighting), deep semantic information is easily lost due to the normalization operation. Additionally, there is still a gap between the automatically generated label descriptions and manually designed ones in terms of semantic accuracy and domain adaptability, making it difficult to fully capture the core semantics of the labels. Summary of the Invention
[0009] Based on this, it is necessary to provide a zero-shot text classification method and device based on cross-modal information completion for the technical problem of insufficient semantic integrity in cross-modal feature alignment.
[0010] A zero-shot text classification method based on cross-modal information completion, the method comprising:
[0011] Obtain a text sample set and an image candidate set corresponding to the labels.
[0012] Calculate the cosine similarity between the image corresponding to the label and the corresponding text sample and take the mean value, and determine the mapping image candidate set for each label according to the obtained average cosine similarity.
[0013] According to the performance of the mapping image candidate set of each label on the validation set, select the image with the best performance as the best mapping image for the label.
[0014] Input the training set samples corresponding to each label into a large-scale pre-trained language model to generate a preset number of candidate keyword phrases, and form a keyword candidate set.
[0015] Perform multi-modal matching and selection according to the keyword candidate set, label mapping image, text sample, and label set to obtain the best keywords adapted to the domain.
[0016] Perform image annotation with enhanced text semantics according to the best mapping image of each label and the best keywords adapted to the domain to obtain an image with cross-modal information completion.
[0017] Process the test sample set and the image with cross-modal information completion using a dual-path enhancement strategy and then perform zero-shot inference through the CLIP model to obtain the text classification result.
[0018] A zero-shot text classification device based on cross-modal information completion, the device comprising:
[0019] A text and label image acquisition module, configured to obtain a text sample set and an image candidate set corresponding to the labels.
[0020] A mapping image candidate set construction module for labels, configured to calculate the cosine similarity between the image corresponding to the label and the corresponding text sample and take the mean value, and determine the mapping image candidate set for each label according to the obtained average cosine similarity.
[0021] The image generation module for cross-modal information completion is used to select the best-performing image as the best mapped image for each label according to the performance of the candidate set of mapped images for each label on the validation set; input the training set samples corresponding to each label into a large-scale pre-trained language model to generate a preset number of candidate keyword phrases to form a keyword candidate set; perform multi-modal matching and selection based on the keyword candidate set, label mapped images, text samples, and label set to obtain the best keywords adapted to the domain; perform image annotation with enhanced text semantics based on the best mapped image for each label and the best keywords adapted to the domain to obtain images with cross-modal information completion.
[0022] The text classification module is used to perform zero-shot inference through the CLIP model after processing the test sample set and the images with cross-modal information completion using a dual-path enhancement strategy to obtain text classification results.
[0023] The above zero-shot text classification method and device based on cross-modal information completion. The method realizes the construction of images with cross-modal information completion by designing a new image label mapping mechanism and generating based on context text. Finally, a multi-modal collaborative inference framework is designed in the inference stage, and the text input is optimized through prompt engineering to enhance the text semantic perception ability, effectively improving the zero-shot inference ability of the CLIP model. Aiming at the semantic deviation problem in automatic label generation, based on context text generation, images with cross-modal information completion are constructed to optimize the integrity of text semantic representation. At the same time, prompt design is introduced in the inference stage to enhance the text semantic perception ability. Further, a two-stage keyword automatic selection mechanism is proposed. First, a keyword candidate set is generated using a large model, and then the best keyword phrases are selected through multi-modal matching and selection as the text information with cross-modal information completion, effectively improving semantic accuracy and domain adaptability. Description of the Drawings
[0024] Figure 1 It is a schematic flowchart of the zero-shot text classification method based on cross-modal information completion in one embodiment;
[0025] Figure 2 It is an overall architecture diagram of the zero-shot text classification method based on cross-modal information completion in one embodiment;
[0026] Figure 3 It is a schematic structural diagram of the two-stage keyword automatic selection mechanism (Selection-CLIPCMIC) in another embodiment;
[0027] Figure 4 It is a schematic diagram of the experimental verification results of the effectiveness of image selection in another embodiment;
[0028] Figure 5Schematic diagram of the comparison of the visualization results of the mapped images obtained by two different methods, namely Text-result and Label-result, on the dataset Trec in another embodiment;
[0029] Figure 6 Schematic diagram of PCA visualization analysis in another embodiment. Detailed implementation manners
[0030] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0031] The zero-shot text classification method based on cross-modal information completion proposed in the present application (CLIP-based Cross-modal Information Completion, abbreviated as: CLIPCMIC) improves the deficiencies of existing research through a multi-modal text semantic enhancement mechanism. Specifically, it includes the following three stages: (1) Semantic alignment label mapping: Based on the cross-modal representation consistency hypothesis, the present application designs a new mapping method to map the text label space to the image visual space by combining the training set text, and constructs an interpretable visual semantic representation for each label; (2) Image construction for cross-modal information completion based on context text generation. Different from the previous method (LABCLIP), the present application uses a large-scale pre-trained language model and combines the training set text to automatically generate domain-related keyword phrases, and then obtains the optimal keywords for each label through an adaptive screening mechanism (combining cosine similarity calculation). Finally, through automatic semantic annotation technology, the screened keywords are deeply fused with the image to construct a cross-modal representation with complete information, thereby enhancing the CLIP model's ability to capture fine-grained semantic knowledge; (3) Multi-modal collaborative reasoning: A dual-path enhancement strategy is designed in the reasoning stage. On the one hand, a prompt (here, a prompt sentence is used) is introduced at the text input end, and a domain-adapted prompt template is constructed through context optimization; on the other hand, the image with completed text semantic information is used as the image input end, and then zero-shot reasoning is performed through the CLIP model.
[0032] In one embodiment, as Figure 1 shown, a zero-shot text classification method based on cross-modal information completion is provided, and the method includes the following steps:
[0033] Step 100: Obtain a text sample set and an image candidate set corresponding to the labels.
[0034] Specifically, in order to more effectively use the CLIP model for zero-shot reasoning, this method converts the original "text-label" pair into a "text-image" pair.
[0035] Given training set (where represents the number of samples, ), label set (where represents the number of labels). By calculating the cosine similarity between each text and its corresponding label , for each label , the top ten samples with the highest similarity are taken as the text sample set, that is , where represents the first sample with the highest similarity to the label . For each label , on the Google search engine, 20 images are manually selected according to the label name to form the image set of this label. Then the image candidate set corresponding to the labels of the entire data set is , where represents the first image manually selected according to the label name.
[0036] The overall architecture of the zero-shot text classification method based on cross-modal information completion is as shown in Figure 2 .
[0037] Step 102: Calculate the cosine similarity between the image corresponding to the label and the corresponding text sample and take the mean. According to the obtained average cosine similarity, determine the mapping image candidate set for each label.
[0038] Specifically, use the text encoder and image encoder of the CLIP model to perform text encoding and image encoding and normalization on the text samples and the images in the image candidate set corresponding to the labels respectively. Then calculate the cosine similarity between each candidate image and all text samples and take the mean to obtain the cosine similarity and take the mean; for each label , a group of the top 5 images with the highest average cosine similarity score can be found as the mapping image candidate set for each label.
[0039] Step 104: According to the performance of the mapping image candidate set of each label on the validation set, select the image with the best performance as the best mapping image of this label.
[0040] Specifically, as shown in Figure 2 , based on the training set , label set (with a total of different class labels), randomly extract samples from the sample subset corresponding to each label as the validation set. The sample subset is , validation set ( ,in yes The sample index randomly selected from When , all samples are selected as the validation set. For each label, the performance of the top 5 images with the highest scores on the validation set is recorded in turn, and the image with the best result is selected as the best mapping image for the label. Therefore, with the help of semantically aligned label mapping and combined with the performance on the validation set, the “text-label” pair Can be mapped as a "text-image" pair ,in represents the optimal mapping image.
[0041] To address the semantic deviation problem caused by short and abstract label texts in the label mapping process, a new image label mapping method is proposed. Secondly, image construction with cross-modal information completion is achieved based on contextual text generation. Finally, a multimodal collaborative reasoning framework is designed in the reasoning stage. The text input is optimized through prompt engineering, which effectively improves the zero-shot reasoning capability of the CLIP model.
[0042] Step 106: Input the training set samples corresponding to each label into a large-scale pre-trained language model to generate a preset number of candidate keyword phrases to form a keyword candidate set.
[0043] Specifically, based on the context text generation, the training set samples corresponding to each label are sent as input data to the large-scale pre-trained language model to enhance its generation capability. For each label, ten candidate keyword phrases are generated by the model to form a keyword candidate set.
[0044] Keyword text generation: The training set samples corresponding to each label are sent to the big model as input data to enhance its generation ability. For each label, the model generates ten candidate keyword phrases to form a candidate set through user instructions.
[0045] Step 108: Perform multimodal matching and selection based on the keyword candidate set, label mapping image, text sample, and label set to obtain the best keyword adapted to the domain.
[0046] Specifically, multimodal matching and selection: by comprehensively evaluating the semantic similarity between keywords and label mapping images, label texts, and typical text samples, a weighted sum strategy is used to select the keyword with the highest comprehensive score as the text semantic information that needs to be completed for the image generated by the cross-modal information completion in step 110.
[0047] Regarding the semantic deviation problem in automated tag generation, in order to reduce the gap between the tag descriptions generated automatically and those designed manually in terms of semantic accuracy and domain adaptability, this application proposes a two-stage keyword automatic selection mechanism (Selection-CLIPCMIC), which optimizes the automatically generated tag descriptions into semantic representations of generated keyword phrases to improve the semantic accuracy of information description. The two-stage keyword automatic selection mechanism (Selection-CLIPCMIC) specifically includes: first, using a large model to generate a candidate set of keywords, and second, through multimodal matching and selection, selecting the best keyword phrase as the text information for cross-modal information complementation, effectively improving semantic accuracy and domain adaptability. Experimental observations show that annotating keywords at different positions on the image will significantly affect the inference performance of the CLIP model. Therefore, a method for automatically selecting the annotation position is proposed to improve the robustness of cross-modal alignment.
[0048] The two-stage keyword automatic selection mechanism (Selection-CLIPCMIC) provides more sufficient semantic knowledge and effectively improves semantic accuracy and domain adaptability by using a pre-trained language model to generate domain-related candidate keywords to replace short tag texts.
[0049] Step 110: Perform image annotation with enhanced text semantics based on the best-mapped image of each tag and the best keyword adapted to the domain, to obtain an image with cross-modal information complementation.
[0050] Specifically, this method constructs an image with cross-modal information complementation by designing a new image tag mapping mechanism and generating based on context text, to optimize the integrity of the text semantic representation, and at the same time introduces prompt design in the inference stage to enhance the text semantic perception ability.
[0051] The automatic semantic annotation technology deeply fuses the filtered keywords with the image to construct an image with enhanced text semantics. The automatic semantic annotation technology not only alleviates the impact of the difference in the embedded tag position on the model performance, but also effectively reduces the dependence on manual design.
[0052] Regarding the information missing problem in cross-modal feature alignment, a new image tag mapping method is proposed, and combined with context text generation to achieve the construction of an image with cross-modal information complementation, significantly improving the zero-shot inference ability of the CLIP model. Further, a multi-modal collaborative inference framework is constructed, and the semantic representation of the text input is optimized through prompts (using prompt sentences), achieving a performance improvement of 1.8% in average (AVG) accuracy.
[0053] Step 112: Process the test sample set and the image with cross-modal information complementation using a dual-path enhancement strategy and then perform zero-shot inference through the CLIP model to obtain the text classification result.
[0054] Specifically, as Figure 2 shown, to effectively complement the deep text semantic information, this method designs a dual-path enhancement strategy in the inference stage: First, introduce prompts in the text processing path, and construct a domain-adapted prompt template through context optimization; Second, use the image data with complemented information as the input in the visual processing path. Specifically, taking the original text input "How far is it from Denver to Aspen?" in the figure as an example, to enhance the text semantic feature extraction ability of the CLIP model and optimize the knowledge transfer effect, this application adds a structured prompt prefix before the original input to generate guiding text "This text is about [subtitle]: How far is it from Denver to Aspen?". Combining with the cross-modal information-complemented image generated in step 110, the enhanced data of the two paths are then jointly input into the CLIP model for zero-shot inference. In addition, the prompt sentences corresponding to the six datasets are shown in Table 1.
[0055] Step 1: Obtain the "text-image" pair ; Step 2: Define it as ; Step 3: Introduce a text prompt, define it as (where represents the text after introducing the prompt, represents the image after cross-modal information complementation). Finally, input the text and image together into the CLIP model for zero-shot prediction, which is defined as follows:
[0056] ;
[0057] Among them, and respectively represent the text encoder and the visual encoder of the CLIP model. In the single-label text classification task, select the label with the highest probability as the final prediction result, and in the multi-label text classification task, select the labels greater than the threshold as the final prediction result.
[0058] Table 1 Prompt sentence templates for different datasets
[0059]
[0060] To further complement the deep text semantic information, a multi-modal collaborative inference framework is proposed in the inference stage, a dual-path enhancement strategy is designed, and the text input is optimized through prompt engineering, significantly improving the zero-shot inference ability of the CLIP model.
[0061] In the above zero-shot text classification method based on cross-modal information completion, the method realizes the construction of images for cross-modal information completion by designing a new image label mapping mechanism and generating based on context text. Finally, in the inference stage, a multi-modal collaborative inference framework is designed, and the text input is optimized through prompt engineering to enhance the text semantic perception ability, effectively improving the zero-shot inference ability of the CLIP model. Aiming at the semantic deviation problem of automatic label generation, images for cross-modal information completion are constructed based on context text generation to optimize the integrity of text semantic representation. At the same time, prompt design is introduced in the inference stage to enhance the text semantic perception ability. Further, a two-stage keyword automatic selection mechanism is proposed. First, a set of keyword candidates is generated using a large model, and then, through multi-modal matching and selection, the best keyword phrase is selected as the text information for cross-modal information completion, effectively improving semantic accuracy and domain adaptability.
[0062] In one embodiment, step 102 includes: using the image encoder and text encoder of the CLIP model to perform text encoding and image encoding on the image corresponding to the label and the corresponding text sample respectively and normalizing them to obtain normalized text encoding and normalized image encoding; calculating the similarity between each candidate image and all text samples using cosine similarity according to the normalized text encoding and normalized image encoding and taking the average value to obtain the average cosine similarity of each candidate image; for each label, selecting the top 5 candidate images from the corresponding image candidate set as the mapping image candidate set of the label according to the average cosine similarity.
[0063] Specifically, CLIP (Contrastive Language-Image Pretraining) is a cross-modal pre-training framework based on contrastive learning. Its core mechanism is to map heterogeneous modal data to a shared semantic space through large-scale text-image pair joint training. Different from the traditional pre-training paradigm that relies on task-specific labeled data, CLIP uses 400 million pairs of large-scale text-image pairs and, through an inter-modal contrastive learning strategy, makes relevant text and images as close as possible in this space, while keeping irrelevant text and images at a distance. The remarkable feature of this framework is reflected in its zero-shot inference ability, that is, it can directly classify by calculating the similarity between text and images without specific task training data. This feature enables it to demonstrate excellent zero-shot transfer ability in fields such as image classification, image generation, downstream visual tasks, and cross-modal text classification.
[0064] The CLIP model adopts a dual-encoder architecture, which consists of a text encoder and an image encoder respectively. The text encoder is mainly implemented based on the Transformer architecture, and pre-trained language models such as BERT and GPT are commonly used text encoders. The image encoder usually uses convolutional neural networks (CNNs), such as classic network structures like ResNet. These networks extract local features from the image through multiple convolutional, pooling operations, and fully connected layers, and gradually combine these features into global features. When the text encoder and the image encoder respectively generate the text embedding vector and the image embedding vector , the model measures the semantic similarity between the text and the image by calculating the normalized dot product. Its calculation formula can be expressed as:
[0065] ;
[0066] where and represent the embedding vectors of the text and the image respectively, represents the L2 norm of the vector, and the superscript T represents the transpose.
[0067] Then, use the text encoder and the image encoder of the CLIP model to perform text encoding and image encoding on the text and the image respectively and normalize them; for each image perform feature encoding:
[0068] ;
[0069] where represents the L2 norm, represents the feature dimension.
[0070] For each label , the feature encoding of its text sample set is:
[0071] ;
[0072] where , , represents the feature dimension, and the superscript T represents the transpose.
[0073] Using the normalized embedding vectors, calculate the similarity between each candidate image and all text samples using cosine similarity and take the average. Specifically, for each candidate image , define its average similarity score as:
[0074] ;
[0075] where Represents the features of the th image, and represents the features of the th group of text samples. The th image and the l th group of text samples have an average cosine similarity of:
[0076] ;
[0077] where represents the average cosine similarity between the th image and the th group of text samples.
[0078] Finally, the optimal image selection is performed. For the label , the Top-5 images are selected according to the average cosine similarity:
[0079] ;
[0080] After the above process, for each label , a group of the top 5 images with the highest average similarity score can be found. In step 104 of this method, each image in the validation set is verified one by one on the validation set, and the best image is selected as the mapped image for the label.
[0081] In one embodiment, step 106 includes: inputting the training set samples corresponding to each label into the GPT-4o large-scale pre-trained language model, and generating a group of relevant keyword phrases according to the user's prompt instructions to form a keyword candidate set.
[0082] Specifically, in order to provide sufficient text semantic knowledge and reduce the gap between the automatically generated label descriptions and the manually designed ones in terms of semantic accuracy and domain adaptability, a two-stage keyword automatic selection mechanism (Selection-CLIPCMIC) is proposed to optimize the automatically generated label descriptions into domain-adapted keyword phrase representations, that is, replacing the label semantic information embedded in the image with keyword semantic information.
[0083] As Figure 3 shown, in the first stage of the two-stage keyword automatic selection mechanism: the training set samples corresponding to each label are sent as input data to the large-scale pre-trained language model to enhance its generation ability. In this method, GPT-4o is used to generate the keyword phrases of the labels, and the prompt format used is: "The above document is the dataset <dataset>The label in <label category name>Training set samples, please summarize all the sample content in the above document with ten keywords", among which <dataset>Indicates the name of the dataset, <label category name>Represents the label names in the dataset. Eventually, each label will have a set of associated candidate keyword phrases.
[0084] In one embodiment, step 108 includes: extracting features from the label mapping image, label, text sample, and keywords in the keyword candidate set to obtain image feature encoding, label feature encoding, text sample feature fusion, and keyword feature encoding; calculating the cosine similarity between the keywords and the label mapping image, label, and text sample respectively to obtain keyword-image semantic similarity, keyword-label semantic similarity, and keyword-text sample semantic similarity; performing weighted summation on the keyword-image semantic similarity, keyword-label semantic similarity, and keyword-text sample semantic similarity to obtain a comprehensive weighted scoring function; and making an optimal keyword decision based on the comprehensive weighted scoring function to obtain the best keywords adapted to the domain.
[0085] Specifically, in the first stage of the two-stage keyword automatic selection mechanism, multi-modal matching and selection are achieved: by comprehensively evaluating the semantic similarity between the keywords and the label mapping image, label text, and typical text samples, a weighted summation strategy is used to select the keyword with the highest comprehensive score as the text semantic information to be complemented in the second stage. The relevant formula definitions are as follows:
[0086] Category label set , which is used to represent the labels of all categories in the dataset;
[0087] Image set , which is used to represent the best mapping image corresponding to the category label ;
[0088] Text sample set , where , which is used to represent the text sample set corresponding to each category label, where ;
[0089] Candidate keyword phrase set , where , which is used to represent the keyword phrases corresponding to each category label, where ;
[0090] Optimal keyword set , which is used to represent an optimal keyword corresponding to each category label.
[0091] (1) Multi-modal feature extraction
[0092] Image feature encoding:
[0093] ;
[0094] Among them, is an image feature, is the result of the image encoder encoding the image corresponding to each label , and \(\mathbb{R}\) represents the set of real numbers. It represents the set of real numbers.
[0095] Label feature encoding:
[0096] ;
[0097] Among them, is the label feature, is the result of the text encoder encoding each label .
[0098] Text sample feature fusion:
[0099] ;
[0100] Among them, is the sample feature, is the result of the text encoder encoding the text sample corresponding to each label , and \(\text{mean}\) is to take the mean value, where ; here .
[0101] Keyword feature encoding:
[0102] ;
[0103] Among them, is the keyword feature, is the result of the text encoder encoding each keyword .
[0104] All of the above feature vectors are row vectors.
[0105] (2) Cosine similarity calculation
[0106] Calculate the cosine similarity between the keyword and the image to measure the semantic consistency between the keyword and the associated image of the category , and the output is a scalar.
[0107] ;
[0108] Among them, is the cosine similarity between the keyword and the image.
[0109] Calculate the cosine similarity between the keyword and the label to measure the text semantic consistency between the keyword and the category , and the output is a scalar.
[0110] ;
[0111] Among them, is the cosine similarity between the keyword and the tag.
[0112] Calculate the cosine similarity between the keyword and the text sample to measure the keyword k and the l semantic consistency of the text sample of the category.
[0113] ;
[0114] Among them, is the cosine similarity between the keyword and the text sample.
[0115] (3) Optimal keyword phrase selection
[0116] The comprehensive weighted scoring function is defined as follows:
[0117] ;
[0118] Among them, is the comprehensive weighted scoring function, are three weights, and the specific values are determined according to the performance of the record model on the validation set. The values of the 6 datasets in the experiment are shown in Table 2.
[0119] Optimal keyword decision:
[0120] ;
[0121] Among them, is the optimal keyword decision.
[0122] Through the above process, for each tag in each dataset, an optimal keyword phrase will be selected as the text semantic information for image semantic completion. Through this method, the automatically generated tag description is optimized into a domain-adapted keyword phrase representation, effectively reducing the gap between the automatically generated tag description and the manually designed one in terms of semantic accuracy and domain adaptability.
[0123] Table 2 Settings of weights and font sizes
[0124]
[0125] To reduce the gap between the automatically generated tag descriptions and the manually designed ones in terms of semantic accuracy and domain adaptability, this application designs a two-stage keyword automatic selection mechanism (Selection-CLIPCMIC) to optimize the automatically generated tag descriptions into domain-adapted keyword phrase representations. The structure of the two-stage keyword automatic selection mechanism (Selection-CLIPCMIC) is as shown in Figure 3 and is divided into two steps:
[0126] (1) Keyword text generation: The training set samples corresponding to each tag are sent as input data to the large model to enhance its generation ability. For each tag, the model generates ten candidate keyword phrases to form a candidate set through user instructions.
[0127] (2) Multimodal matching and selection: By comprehensively evaluating the semantic similarity between the keywords and the mapped images, tag texts, and typical text samples of the tags, a weighted summation strategy is adopted to select the keyword with the highest comprehensive score as the text semantic information to be complemented in step 104.
[0128] In addition, experimental observations show that annotating keywords at different positions on the image will significantly affect the inference performance of the CLIP model. Therefore, a method for automatically selecting the annotation position is proposed to improve the cross-modal alignment robustness.
[0129] In one embodiment, step 110 includes: determining the best position mapping of the best keyword adapted to the domain on the corresponding best mapped image according to the adaptive position annotation selection method; quantitatively evaluating and comparatively analyzing the performance of different font sizes on the validation set to determine the optimal font parameter configuration of the best keyword adapted to the domain; and completing the image annotation according to the best position mapping and the optimal font parameter configuration of the best keyword adapted to the domain to obtain an image with cross-modal information complemented.
[0130] In one embodiment, the specific steps of the adaptive position annotation selection method include: using the CLIP model to encode the text and image features, and performing post-processing of normalization to obtain the text feature matrix and the image feature vector; calculating the cosine similarity between the image and all texts according to the image feature vector and the text feature matrix, and taking its average value to obtain the final evaluation index; and selecting the position with the highest semantic consistency between the annotated image and the text sample set as the best position mapping of the text on the image according to the final evaluation index and the annotation effects of different preset positions.
[0131] Specifically, experimental observations find that the annotation position of keyword phrases has a significant impact on the inference performance of the CLIP model. Therefore, a method for automatically selecting the annotation position is proposed to optimize the cross-modal alignment.
[0132] In one of the embodiments, the expression of the optimal position mapping is as follows:
[0133] ;
[0134] wherein, represents the optimal position mapping from the category label set to the preset position set, represents the process of annotating the keyword at the position p on the image , represents the text sample set, represents the final evaluation metric.
[0135] In one of the embodiments, the performance of different font sizes is quantitatively evaluated and comparatively analyzed on the validation set to determine the optimal font parameter configuration of the optimal keyword adapted to the domain, including: designing the objective function and constraint conditions for dynamic font size adjustment; the expressions of the objective function and constraint conditions for dynamic font size adjustment are as follows:
[0136] ;
[0137] ;
[0138] wherein, is the objective function, represents the font size, , t represents the text content to be annotated (i.e., keyword phrase or text sample), which is the key variable in the constraint condition, and represent the text width and height, W , H represents the image size.
[0139] When dynamic adjustment is enabled, the most suitable font size is optimized based on the constraint conditions as the optimal font parameter configuration of the optimal keyword adapted to the domain; if the performance of the model is quantitatively evaluated through the cross-validation set, and if the inference performance of the model is not significantly improved after dynamic adjustment, then the font size is manually adjusted to determine the optimal keyword adapted to the domain.
[0140] Specifically, image annotation with enhanced text semantics is finally performed to achieve the construction of images for cross-modal information completion. Through experiments, it is found that the method of directly embedding keyword phrases into the mapped image space has limited effect on improving the zero-shot inference ability of the CLIP model, because for the entire dataset, the best effect can be achieved not only at the center of the image. Further analysis reveals that visual parameters such as the annotation position and font size of keyword phrases have a significant impact on model performance. To solve the above problems, an adaptive position annotation selection method is proposed: First, the spatial layout of semantic information is optimized by designing an adaptive annotation position selection algorithm; Second, regarding the impact of font size on model inference, systematic experiments are carried out on the validation set to quantitatively evaluate and comparatively analyze the performance of different font sizes, and finally the optimal font parameter configuration is determined. (Note: All the above methods are implemented through the CLIP model, and the text sample set mentioned in the algorithm is the same as that in Step 100). The specific implementation process is as follows:
[0141] The relevant formula definitions are as follows:
[0142] Category label set , which is used to represent the labels of all categories in the dataset;
[0143] Preset position set , which is used to represent the preset positions of text annotations;
[0144] Image set , which is used to represent the best mapped image corresponding to each category label;
[0145] Keyword phrase set , which is used to represent 1 keyword phrase corresponding to each category label;
[0146] Text sample set , which is used to represent the text sample set corresponding to each category label and is used to calculate the similarity between images and texts;
[0147] Best position mapping , which is used to represent the position that selects the highest semantic consistency between the annotated image and the text sample;
[0148] Optimized image set , which contains the annotated images optimized for all category labels, enhances semantic expression and supports downstream tasks.
[0149] (1) Image brightness calculation
[0150] It is used to adaptively select the text color to ensure that the text is clearly visible on the image.
[0151] ;
[0152] wherein is the input image, is the RGB-to-luminance conversion function, W , H represents the image size, is the average luminance scalar, and (x, y) represents the coordinates of the image pixels.
[0153] (2) Font size constraint optimization
[0154] Design the font size constraint optimization strategy described by the objective function and constraint expression for dynamic font size adjustment, and adjust the font size. Dynamic font size adjustment aims to ensure that text information is clearly distinguishable within the image area and does not exceed the boundaries. When dynamic adjustment is enabled, the system optimizes based on the constraints to select the most appropriate font size. To verify the effectiveness of the method, the model performance is quantitatively evaluated through a cross-validation set. If the inference performance of the model has not been significantly improved after dynamic adjustment, manual adjustment is selected (Note: When dynamically adjusting, the fonts of the keyword phrases annotated for all labeled images in each dataset will change, and manual adjustment is a unified change, and the specific values are shown in Table 2).
[0155] (3) Text image feature encoding
[0156] Use the CLIP model to encode text and image features and map them to the same feature space for cosine similarity calculation. The image feature and text feature matrices are:
[0157] ;
[0158] ;
[0159] wherein, represents the image feature (row vector), represents the text feature matrix (each row is a text feature, and the text comes from ), represents the number of text samples, where , represents the feature dimension, represents the L2 norm.
[0160] (4) Multimodal cosine similarity calculation
[0161] Measures the semantic consistency between the image and the text and is used to optimize the text position.
[0162] ;
[0163] ;
[0164] Among them, represents the cosine similarity matrix of the image and all texts, represents taking the mean of the cosine similarity scores of m texts as the final evaluation metric, represents the label, which is used to represent the set of text samples that are clearly related to .
[0165] (5) Optimal mapping position
[0166] By evaluating the annotation effects of different preset positions, select the position that maximizes the semantic consistency between the annotated image and the set of text samples . The optimal position mapping is as shown in the expression of the above optimal position mapping.
[0167] (6) Optimized image set
[0168] Through keyword phrase annotation, the optimized annotated images containing all category labels not only enhance the semantic expression of the images but also support downstream tasks (such as classification, etc.).
[0169] ;
[0170] Among them, represents the optimized image set, represents the final optimized image of the category.
[0171] This method realizes the construction of an image for cross-modal information completion, generating an image representation with enhanced text semantic features. This image representation will be used as the visual input of the multi-modal collaborative reasoning module in the third stage, jointly modeled with the corresponding text features, and finally realizing the cross-modal reasoning task under zero-shot conditions.
[0172] It should be understood that although Figure 1 the steps in the flowchart of Figure 1 are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover,
[0173] In a verification example, to comprehensively evaluate the performance of CLIPCMIC in zero-shot text classification tasks, this study conducts experimental verification on six publicly available benchmark datasets, covering various task types such as multi-label classification, intent recognition, and topic classification. The characteristics of each dataset are as follows:
[0174] Yahoo! Answers: A multi-label text classification dataset containing 10 significantly semantically different topic categories (such as science, health, sports, etc.) in the Yahoo! Q&A community. This dataset features the language diversity of user-generated content and high distinctiveness between categories, and is suitable for verifying the zero-shot classification ability of the model in open-domain scenarios.
[0175] TREC (Text REtrieval Conference): A fine-grained question classification benchmark containing 6 types of question intents (such as fact-based, person-based, etc.). As an authoritative evaluation set released by the TREC conference, its annotation standardization and semantic complexity can be effectively used to evaluate the model's ability to understand question intents.
[0176] AG’s News: A news classification dataset covering 4 major news categories of business, politics, sports, and technology. Its large-scale and high-quality professional news corpus provides a reliable benchmark for evaluating the domain generalization ability of the model in formal texts.
[0177] Subj: A subjectivity classification dataset containing two types of texts, subjective (such as comments, opinions) and objective (such as factual statements). With a moderate data scale and concise and clear texts, it can effectively verify the model's discriminative ability for language expression attributes.
[0178] SNIPS: An intent detection dataset containing 7 types of user intent texts (such as querying the weather, playing music, booking a restaurant, etc.). The data is sourced from the text input of voice assistants and is suitable for testing the model's intent inference ability in colloquial expressions.
[0179] ATIS: An intent classification benchmark in the aviation field containing 15 professional intent categories such as flight queries and ticket bookings. As a classic evaluation set in the field of dialogue systems, its high density of professional terms and domain specificity provides an effective test platform for verifying the model's domain adaptation ability.
[0180] (1) Baseline methods
[0181] Compare the method CLIPCMIC of this application with the following benchmark models:
[0182] 1) RTE: A classification method based on text entailment, regarding the input text and label as an entailment problem, and training a bert-base-uncased model based on the RTE dataset for inference.
[0183] 2) FEVER: Similar to RTE, it uses a model based on bert-base-uncased and is pre-trained on the FEVER dataset to enhance text entailment ability.
[0184] 3) MNLI: Also based on the idea of text entailment, it uses the bert-base-uncased model and is pre-trained through the MNLI dataset to improve generalization ability.
[0185] 4) NSP: A method for zero-shot text classification that directly uses the next sentence prediction pre-training task of BERT.
[0186] 5) NSP (Reverse): Since NSP is not designed for directed semantic entailment prediction, Ma et al. explored a variant that reverses the order of all text pairs and calls it NSP (Reverse).
[0187] 6) GPT-2: A zero-shot text classification method based on a generative pre-trained model that completes the classification task by directly generating the corresponding label output.
[0188] 7) CLIP Text Encoder: A Transformer-based text encoder that maps text features to a multimodal embedding space to achieve text classification tasks.
[0189] 8) CLIPTEXT: By reformulating zero-shot text classification as a text-image matching problem, it uses the cross-modal alignment ability of the CLIP model for classification.
[0190] 9) LABCLIP: By embedding label semantic information in images, it significantly enhances the ability to express label semantics.
[0191] 10) CLIPMulti (MF): A multimodal enhanced CLIP framework based on the text-image & text matching paradigm and through matrix fusion.
[0192] (2) Experimental results
[0193] All datasets in the experiment were evaluated using the accuracy (acc). The experimental results show that (as shown in Table 3), except for the Yahoo! Answers dataset, the proposed method outperforms all baseline methods on the other five benchmark datasets. Among them, the ATIS dataset achieved the most significant performance improvement, with the classification accuracy increasing by 11.7% compared to the best baseline method; followed by the Subj dataset, with a 6.0% improvement. Regarding the issue of label distribution differences in the ATIS dataset, it was found through analysis that its test set only contains samples of 2 atis_day_name categories, and this category does not appear in the public training set. To maintain the rigor of the experiment, this study evaluated the model under 15 standard intent categories and re-obtained the classification accuracy of the ATIS dataset according to the experimental structure of Zhang et al. to ensure the fairness of the comparative experiment. It is worth noting that under the same experimental conditions, the proposed method achieved significant improvements of 4.5% and 3.7% respectively compared to the LABCLIP and CLIPMulti baseline models in terms of the average accuracy (AVG) metric. This result verifies that the proposed method significantly enhances the CLIP model's ability to capture text semantic features by effectively integrating cross-modal semantic information, thus achieving better generalization performance in the zero-shot text classification task.
[0194] Table 3 Experimental Results
[0195]
[0196] Among them, # indicates that the experimental results are from Zhang et al., & indicates that it is reproduced on the public code, and other results are from the original paper. AVG represents the average Acc score, and CLIPCMIC* indicates alignment with the CLIPMulti method, calculating the AVG of the same 5 datasets for comparison.
[0197] (3)Analysis of Experimental Results
[0198] To systematically verify the effectiveness of the proposed method, the following experiments were conducted for analysis: First, verify the effectiveness of the image selection strategy; second, examine the performance of the automatic annotation position selection mechanism; then, explore the contributions of each module through ablation experiments; finally, conduct multi-modal feature analysis on the Subj dataset through visualization methods.
[0199] 1) Verification of Image Selection Effectiveness
[0200] In the semantic label mapping stage, this study proposes a text sample-based method to select the best image by calculating the cosine similarity between the image and its corresponding text sample and taking the mean. To verify the effectiveness of this method, it is compared with the method proposed by Qin et al. that calculates the similarity between the image and the corresponding label name based on the label name in the experiment. Both methods are experimented on the same validation set, and no additional text information or hint mechanism is introduced during the experiment.
[0201] The experimental results are as Figure 4 shown, where Label-result represents the image obtained by the original method, where the original method refers to the method of calculating similarity based on the label name, Text-result represents the image obtained by this method, where this method refers to the method of calculating cosine similarity based on text samples proposed in this application, and AVG represents the corresponding average accuracy metric. It can be intuitively observed that the method proposed in this study is significantly better than the label name-based baseline method on all six datasets (with an average improvement rate of 15.8%). It verifies that text samples have a richer semantic representation ability than label names and can effectively enhance the accuracy of cross-modal mapping. This method not only does not rely on label names but also reduces the semantic deviation between the image and the real text, improving the model classification accuracy. It should be noted that on the Trec and ATIS datasets, there is a performance difference between the validation set and the test set in the experimental results of Label-result, resulting in the test results being significantly lower than Text-result. After analysis, it is found that this phenomenon is due to the overfitting problem of the label name-based baseline method during this experiment.
[0202] The visualization results are as Figure 5 shown. Taking the Trec dataset as an example, the cross-modal feature distribution is demonstrated. In this embodiment, the LDA method is used to perform dimensionality reduction and visualization analysis on the test set samples (circular markers) processed by CLIP, the images obtained by this method (star markers), and the images obtained by the original method (diamond markers). In addition, × represents the clustering center. The experimental results show that the images obtained by this method (star markers) are closer to the clustering center in the feature space than the images obtained by the original method (diamond markers), indicating that text samples can more accurately capture the core features of the category through their richer semantic representation ability. This phenomenon reveals that the text-guided feature space has higher intra-class aggregation and inter-class discrimination, thus strengthening the robustness of cross-modal semantic alignment and providing a theoretical basis for the improvement of classification performance. It should be noted that since the results of the two mapped images corresponding to the ENTY label are the same, they appear in a superimposed state in the visualization diagram.
[0203] 2) Verification of the effectiveness of automatically selecting the annotation position
[0204] It was found in the experiment that different positions of keyword phrase annotation have a significant impact on the model performance. Therefore, an automatic semantic annotation technique was proposed, which can automatically select the annotation position. To verify the effectiveness of this method, this study conducted experimental verification on three datasets. The specific positions of injecting keyword phrases in the image, including: center, top_left, top_right, bottom_left, bottom_right, top_center, bottom_center.
[0205] The experimental results are shown in Table 4. This method was compared with seven fixed annotation position strategies on three benchmark datasets. The results show that automatically selecting the annotation position is significantly better than the method with fixed annotation positions in terms of the Acc metric. It shows that the annotation position plays a key role in cross-modal information fusion, and at the same time proves that it is difficult for the fixed annotation position method to achieve optimal semantic injection. While automatically selecting the annotation position can effectively capture the context-sensitive annotation position, effectively improve the model inference ability, and at the same time reduce the dependence on the artificially designed label position.
[0206] Table 4 Experimental results for verifying the method of automatically selecting annotation positions
[0207]
[0208] 3) Ablation experiment
[0209] To evaluate the contributions of each module in the CLIPCMIC framework, relevant ablation experiments were designed. The experimental results are shown in Table 5. First, when the keyword phrase annotation module was removed, the performance of the CLIPCMIC framework decreased significantly, and the average accuracy (AVG) decreased by 3.4%. This indicates that the image construction module that realizes cross-modal information complementation based on context text generation helps to reduce the information loss problem in the cross-modal feature alignment process and, to a certain extent, complements the text semantic information, thus improving the zero-shot inference ability of the model. At the same time, this result also proves that the two-stage keyword automatic selection mechanism ((Selection-CLIPCMIC)) optimizes the automatically generated label descriptions into domain-adapted keyword phrase representations, significantly narrowing the gap in semantic accuracy and domain adaptability between the automatically generated label descriptions and the artificial design. Second, when the prompts were removed, the average accuracy (AVG) decreased by 1.8%, confirming that the domain-adapted prompts can effectively enhance the semantic understanding ability of the CLIP model. Finally, when both the keyword phrase annotation and the prompts were removed simultaneously, the average accuracy of CLIPCMIC on six datasets decreased by 4.5%, with the highest average performance decline among the three groups of experiments, indicating that in the CLIPCMIC method, neither module can be missing, and they have a synergistic enhancement effect in this framework. Additionally, it also shows that this method can effectively reduce the deep semantic information lost due to the normalization operation in the multi-modal fusion method.
[0210] Table 5 Ablation Experiments
[0211]
[0212] Among them, w / o means removing this module, prompts means prompts, and keyword-pa means keyword phrase annotation.
[0213] 4) Visualization Analysis
[0214] Taking the Subj benchmark dataset as the experimental object, the CLIP model was used to conduct a joint characterization analysis of multi-modal data features, specifically including the test set text, semantically aligned label mapping images, images after cross-modal information complementation, and the best keyword phrases. After extracting the feature vectors of each modality, principal component analysis (PCA) was used for dimensionality reduction processing and visualization analysis. The schematic diagram of the PCA visualization analysis is as Figure 6 shown. Figure 6 The text in the test set is represented by dots, the image after cross-modal information completion is represented by pentagrams, the label mapping image with semantic alignment is represented by diamonds, the best keyword phrases are represented by squares, and the cluster centers are represented by ×. Despite the existence of discrete data points, overall, the feature vectors of different modalities show obvious clustering characteristics in the embedding space, indicating that the CLIP model has strong multi-modal alignment capabilities. It is worth noting that the feature vectors of the images after cross-modal text semantic completion (the cluster centers marked with ×) are more spatially distributed towards the cluster centers compared to the feature vectors of the label mapping images without semantic completion. This shows that the cross-modal information completion strategy proposed in this paper can significantly improve the semantic discrimination ability of visual representations, that is, the images after text semantic information completion can capture the core semantic information of the target category more accurately, thus achieving a consistent representation of the image and text in the semantic space.
[0215] The experimental results show that this method has achieved performance significantly superior to existing baseline methods on six publicly available benchmark datasets. Compared with LABCLIP and CLIPMulti, the average accuracy has increased by 4.5% and 3.7% respectively. In addition, to systematically verify the effectiveness of the designed method in this study, extensive experimental verifications were carried out in this embodiment to prove the effectiveness of the method designed in this application.
[0216] In one embodiment, a zero-shot text classification device based on cross-modal information completion is provided, including: a text and label image acquisition module, a mapping image candidate set construction module for labels, an image generation module for cross-modal information completion, and a text classification module, where:
[0217] The text and label image acquisition module is used to acquire a text sample set and an image candidate set corresponding to the labels;
[0218] The mapping image candidate set construction module for labels is used to calculate the cosine similarity between the image corresponding to the label and the corresponding text sample and take the average value, and determine the mapping image candidate set for each label according to the obtained average cosine similarity;
[0219] The image generation module for cross-modal information completion is used to select the image with the best performance as the best mapping image for the label according to the performance of the mapping image candidate set for each label on the validation set; input the training set samples corresponding to each label into a large-scale pre-trained language model to generate a preset number of candidate keyword phrases to form a keyword candidate set; perform multi-modal matching and selection according to the keyword candidate set, label mapping image, text sample, and label set to obtain the best keywords adapted to the domain; perform image annotation with text semantic enhancement according to the best mapping image for each label and the best keywords adapted to the domain to obtain the image with cross-modal information completion;
[0220] A text classification module for performing zero-shot inference on a test sample set and images with cross-modal information completion through a CLIP model after processing using a dual-path enhancement strategy to obtain text classification results.
[0221] In one embodiment, the mapping image candidate set construction module for labels is further configured to perform text encoding and image encoding and normalization on the image corresponding to the label and the corresponding text sample respectively using the image encoder and text encoder of the CLIP model to obtain normalized text encoding and normalized image encoding; calculate the similarity between each candidate image and all text samples using cosine similarity based on the normalized text encoding and normalized image encoding, and take the average value to obtain the average cosine similarity of each candidate image; for each label, select the top 5 candidate images from the corresponding image candidate set as the mapping image candidate set for the label according to the average cosine similarity.
[0222] In one embodiment, the cross-modal information completion image generation module is further configured to input the training set samples corresponding to each label into the GPT-4o large-scale pre-trained language model, and generate a set of relevant keyword phrases according to the user's prompt instructions to form a keyword candidate set.
[0223] In one embodiment, the cross-modal information completion image generation module is further configured to extract features from the keyword in the label mapping image, label, text sample, and keyword candidate set to obtain image feature encoding, label feature encoding, text sample feature fusion, and keyword feature encoding; calculate the cosine similarity between the keyword and the label mapping image, label, and text sample respectively to obtain keyword-image semantic similarity, keyword-label semantic similarity, and keyword-text sample semantic similarity; perform weighted summation on the keyword-image semantic similarity, keyword-label semantic similarity, and keyword-text sample semantic similarity to obtain a comprehensive weighted scoring function; adopt an optimal keyword decision according to the comprehensive weighted scoring function to obtain the best keyword adapted to the domain.
[0224] In one embodiment, the cross-modal information completion image generation module is further configured to determine the best position mapping of the best keyword adapted to the domain on the corresponding best mapping image according to the adaptive position annotation selection method; quantitatively evaluate and compare the performance of different font sizes on the validation set to determine the optimal font parameter configuration of the best keyword adapted to the domain; complete the image annotation according to the best position mapping and optimal font parameter configuration of the best keyword adapted to the domain to obtain the cross-modal information completion image.
[0225] In one embodiment, the specific steps of the adaptive position annotation selection method in the image generation module for cross-modal information completion include: using the CLIP model to encode text and image features, and performing post-processing of normalization to obtain a text feature matrix and an image feature vector; calculating the cosine similarity between the image and all texts based on the image feature vector and the text feature matrix, and taking the average value to obtain a final evaluation index; selecting the position with the highest semantic consistency between the annotated image and the text sample set as the best position mapping of the text on the image according to the final evaluation index and the annotation effects of different preset positions.
[0226] In one embodiment, the best position mapping is as shown in the expression of the best position mapping above.
[0227] In one embodiment, the performance of different font sizes is quantitatively evaluated and comparatively analyzed on the validation set to determine the optimal font parameter configuration of the best keywords adapted to the domain, including: designing the objective function and constraint conditions for dynamic font size adjustment as shown in the expressions of the above objective function and constraint conditions.
[0228] When dynamic adjustment is enabled, the most suitable font size is optimally selected based on the constraint conditions as the optimal font parameter configuration of the best keywords adapted to the domain; if the performance of the model is quantitatively evaluated through the cross-validation set, and if the inference performance of the model is still not significantly improved after dynamic adjustment, then the font size is manually adjusted to determine the best keywords adapted to the domain.
[0229] For the specific limitations of the zero-shot text classification device based on cross-modal information completion, reference can be made to the limitations of the zero-shot text classification method based on cross-modal information completion in the above text, which will not be elaborated here. Each module in the above zero-shot text classification device based on cross-modal information completion can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or be independent of it, or can be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0230] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0231] The embodiments described above merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.< / label> < / dataset> < / label> < / dataset>
Claims
1. A zero-shot text classification method based on cross-modal information completion, characterized in that The method includes: Obtain a text sample set and an image candidate set corresponding to labels; Calculate the cosine similarity between the image corresponding to a label and the corresponding text sample and take the average value. According to the obtained average cosine similarity, determine the mapping image candidate set for each label; According to the performance of the mapping image candidate set of each label on the validation set, select the image with the best performance as the best mapping image for the label; Input the training set samples corresponding to each label into a large-scale pre-trained language model to generate a preset number of candidate keyword phrases, constituting a keyword candidate set; Perform multimodal matching and selection according to the keyword candidate set, label mapping image, text sample, and label set to obtain the best keywords adapted to the domain; Perform image annotation with enhanced text semantics according to the best mapping image of each label and the best keywords adapted to the domain to obtain an image with cross-modal information complemented; Process the test sample set and the image with cross-modal information complemented using a dual-path enhancement strategy and then perform zero-shot inference through a CLIP model to obtain a text classification result.
2. The zero-shot text classification method based on cross-modal information completion according to claim 1, wherein Calculating the cosine similarity between the image corresponding to a label and the corresponding text sample and taking the average value. According to the obtained average cosine similarity, determining the mapping image candidate set for each label includes: Use the image encoder and text encoder of the CLIP model to perform image encoding and text encoding on the image corresponding to the label and the corresponding text sample respectively and normalize them to obtain normalized text encoding and normalized image encoding; Calculate the similarity between each candidate image and all text samples using cosine similarity according to the normalized text encoding and the normalized image encoding, and take the average value to obtain the average cosine similarity of each candidate image; For each label, select the top 5 candidate images from the corresponding image candidate set as the mapping image candidate set for the label according to the average cosine similarity.
3. The zero-sample text classification method based on cross-modal information completion according to claim 1, characterized in that Inputting the training set samples corresponding to each label into a large-scale pre-trained language model to generate a preset number of candidate keyword phrases, constituting a keyword candidate set, includes: Input the training set samples corresponding to each label into the GPT-4o large-scale pre-trained language model, and generate a set of relevant keyword phrases according to the user's prompt instructions to constitute a keyword candidate set.
4. The zero-sample text classification method based on cross-modal information completion according to claim 1, characterized in that Performing multimodal matching and selection according to the keyword candidate set, label mapping image, text sample, and label set to obtain the best keywords adapted to the domain includes: Extract features from the label mapping image, label, text sample, and keywords in the keyword candidate set to obtain image feature encoding, label feature encoding, text sample feature fusion, and keyword feature encoding; Calculate the cosine similarity between the keyword and the label mapping image, label, and text sample respectively to obtain keyword-image semantic similarity, keyword-label semantic similarity, and keyword-text sample semantic similarity; Perform weighted summation on the keyword-image semantic similarity, keyword-label semantic similarity, and keyword-text sample semantic similarity to obtain a comprehensive weighted scoring function; Optimal keyword decision is made according to the comprehensive weighted scoring function to obtain the best keywords adapted to the domain.
5. The zero-shot text classification method based on cross-modal information completion according to claim 1, characterized in that Based on the best mapped images of each label and the best keywords adapted to the domain, image annotation with enhanced text semantics is performed to obtain images with cross-modal information complementation, including: Determine the best position mapping of the best keywords adapted to the domain on the corresponding best mapped images according to the adaptive position annotation selection method; Quantitatively evaluate and comparatively analyze the performance of different font sizes on the validation set to determine the optimal font parameter configuration of the best keywords adapted to the domain; Complete the image annotation according to the best position mapping and the optimal font parameter configuration of the best keywords adapted to the domain to obtain images with cross-modal information complementation.
6. The zero-shot text classification method based on cross-modal information completion according to claim 5, wherein The specific steps of the adaptive position annotation selection method include: Use the CLIP model to encode text and image features, and perform post-normalization processing to obtain a text feature matrix and an image feature vector; Calculate the cosine similarity between the image and all texts according to the image feature vector and the text feature matrix, and take the average value to obtain the final evaluation index; According to the final evaluation index and the annotation effects of different preset positions, select the position with the highest semantic consistency between the annotated image and the text sample set as the best position mapping of the text on the image.
7. The zero-shot text classification method based on cross-modal information completion according to claim 5, characterized in that The best position mapping is: Among them, represents the best position mapping from the category label set to the preset position set, represents a label, represents on the image at the position p annotating keywords of the process, represents the text sample set, represents the final evaluation metric.
8. The zero-sample text classification method based on cross-modal information completion according to claim 5, characterized in that Quantitatively evaluate and comparatively analyze the performance of different font sizes on the validation set to determine the optimal font parameter configuration of the best keywords adapted to the domain, including: Design the objective function and constraint conditions for dynamic font size adjustment as: Among them, is the objective function, represents the font size, , and represent the text width and height, t represents the text content to be labeled, W , H represents the image size; When dynamic adjustment is enabled, optimize and select the most suitable font size as the optimal font parameter configuration of the best keywords adapted to the domain based on the constraint conditions; if the performance of the model is quantitatively evaluated through the cross-validation set and the inference performance of the model is not significantly improved after dynamic adjustment, then select to manually adjust the font size to determine the optimal font parameter configuration of the best keywords adapted to the domain.
9. A zero-sample text classification device based on cross-modal information completion, characterized in that, The device includes: A text and label image acquisition module for acquiring a text sample set and an image candidate set corresponding to the labels; A mapped image candidate set construction module for calculating the cosine similarity between the image corresponding to the label and the corresponding text sample and taking the mean value, and determining the mapped image candidate set of each label according to the obtained average cosine similarity; A cross-modal information complemented image generation module for selecting the image with the best performance as the best mapped image of the label according to the performance of the mapped image candidate set of each label on the validation set; inputting the training set samples corresponding to each label into a large-scale pre-trained language model to generate a preset number of candidate keyword phrases to form a keyword candidate set; performing multi-modal matching and selection according to the keyword candidate set, label mapped image, text sample, and label set to obtain the best keywords adapted to the domain; performing text-semantic-enhanced image annotation according to the best mapped image of each label and the best keywords adapted to the domain to obtain images with cross-modal information complementation; A text classification module, which is used to perform zero-shot inference on the test sample set and the images with cross-modal information completion processed by a dual-path enhancement strategy through a CLIP model to obtain text classification results.
Citation Information
Patent Citations
Zero sample text classification method and system based on CLIP and medium
CN116701637A
Multimodal few-shot learning with frozen language models
US20240282094A1