Zero sample text classification method and device based on cross-modal information completion

By selecting the best keywords for the best mapping image and the generation field adaptation based on cross-modal information completion methods, the challenges of image selection and label semantic embedding in the prior art are solved, and the inference ability and semantic representation integrity of the zero-sample text classification model are significantly improved.

CN120144764AActive Publication Date: 2025-06-13NAT UNIV OF DEFENSE TECH

Patent Information

Application Number
CN202510621202.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-06-13
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

The existing zero-sample text classification methods have challenges in image selection, label semantic embedding and multimodal fusion, which makes it difficult for the model to clearly distinguish categories during the inference process, resulting in uncertain or error in the classification results.

Method used

Using a method based on cross-modal information completion, the cosine similarity between the corresponding image of the label and the corresponding text sample is calculated, the best mapped image is selected, and a large-scale pre-trained language model is used to generate candidate keyword phrases, perform multimodal matching and selection, obtain the best keywords that are adapted to the domain, and then perform text semantic enhancement image annotation to construct a cross-modal information completion image.

Benefits of technology

It effectively improves the zero-sample reasoning capability of the CLIP model, optimizes the completeness of text semantic representation, improves semantic accuracy and domain adaptability, and significantly improves the model's performance in cross-modal feature alignment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144764A_ABST
    Figure CN120144764A_ABST
Patent Text Reader

Abstract

The invention relates to a zero sample text classification method and device based on cross-modal information completion. According to the method, cross-modal information complemented image construction is realized by designing a new image label mapping mechanism and based on context text generation, finally, a multi-modal collaborative reasoning framework is designed in a reasoning stage, text semantic perception capability is enhanced by prompting project optimization text input, and zero sample reasoning capability of a CLIP model is effectively improved. Aiming at the semantic deviation problem of automatic label generation, a cross-modal information complemented image is constructed based on context text generation so as to optimize the text semantic representation integrity, and meanwhile, a prompt design is introduced in a reasoning stage to enhance the text semantic perception capability; according to the two-stage keyword automatic selection mechanism, firstly, a keyword candidate set is generated through a large model, secondly, through multi-modal matching and selection, the optimal keyword phrases are selected to serve as text information complemented by cross-modal information, and semantic accuracy and field adaptability are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of zero-sample text classification, and in particular to a zero-sample text classification method and device based on cross-modal information completion. Background Art

[0002] Text classification is a fundamental task in the field of natural language processing (NLP), and plays an important role in many application scenarios such as news recommendation, sentiment analysis, and medical diagnosis. Its core goal is to map text to a predefined category space through semantic understanding. However, due to the difficulty in obtaining category annotation data, especially in specific professional fields, category annotation data usually requires the participation of domain experts, resulting in high manual annotation costs, making it difficult to effectively construct traditional supervised learning methods in the case of scarce data.

[0003] To alleviate the problem of insufficient category annotation data, few-shot and zero-shot text classification methods have emerged. Among them, few-shot text classification methods usually adopt methods such as transfer learning, meta-learning, or data augmentation to achieve classification of new categories with only a small amount of labeled data. Although few-shot text classification methods reduce the dependence on a large amount of labeled data, they still require a certain number of labeled samples, and their performance is difficult to guarantee in the case of rapid category changes or complete lack of labels. With the increasing requirements for model generalization ability in practical applications and the exacerbation of the data scarcity problem, zero-shot text classification has gradually become a research hotspot. Different from few-shot text classification that relies on a small number of labeled samples, zero-shot text classification requires the model to achieve effective classification through cross-modal knowledge transfer or external knowledge injection in the case of completely lacking labeled data in the target domain, which poses higher requirements for the model's semantic understanding and knowledge transfer ability. The mainstream methods are mainly based on the text-text semantic matching paradigm, that is, by converting the category label into descriptive text and calculating its semantic similarity with the input text for classification. Although this method can effectively alleviate the data scarcity problem, its performance is limited by the semantic integrity of the label description, and it fails to fully utilize the synergistic effect of cross-modal information.

[0004] To address the above limitations, some researchers have started to explore the enhancing effect of visual modality information on text classification tasks. The CLIPTEXT framework proposed by Qin et al. was the first to apply the cross-modal ability of the CLIP model to zero-shot text classification, using images to enhance category information and verifying the effectiveness of the text-image matching method in zero-shot text classification. However, this method only focuses on visual image information and ignores the semantic knowledge embedded in text labels. Therefore, Zhang et al. further proposed the LABCLIP method, verifying the effectiveness of embedding label semantic information in image visual information in improving the performance of zero-shot text classification. In addition, to further improve the model performance, Wang et al. proposed the CLIPMulti framework, which innovatively adopts the text-image & text matching paradigm, aiming to achieve deeper multi-modal fusion by combining the image features and text features of labels. At the same time, to avoid the time cost of manually selecting text and reduce the uncertainty brought by manually writing text, the automated label generation strategy they proposed effectively reduces the dependence on manually designed label templates.

[0005] Although these methods have verified the effectiveness of combining visual image information in zero-shot text classification tasks to some extent, there are still challenges in aspects such as image source, the way of embedding label semantics in images, and multi-modal fusion: (1) When selecting images, the above methods mainly calculate the similarity between the image candidate set and the corresponding label name and select the image with the best performance as the mapped image of the label through verification in the validation set. However, label names are usually short and abstract, making it difficult to cover the diverse semantics in actual data. When relying solely on label names for image selection, there are text semantic deviations between some mapped images and real texts, resulting in the model being unable to clearly distinguish certain categories during the reasoning process, thus leading to uncertain or incorrect classification results. For example, when mapping the "Business" label to general business scene pictures (such as offices, stock charts), the model may not be able to distinguish the semantics of other subclasses; (2) When embedding label semantics in images, Zhang et al. combined semantic knowledge by embedding fixed label texts at fixed positions in the images. Although the classification performance was improved to some extent, it relied on manual design of label positions and contents, and the short label texts could not provide sufficient semantic knowledge to achieve the effect of reducing semantic deviation. In addition, different embedding label positions would affect the classification results of the model. For the entire dataset, the best effect was not achieved only at the center of the image; (3) In multi-modal fusion methods (such as matrix weighting), deep semantic information is easily lost due to normalization operations. In addition, there is still a gap between the automatically generated label descriptions and manual design in terms of semantic accuracy and domain adaptability, making it difficult to fully capture the core semantics of labels. Summary of the Invention

[0006] Based on this, it is necessary to provide a zero-shot text classification method and device based on cross-modal information completion for the technical problem of insufficient semantic integrity in cross-modal feature alignment.

[0007] A zero-shot text classification method based on cross-modal information completion, the method includes: Obtain a text sample set and an image candidate set corresponding to the labels.

[0008] Calculate the cosine similarity between the image corresponding to the label and the corresponding text sample and take the average value, and determine the mapping image candidate set for each label according to the obtained average cosine similarity.

[0009] According to the performance of the mapping image candidate set of each label on the validation set, select the image with the best performance as the best mapping image for the label.

[0010] Input the training set samples corresponding to each label into a large-scale pre-trained language model to generate a preset number of candidate keyword phrases, and form a keyword candidate set.

[0011] Perform multi-modal matching and selection according to the keyword candidate set, label mapping image, text sample, and label set to obtain the best keywords adapted to the domain.

[0012] According to the best mapping image of each label and the best keywords adapted to the domain, perform image annotation for text semantic enhancement to obtain an image with cross-modal information completion.

[0013] Process the test sample set and the image with cross-modal information completion using a dual-path enhancement strategy and then perform zero-shot inference through the CLIP model to obtain the text classification result.

[0014] A zero-shot text classification device based on cross-modal information completion, the device includes: A text and label image acquisition module for obtaining a text sample set and an image candidate set corresponding to the labels.

[0015] A mapping image candidate set construction module for the label, which is used to calculate the cosine similarity between the image corresponding to the label and the corresponding text sample and take the average value, and determine the mapping image candidate set for each label according to the obtained average cosine similarity.

[0016] The image generation module for cross-modal information completion is used to select the best-performing image as the best mapped image for the label according to the performance of the candidate set of mapped images for each label on the validation set; input the training set samples corresponding to each label into a large-scale pre-trained language model to generate a preset number of candidate keyword phrases to form a keyword candidate set; perform multi-modal matching and selection based on the keyword candidate set, the label mapped image, the text sample, and the label set to obtain the best keywords adapted to the domain; perform image annotation with enhanced text semantics based on the best mapped image for each label and the best keywords adapted to the domain to obtain an image with cross-modal information completion.

[0017] The text classification module is used to perform zero-shot inference through the CLIP model after processing the test sample set and the image with cross-modal information completion using a dual-path enhancement strategy to obtain the text classification result.

[0018] The above zero-shot text classification method and device based on cross-modal information completion. The method realizes the construction of an image with cross-modal information completion by designing a new image label mapping mechanism and generating based on context text. Finally, in the inference stage, a multi-modal collaborative inference framework is designed, and the text input is optimized through prompt engineering to enhance the text semantic perception ability, effectively improving the zero-shot inference ability of the CLIP model. Aiming at the semantic deviation problem of automatic label generation, based on context text generation, an image with cross-modal information completion is constructed to optimize the integrity of text semantic representation. At the same time, prompt design is introduced in the inference stage to enhance the text semantic perception ability. Further, a two-stage keyword automatic selection mechanism is proposed. First, a keyword candidate set is generated using a large model, and secondly, through multi-modal matching and selection, the best keyword phrase is selected as the text information with cross-modal information completion, effectively improving the semantic accuracy and domain adaptability. Brief Description of the Drawings

[0019] Figure 1 It is a schematic flowchart of the zero-shot text classification method based on cross-modal information completion in one embodiment; Figure 2 It is an overall architecture diagram of the zero-shot text classification method based on cross-modal information completion in one embodiment; Figure 3 It is a schematic structural diagram of the two-stage keyword automatic selection mechanism (Selection-CLIPCMIC) in another embodiment; Figure 4 It is a schematic diagram of the experimental verification results of the effectiveness of image selection in another embodiment; Figure 5 It is a schematic diagram of the comparison of the visualization results of the mapped images on the dataset Trec obtained by two different methods, Text-result and Label-result, in another embodiment; Figure 6 It is a schematic diagram of PCA visual analysis in another embodiment. Detailed implementation manners

[0020] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0021] The zero-shot text classification method based on cross-modal information completion proposed in the present application (CLIP-based Cross-modal Information Completion, abbreviated as: CLIPCMIC) improves the deficiencies of existing research through a multi-modal text semantic enhancement mechanism. Specifically, it includes the following three stages: (1) Semantic alignment label mapping: Based on the cross-modal representation consistency hypothesis, the present application designs a new mapping method, which maps the text label space to the image visual space in combination with the training set text, and constructs an interpretable visual semantic representation for each label; (2) Image construction for cross-modal information completion based on context text generation. Different from the previous method (LABCLIP), the present application uses a large-scale pre-trained language model and combines the training set text to automatically generate domain-related keyword phrases, and then obtains the optimal keywords for each label through an adaptive screening mechanism (combining cosine similarity calculation). Finally, through the automatic semantic annotation technology, the screened keywords are deeply fused with the image to construct a cross-modal representation with complete information, thereby enhancing the CLIP model's ability to capture fine-grained semantic knowledge; (3) Multi-modal collaborative reasoning: A dual-path enhancement strategy is designed in the reasoning stage. On the one hand, a prompt (here a prompt sentence is used) is introduced at the text input end, and a domain-adapted prompt template is constructed through context optimization; on the other hand, the image after text semantic information completion is used as the image input end, and then zero-shot reasoning is performed through the CLIP model.

[0022] In one embodiment, as Figure 1 shown, a zero-shot text classification method based on cross-modal information completion is provided, and the method includes the following steps: Step 100: Obtain a text sample set and an image candidate set corresponding to the labels.

[0023] Specifically, in order to use the CLIP model for zero-shot reasoning more effectively, this method converts the original "text-label" pair into a "text-image" pair.

[0024] Given a training set (where represents the number of samples, ), label set (where represents the number of tags). By calculating the cosine similarity between each text and its corresponding tag , for each tag , the top ten samples with the highest similarity are taken as the text sample set, that is , where represents the first sample with the highest similarity to the tag . For each tag , on the Google search engine, 20 images are manually selected according to the tag name to form the image set of this tag. Then the image candidate set corresponding to the tags of the entire data set is , where represents the first image manually selected according to the tag name.

[0025] The overall architecture of the zero-shot text classification method based on cross-modal information completion is as shown in Figure 2 .

[0026] Step 102: Calculate the cosine similarity between the image corresponding to the tag and the corresponding text sample and take the average. According to the obtained average cosine similarity, determine the mapping image candidate set for each tag.

[0027] Specifically, use the text encoder and image encoder of the CLIP model to perform text encoding and image encoding and normalization on the text samples and the images in the image candidate set corresponding to the tags respectively. Then calculate the cosine similarity between each candidate image and all text samples and take the average to obtain the cosine similarity and take the average; for each tag , a set of the top 5 images with the highest average cosine similarity score can be found as the mapping image candidate set for each tag.

[0028] Step 104: According to the performance of the mapping image candidate set of each tag on the validation set, select the image with the best performance as the best mapping image for this tag.

[0029] Specifically, as shown in Figure 2 , based on the training set , the label set (there are different class labels in total), randomly extract samples from the sample subset corresponding to each label as the validation set. The sample subset is , and the validation set ( , where is the sample index randomly selected from ), when the sample subset When this is the case, all samples are selected as the validation set. For each label, the performance of the top 5 images with the highest scores is recorded in the validation set in turn, and the image with the best result is selected as the best mapping image for that label. Therefore, with the help of the semantically aligned label mapping and combined with the performance in the validation set, the "text-label" pair can be mapped to the "text-image" pair , where represents the optimal mapping image.

[0030] Aiming at the semantic deviation problem caused by short and abstract label texts in the label mapping process, a new image label mapping method is proposed. Secondly, based on the context text generation, images for cross-modal information completion are constructed. Finally, a multi-modal collaborative reasoning framework is designed in the inference stage, and the text input is optimized through prompt engineering, effectively improving the zero-shot inference ability of the CLIP model.

[0031] Step 106: Input the training set samples corresponding to each label into a large-scale pre-trained language model to generate a preset number of candidate keyword phrases, forming a keyword candidate set.

[0032] Specifically, based on the context text generation, the training set samples corresponding to each label are sent as input data to the large-scale pre-trained language model to enhance its generation ability. For each label, ten candidate keyword phrases are generated by this model to form a keyword candidate set.

[0033] Keyword text generation: The training set samples corresponding to each label are sent as input data to the large model to enhance its generation ability. For each label, the model generates ten candidate keyword phrases to form a candidate set through user instructions.

[0034] Step 108: Perform multi-modal matching and selection based on the keyword candidate set, label mapping images, text samples, and label set to obtain the best keywords adapted to the domain.

[0035] Specifically, for multi-modal matching and selection: By comprehensively evaluating the semantic similarity between keywords and label mapping images, label texts, and typical text samples, a weighted summation strategy is adopted to select the keyword with the highest comprehensive score as the text semantic information to be complemented for the cross-modal information completion image generated in Step 110.

[0036] Regarding the semantic deviation problem in automated tag generation, to reduce the gap between the automatically generated tag descriptions and those designed manually in terms of semantic accuracy and domain adaptability, this application proposes a two-stage keyword automatic selection mechanism (Selection-CLIPCMIC), which optimizes the automatically generated tag descriptions into semantic representations of generated keyword phrases to enhance the semantic accuracy of information description. The two-stage keyword automatic selection mechanism (Selection-CLIPCMIC) specifically includes: first, using a large model to generate a candidate set of keywords, and second, through multimodal matching and selection, selecting the best keyword phrase as the text information for cross-modal information completion, effectively improving semantic accuracy and domain adaptability. Experimental observations show that the inference performance of the CLIP model is significantly affected by the different positions where keywords are annotated on images. Therefore, a method for automatically selecting the annotation position is proposed to enhance the robustness of cross-modal alignment.

[0037] The two-stage keyword automatic selection mechanism (Selection-CLIPCMIC) replaces short tag texts by using a pre-trained language model to generate domain-related candidate keywords, thereby providing more sufficient semantic knowledge and effectively improving semantic accuracy and domain adaptability.

[0038] Step 110: Based on the best-mapped image of each tag and the best keyword adapted to the domain, perform image annotation with enhanced text semantics to obtain an image with cross-modal information completion.

[0039] Specifically, this method constructs an image with cross-modal information completion by designing a new image tag mapping mechanism and generating based on context text to optimize the integrity of text semantic representation. At the same time, a prompt design is introduced in the inference stage to enhance the text semantic perception ability.

[0040] The automatic semantic annotation technology deeply fuses the filtered keywords with the image to construct an image with enhanced text semantics. The automatic semantic annotation technology not only alleviates the impact of the difference in the embedded tag position on the model performance but also effectively reduces the dependence on manual design.

[0041] Regarding the problem of information loss in cross-modal feature alignment, a new image tag mapping method is proposed, and the construction of an image with cross-modal information completion is realized by combining context text generation, significantly enhancing the zero-shot inference ability of the CLIP model. Further, a multi-modal collaborative inference framework is constructed, and the semantic representation of the text input is optimized through prompts (using prompt sentences), achieving a 1.8% performance improvement in the average (AVG) accuracy.

[0042] Step 112: After processing the test sample set and the image with cross-modal information completion using a dual-path enhancement strategy, perform zero-shot inference through the CLIP model to obtain the text classification result.

[0043] Specifically, as Figure 2 shown, to effectively complement the deep-level text semantic information, this method designs a dual-path enhancement strategy in the inference stage: First, introduce prompts in the text processing path and construct a domain-adapted prompt template through context optimization; Second, use the image data with complemented information as the input in the visual processing path. Specifically, taking the original text input "How far is it from Denver to Aspen?" in the figure as an example, to enhance the CLIP model's ability to extract text semantic features and optimize the knowledge transfer effect, this application adds a structured prompt prefix before the original input to generate guiding text "This text is about [subtitle]: How far is it from Denver to Aspen?". Combining with the cross-modal information-complemented image generated in step 110, the enhanced data of the two paths are then jointly input into the CLIP model for zero-shot inference. In addition, the prompt sentences corresponding to the six datasets are shown in Table 1.

[0044] Step 1: Obtain the "text-image" pair ; Step 2: Define it as ; Step 3: Introduce a text prompt, define it as (where represents the text after introducing the prompt, and represents the image after cross-modal information complementation). Finally, input the text and image together into the CLIP model for zero-shot prediction, defined as follows: ; Among them, and represent the text encoder and visual encoder of the CLIP model respectively. In the single-label text classification task, select the label with the highest probability as the final prediction result, and in the multi-label text classification task, select the labels greater than the threshold as the final prediction result.

[0045] Table 1 Prompt sentence templates for different datasets

[0046] To further complement the deep-level text semantic information, a multi-modal collaborative inference framework is proposed in the inference stage, a dual-path enhancement strategy is designed, and the text input is optimized through prompt engineering, significantly improving the zero-shot inference ability of the CLIP model.

[0047] In the above zero-shot text classification method based on cross-modal information completion, the method realizes the construction of images for cross-modal information completion by designing a new image label mapping mechanism and generating based on context text. Finally, in the inference stage, a multi-modal collaborative inference framework is designed to optimize the text input through prompt engineering to enhance the text semantic perception ability, effectively improving the zero-shot inference ability of the CLIP model. Aiming at the semantic deviation problem of automatic label generation, images for cross-modal information completion are constructed based on context text generation to optimize the integrity of text semantic representation. At the same time, prompt design is introduced in the inference stage to enhance the text semantic perception ability. Further, a two-stage keyword automatic selection mechanism is proposed. First, a keyword candidate set is generated using a large model. Second, through multi-modal matching and selection, the best keyword phrase is selected as the text information for cross-modal information completion, effectively improving semantic accuracy and domain adaptability.

[0048] In one embodiment, step 102 includes: using the image encoder and text encoder of the CLIP model to perform text encoding and image encoding on the image corresponding to the label and the corresponding text sample respectively and normalizing them to obtain normalized text encoding and normalized image encoding; calculating the similarity between each candidate image and all text samples using cosine similarity according to the normalized text encoding and normalized image encoding and taking the average value to obtain the average cosine similarity of each candidate image; for each label, selecting the top 5 candidate images from the corresponding image candidate set as the mapping image candidate set of the label according to the average cosine similarity.

[0049] Specifically, CLIP (Contrastive Language-Image Pretraining) is a cross-modal pre-training framework based on contrastive learning. Its core mechanism is to map heterogeneous modal data to a shared semantic space through large-scale text-image pair joint training. Different from the traditional pre-training paradigm that relies on task-specific labeled data, CLIP uses 400 million pairs of large-scale text-image pairs and, through the inter-modal contrastive learning strategy, makes relevant text and images as close as possible in this space, while keeping irrelevant text and images at a distance. The remarkable feature of this framework is reflected in its zero-shot inference ability, that is, it can directly classify by calculating the similarity between text and images without specific task training data. This feature enables it to demonstrate excellent zero-shot transfer ability in fields such as image classification, image generation, downstream visual tasks, and cross-modal text classification.

[0050] The CLIP model adopts a dual-encoder architecture, which consists of a text encoder and an image encoder respectively. The text encoder is mainly implemented based on the Transformer architecture, and pre-trained language models such as BERT and GPT are commonly used text encoders. The image encoder usually uses convolutional neural networks (CNNs), such as classic network structures like ResNet. These networks extract local features from images through multiple convolutional, pooling operations and fully connected layers, and gradually combine these features into global features. When the text encoder and the image encoder respectively generate the text embedding vector and the image embedding vector respectively, the model measures the semantic similarity between the text and the image by calculating the normalized dot product. Its calculation formula can be expressed as: ; where and represent the embedding vectors of the text and the image respectively, represents the L2 norm of the vector, and the superscript T represents the transpose.

[0051] Then, use the text encoder and the image encoder of the CLIP model to perform text encoding and image encoding on the text and the image respectively and normalize them; for each image perform feature encoding: ; where represents the L2 norm, represents the feature dimension.

[0052] For each label , the feature encoding of its text sample set is: ; where , , represents the feature dimension, and the superscript T represents the transpose.

[0053] Using the normalized embedding vectors, calculate the similarity between each candidate image and all text samples using cosine similarity and take the average. Specifically, for each candidate image , define its average similarity score as: ; where represents the feature of the th image, represents the feature of the th group of text samples. The th image and the lThe average cosine similarity of the group of text samples is: ; where represents the average cosine similarity between the th image and the th group of text samples.

[0054] Finally, perform optimal image selection. For the label , select the Top-5 images according to the average cosine similarity: ; After the above process, for each label , a group of the top 5 images with the highest average similarity score can be found. In step 104 of this method, each image in the validation set is verified one by one on the validation set, and the best image is selected as the mapped image for the label.

[0055] In one embodiment, step 106 includes: inputting the training set samples corresponding to each label into the GPT-4o large-scale pre-trained language model, and generating a group of relevant keyword phrases according to the user's prompt instructions to form a keyword candidate set.

[0056] Specifically, in order to provide sufficient text semantic knowledge and reduce the gap between the automatically generated label descriptions and the manually designed ones in terms of semantic accuracy and domain adaptability, a two-stage keyword automatic selection mechanism (Selection-CLIPCMIC) is proposed to optimize the automatically generated label descriptions into domain-adapted keyword phrase representations, that is, replacing the label semantic information embedded in the image with keyword semantic information.

[0057] As Figure 3 shown, in the first stage of the two-stage keyword automatic selection mechanism: the training set samples corresponding to each label are sent as input data to the large-scale pre-trained language model to enhance its generation ability. In this method, GPT-4o is used to generate the keyword phrases of the labels, and the prompt format used is: "The above document is the dataset <dataset>The label in <label category name>Training set samples, please summarize all the sample content in the above document with ten keywords”, among which <dataset>Indicates the dataset name, <label category name>Represents the label names in the dataset. Eventually, each label will have a set of associated keyword phrase candidates.

[0058] In one embodiment, step 108 includes: extracting features from the label mapping image, label, text sample, and keywords in the keyword candidate set to obtain image feature encoding, label feature encoding, text sample feature fusion, and keyword feature encoding; calculating the cosine similarity between the keywords and the label mapping image, label, and text sample respectively to obtain keyword-image semantic similarity, keyword-label semantic similarity, and keyword-text sample semantic similarity; performing weighted summation on the keyword-image semantic similarity, keyword-label semantic similarity, and keyword-text sample semantic similarity to obtain a comprehensive weighted scoring function; and making an optimal keyword decision based on the comprehensive weighted scoring function to obtain the best keywords adapted to the domain.

[0059] Specifically, in the first stage of the two-stage keyword automatic selection mechanism, multi-modal matching and selection are achieved: by comprehensively evaluating the semantic similarity between the keywords and the label mapping image, label text, and typical text samples, and adopting a weighted summation strategy to select the keyword with the highest comprehensive score as the text semantic information to be complemented in the second stage. The relevant formula definitions are as follows: Category label set , used to represent the labels of all categories in the dataset; Image set , used to represent the category label corresponding to the best mapping image; Text sample set , where , used to represent the text sample set corresponding to each category label, where ; Candidate keyword phrase set , where , used to represent the keyword phrases corresponding to each category label, where ; Optimal keyword set , used to represent an optimal keyword corresponding to each category label.

[0060] (1) Multi-modal feature extraction Image feature encoding: ; Where, is the image feature, is the result of the image encoder encoding the image corresponding to each label , represents the set of real numbers.

[0061] Label feature encoding: ; Among them, is the label feature, is the result of encoding each label by the text encoder.

[0062] Text sample feature fusion: ; Among them, is the sample feature, is the result of encoding the corresponding text sample of each label by the text encoder, is to take the mean, and here .

[0063] Keyword feature encoding: ; Among them, is the keyword feature, is the result of encoding each keyword by the text encoder, .

[0064] All of the above feature vectors are row vectors.

[0065] (2) Cosine similarity calculation Calculate the cosine similarity between the keyword and the image to measure the semantic consistency between the keyword and the associated image of the category , and the output is a scalar.

[0066] ; Among them, is the cosine similarity between the keyword and the image.

[0067] Calculate the cosine similarity between the keyword and the label to measure the text semantic consistency between the keyword and the category , and the output is a scalar.

[0068] ; Among them, is the cosine similarity between the keyword and the label.

[0069] Calculate the cosine similarity between the keyword and the text sample to measure the semantic consistency between the keyword k and the text sample of the category l .

[0070] ; Among them, is the cosine similarity between the keyword and the text sample.

[0071] (3)Optimal keyword phrase selection The comprehensive weighted scoring function is defined as follows: ; where is the comprehensive weighted scoring function, are three weights, and the specific values are determined according to the performance of the record model on the validation set. The values of the 6 datasets in the experiment are shown in Table 2.

[0072] Optimal keyword decision: ; where is the optimal keyword decision.

[0073] Through the above process, an optimal keyword phrase will be selected for each label in each dataset as the text semantic information for image semantic completion. Through this method, the automatically generated label descriptions are optimized into domain-adapted keyword phrase representations, effectively reducing the gap between the automatically generated label descriptions and the manually designed ones in terms of semantic accuracy and domain adaptability.

[0074] Table 2 Settings of weights and font sizes

[0075] To reduce the gap between the automatically generated label descriptions and the manually designed ones in terms of semantic accuracy and domain adaptability, this application designs a two-stage keyword automatic selection mechanism (Selection-CLIPCMIC) to optimize the automatically generated label descriptions into domain-adapted keyword phrase representations. The structure of the two-stage keyword automatic selection mechanism (Selection-CLIPCMIC) is as Figure 3 shown, and it is divided into two steps: (1)Keyword text generation: The training set samples corresponding to each label are sent as input data to the large model to enhance its generation ability. For each label, the model generates ten candidate keyword phrases to form a candidate set through user instructions.

[0076] (2)Multimodal matching and selection: By comprehensively evaluating the semantic similarities between the keywords and the mapped images of the labels, the label texts, and the typical text samples, the weighted summation strategy is used to select the keyword with the highest comprehensive score as the text semantic information to be complemented in step 104.

[0077] In addition, experimental observations show that the position where keywords are labeled on the image significantly affects the inference performance of the CLIP model. Therefore, a method for automatically selecting the labeling position is proposed to improve the robustness of cross-modal alignment.

[0078] In one embodiment, step 110 includes: determining the optimal position mapping of the domain-adapted best keywords on the corresponding best mapping image according to the adaptive position annotation selection method; quantitatively evaluating and comparatively analyzing the performance of different font sizes on the validation set to determine the optimal font parameter configuration of the domain-adapted best keywords; completing the annotation of the image according to the optimal position mapping and the optimal font parameter configuration of the domain-adapted best keywords to obtain an image with cross-modal information complemented.

[0079] In one embodiment, the specific steps of the adaptive position annotation selection method include: using the CLIP model to perform text and image feature encoding, and performing post-processing of normalization to obtain a text feature matrix and an image feature vector; calculating the cosine similarity between the image and all texts according to the image feature vector and the text feature matrix, and taking the average value to obtain the final evaluation index; selecting the position with the highest semantic consistency between the labeled image and the text sample set as the optimal position mapping of the text on the image according to the final evaluation index and the annotation effects of different preset positions.

[0080] Specifically, experimental observations find that the keyword phrase annotation position has a significant impact on the inference performance of the CLIP model. Therefore, a method for automatically selecting the annotation position is proposed to optimize cross-modal alignment.

[0081] In one embodiment, the expression of the optimal position mapping is: ; Among them, represents the optimal position mapping from the category label set to the preset position set, represents the process of labeling the keyword at the position p on the image , represents the text sample set, represents the final evaluation index.

[0082] In one embodiment, quantitatively evaluating and comparatively analyzing the performance of different font sizes on the validation set to determine the optimal font parameter configuration of the domain-adapted best keywords includes: designing the objective function and constraint conditions for dynamic font size adjustment; the expressions of the objective function and constraint conditions for dynamic font size adjustment are: ; ; Among them, is the objective function, represents the font size, , t represents the text content to be labeled (i.e., keyword phrase or text sample), which is a key variable in the constraint conditions, and represent the text width and height, W , H represents the image size.

[0083] When dynamic adjustment is enabled, the most suitable font size is optimized based on the constraint conditions as the optimal font parameter configuration for the best keywords adapted to the domain; if the model performance is quantitatively evaluated through the cross-validation set, and if the inference performance of the model still has not been significantly improved after dynamic adjustment, then the font size is manually adjusted to determine the best key adapted to the domain.

[0084] Specifically, finally, image annotation with text semantic enhancement is performed to achieve the construction of cross-modal information-complemented images. Through experiments, it is found that the method of directly embedding keyword phrases into the mapped image space has limited effect on improving the zero-shot inference ability of the CLIP model because for the entire dataset, the best effect can be achieved not only at the center of the image. In further analysis, it is found that the annotation position of keyword phrases and visual parameters such as font size have a significant impact on model performance. To solve the above problems, an adaptive position annotation selection method is proposed: First, the spatial layout of semantic information is optimized by designing an adaptive annotation position selection algorithm; second, regarding the impact of font size on model inference, systematic experiments are carried out on the validation set to quantitatively evaluate and comparatively analyze the performance of different font sizes, and finally the optimal font parameter configuration is determined. (Note: All the above methods are implemented through the CLIP model, and the text sample set mentioned in the algorithm is the same as the in step 100). The specific implementation process is as follows: The relevant formula definitions are as follows: The category label set , which is used to represent the labels of all categories in the dataset; The preset position set , which is used to represent the preset positions of text annotations; The image set , which is used to represent the best mapped image corresponding to each category label; The keyword phrase set , which is used to represent 1 keyword phrase corresponding to each category label; The text sample set , which is used to represent the text sample set corresponding to each category label and is used to calculate the similarity between the image and the text; Optimal position mapping , used to represent the position that selects the highest semantic consistency between the annotated image and the text sample; Optimized image set , including the annotated images with all category labels optimized, enhancing semantic expression and supporting downstream tasks.

[0085] (1) Image brightness calculation Used to adaptively select the text color to ensure that the text is clearly visible on the image.

[0086] ; Wherein is the input image, is the RGB-to-brightness conversion function, W , H represents the image size, is the average brightness scalar, and (x, y) represents the coordinates of the image pixel.

[0087] (2) Font size constraint optimization Design the font size constraint optimization strategy described by the objective function and constraint expression for dynamic font size adjustment to adjust the font size. The dynamic font size adjustment aims to ensure that the text information is clearly distinguishable within the image area and does not exceed the boundary. When dynamic adjustment is enabled, the system optimizes based on the constraints to select the most appropriate font size. To verify the effectiveness of the method, the model performance is quantitatively evaluated through a cross-validation set. If the inference performance of the model is not significantly improved after dynamic adjustment, manual adjustment is selected (Note: When dynamically adjusting, the font of the keyword phrase annotation of all label mapping images in each dataset will change, and manual adjustment is a unified change, and the specific values are shown in Table 2).

[0088] (3) Text image feature encoding Use the CLIP model to encode text and image features and map them to the same feature space for cosine similarity calculation. The image feature and text feature matrices are: ; ; Among them, represents the image feature (row vector), represents the text feature matrix (each row is a text feature, and the text comes from ), represents the number of text samples, where , represents the feature dimension, represents the L2 norm.

[0089] (4) Multi-modal Cosine Similarity Calculation Measure the semantic consistency between images and texts for optimizing text positions.

[0090] ; ; Among them, represents the cosine similarity matrix between the image and all texts, represents taking the mean of the cosine similarity scores for m texts as the final evaluation metric, represents the label used to indicate the set of text samples that are clearly related .

[0091] (5) Optimal Mapping Position By evaluating the annotation effects at different preset positions, select the position that maximizes the semantic consistency between the annotated image and the set of text samples . The optimal position mapping is as shown in the expression of the above optimal position mapping.

[0092] (6) Optimized Image Set Through keyword phrase annotation, the optimized annotated images containing all category labels not only enhance the semantic expression of the images but also support downstream tasks (such as classification, etc.).

[0093] ; Among them, represents the optimized image set, represents the category of the final optimized image.

[0094] This method realizes the construction of cross-modal information-complemented images, generating image representations with enhanced text semantics. This image representation will be used as the visual input of the multi-modal collaborative reasoning module in the third stage, jointly modeled with the corresponding text features, and finally realizing the cross-modal reasoning task under zero-shot conditions.

[0095] It should be understood that although each step in the flowchart of Figure 1 is shown in sequence according to the arrow indication, these steps are not necessarily executed in the order indicated by the arrow. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, Figure 1 At least some of the steps may include multiple sub-steps or multiple stages, which do not necessarily need to be completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least some of the sub-steps or stages of other steps or other steps.

[0096] In a validation example, to comprehensively evaluate the performance of CLIPCMIC in zero-shot text classification tasks, this study conducted experimental verification on six publicly available benchmark datasets, covering various task types such as multi-label classification, intent recognition, and topic classification. The characteristics of each dataset are as follows: Yahoo! Answers: A multi-label text classification dataset containing 10 significantly semantically different topic categories (such as science, health, sports, etc.) in the Yahoo! Q&A community. This dataset has the characteristics of language diversity of user-generated content and high distinctiveness between categories, and is suitable for verifying the zero-shot classification ability of the model in an open-domain scenario.

[0097] TREC (Text REtrieval Conference): A fine-grained question classification benchmark containing 6 types of question intents (such as fact-based, person-based, etc.). As an authoritative evaluation set released by the TREC conference, its annotation standardization and semantic complexity can be effectively used to evaluate the model's ability to understand question intents.

[0098] AG’s News: A news classification dataset covering 4 major news categories: business, politics, sports, and technology. Its large-scale and high-quality professional news corpus provides a reliable benchmark for evaluating the model's domain generalization ability in formal texts.

[0099] Subj: A subjectivity classification dataset containing two types of texts: subjective (such as comments, opinions) and objective (such as factual statements). The data size is moderate, and the texts are concise and clear, which can effectively verify the model's discriminative ability for language expression attributes.

[0100] SNIPS: An intent detection dataset containing 7 types of user intent texts (such as querying the weather, playing music, booking a restaurant, etc.). The data is sourced from text inputs of voice assistants and is suitable for testing the model's intent reasoning ability in colloquial expressions.

[0101] ATIS: An intent classification benchmark in the aviation field containing 15 professional intent categories such as flight queries and ticket bookings. As a classic evaluation set in the field of dialogue systems, its high density of professional terms and domain specificity provide an effective test platform for verifying the model's domain adaptation ability.

[0102] (1) Baseline method The method CLIPCMIC of this application is compared with the following benchmark models: 1) RTE: A classification method based on text entailment. The input text and label are regarded as an entailment problem, and a bert-base-uncased model is trained based on the RTE dataset for inference.

[0103] 2) FEVER: Similar to RTE, it uses a model based on bert-base-uncased and is pre-trained on the FEVER dataset to enhance the text entailment ability.

[0104] 3) MNLI: Also based on the idea of text entailment, it uses a bert-base-uncased model and is pre-trained through the MNLI dataset to improve the generalization ability.

[0105] 4) NSP: A zero-shot text classification method that directly uses the next sentence prediction pre-training task of BERT.

[0106] 5) NSP (Reverse): Since NSP is not for predicting directed semantic entailment, Ma et al. explored a variant that reverses the order of all text pairs and called it NSP (Reverse).

[0107] 6) GPT-2: A zero-shot text classification method based on a generative pre-training model that completes the classification task by directly generating the corresponding label output.

[0108] 7) CLIP Text Encoder: A text encoder based on Transformer that maps text features to a multimodal embedding space to achieve text classification tasks.

[0109] 8) CLIPTEXT: By reformulating zero-shot text classification as a text-image matching problem, it uses the cross-modal alignment ability of the CLIP model for classification.

[0110] 9) LABCLIP: By embedding label semantic information in images, it significantly enhances the ability to express label semantics.

[0111] 10) CLIPMulti (MF): A multimodal enhanced CLIP framework based on the text-image & text matching paradigm and through matrix fusion.

[0112] (2) Experimental results All datasets in the experiment were evaluated using the accuracy (acc). The experimental results show that (as shown in Table 3), except for the Yahoo! Answers dataset, the proposed method outperforms all baseline methods on the other five benchmark datasets. Among them, the ATIS dataset achieved the most significant performance improvement, with the classification accuracy increasing by 11.7% compared to the best baseline method; followed by the Subj dataset, with a 6.0% improvement. Regarding the issue of label distribution differences in the ATIS dataset, it was found through analysis that only 2 samples of the atis_day_name category were included in its test set, and this category did not appear in the public training set. To maintain the rigor of the experiment, this study evaluated the model under 15 standard intent categories and re-obtained the classification accuracy of the ATIS dataset according to the experimental structure of Zhang et al. to ensure the fairness of the comparative experiment. It is worth noting that under the same experimental conditions, compared with the LABCLIP and CLIPMulti baseline models, the proposed method achieved significant improvements of 4.5% and 3.7% respectively in terms of the average accuracy (AVG) metric. This result verifies that the proposed method significantly enhances the CLIP model's ability to capture text semantic features by effectively integrating cross-modal semantic information, thus achieving better generalization performance in the zero-shot text classification task.

[0113] Table 3 Experimental Results

[0114] Among them, # indicates that the experimental results are from Zhang et al., & indicates that it is reproduced on the public code, and other results are from the original paper. AVG represents the average score of Acc, and CLIPCMIC* indicates alignment with the CLIPMulti method, and the AVG of the same 5 datasets is calculated for comparison.

[0115] (3)Analysis of Experimental Results To systematically verify the effectiveness of the proposed method, the following experiments were carried out for analysis: First, verify the effectiveness of the image selection strategy; second, examine the performance of the automatic annotation position selection mechanism; then, explore the contributions of each module through ablation experiments; finally, conduct multi-modal feature analysis on the Subj dataset through visualization methods.

[0116] 1) Verification of Image Selection Effectiveness In the semantic tag mapping stage, this study proposes a text sample-based method to select the best image by calculating the cosine similarity between the image and its corresponding text sample and taking the average value. To verify the effectiveness of this method, it is compared with the method proposed by Qin et al. that calculates the similarity between the image and the corresponding tag name based on the tag name in the experiment. Both methods are experimented on the same validation set, and no additional text information or hint mechanism is introduced during the experiment.

[0117] The experimental results are as Figure 4 shown, where Label-result represents the image obtained by the original method, where the original method refers to the method of calculating similarity based on the tag name, Text-result represents the image obtained by this method, where this method refers to the method of calculating cosine similarity based on text samples proposed in this application, and AVG represents the corresponding average accuracy metric. It can be intuitively observed that the method proposed in this study is significantly better than the baseline method based on tag names on six datasets (the average improvement rate reaches 15.8%). It is verified that text samples have richer semantic representation capabilities than tag names and can effectively enhance the accuracy of cross-modal mapping. This method not only does not rely on tag names, but also reduces the semantic deviation between images and real texts and improves the model classification accuracy. It should be noted that on the Trec and ATIS datasets, there is a performance difference between the validation set and the test set in the experimental results of Label-result, resulting in the test results being significantly lower than Text-result. After analysis, it is found that this phenomenon stems from the overfitting problem of the baseline method based on tag names during this experiment.

[0118] The visualization results are as Figure 5 shown. Taking the Trec dataset as an example, the cross-modal feature distribution is presented. In this embodiment, the LDA method is used to perform dimensionality reduction and visualization analysis on the test set samples (circular markers) processed by CLIP, the images obtained by this method (pentagram markers), and the images obtained by the original method (diamond markers). In addition, × represents the clustering center. The experimental results show that the images obtained by this method (pentagram markers) are closer to the clustering center in the feature space than the images obtained by the original method (diamond markers), which indicates that text samples can more accurately capture the core features of categories through their richer semantic representation capabilities. This phenomenon reveals that the text-guided feature space has higher intra-class aggregation and inter-class discrimination, thus strengthening the robustness of cross-modal semantic alignment and providing a theoretical basis for the improvement of classification performance. It should be noted that since the results of the two mapped images corresponding to the ENTY tag are the same, they are presented in a superimposed state in the visualization graph.

[0119] 2) Verification of the effectiveness of automatically selecting the annotation position It was found in the experiment that different positions of keyword phrase annotation had a significant impact on the model performance. Therefore, an automatic semantic annotation technique was proposed, which could automatically select the annotation position. To verify the effectiveness of this method, this study conducted experimental verification on three datasets. The specific positions where keyword phrases were injected into the images included: center, top_left, top_right, bottom_left, bottom_right, top_center, and bottom_center.

[0120] The experimental results are shown in Table 4. The method was compared with seven fixed annotation position strategies on three benchmark datasets. The results showed that automatically selecting the annotation position was significantly better than the method with fixed annotation positions in terms of the Acc metric. It indicated that the annotation position played a key role in cross-modal information fusion, and at the same time proved that it was difficult for the method with fixed annotation positions to achieve optimal semantic injection. While automatically selecting the annotation position could effectively capture context-sensitive annotation positions, effectively improve the model's reasoning ability, and reduce the dependence on manually designed label positions.

[0121] Table 4 Experimental results for verifying the method of automatically selecting annotation positions

[0122] 3) Ablation experiment To evaluate the contributions of each module of the CLIPCMIC framework, relevant ablation experiments were designed. The experimental results are shown in Table 5. First, when the keyword phrase annotation module was removed, the performance of the CLIPCMIC framework decreased significantly, and the average accuracy (AVG) decreased by 3.4%. This indicates that the image construction module that realizes cross-modal information complementation based on context text generation helps to reduce the information loss problem in the cross-modal feature alignment process and, to a certain extent, complements the text semantic information, thus improving the zero-shot inference ability of the model. At the same time, this result also proves that the two-stage keyword automatic selection mechanism ((Selection-CLIPCMIC)) optimizes the automatically generated label descriptions into domain-adapted keyword phrase representations, significantly narrowing the gap in semantic accuracy and domain adaptability between the automatically generated label descriptions and the artificial design. Second, when the prompts were removed, the average accuracy (AVG) decreased by 1.8%, confirming that the domain-adapted prompts can effectively enhance the semantic understanding ability of the CLIP model. Finally, when both the keyword phrase annotation and the prompts were removed simultaneously, the average accuracy of CLIPCMIC on six datasets decreased by 4.5%, with the highest average performance decrease among the three groups of experiments, indicating that in the CLIPCMIC method, the two modules are indispensable and have a synergistic enhancement effect in this framework. It also shows that this method can effectively reduce the deep semantic information lost due to the normalization operation in the multi-modal fusion method.

[0123] Table 5 Ablation Experiments

[0124] Among them, w / o means removing the module, prompts means prompts, and keyword-pa means keyword phrase annotation.

[0125] 4) Visualization Analysis Taking the Subj benchmark dataset as the experimental object, the CLIP model was used to conduct a joint characterization analysis of multi-modal data features, specifically including the test set text, the semantically aligned label mapping images, the images after cross-modal information complementation, and the best keyword phrases. After extracting the feature vectors of each modality, principal component analysis (PCA) was used for dimensionality reduction processing and visualization analysis. The schematic diagram of the PCA visualization analysis is as Figure 6 shown. Figure 6 The text in the test set is represented by dots, the image after cross-modal information completion is represented by pentagrams, the label mapping image with semantic alignment is represented by diamonds, the best keyword phrases are represented by squares, and the cluster centers are represented by ×. Despite the existence of discrete data points, overall, the feature vectors of different modalities show obvious clustering characteristics in the embedding space, indicating that the CLIP model has strong multi-modal alignment capabilities. It is worth noting that the feature vectors of the images after cross-modal text semantic completion (the cluster centers marked with ×) are more spatially distributed closer to the cluster centers compared to the feature vectors of the label mapping images without semantic completion. This shows that the cross-modal information completion strategy proposed in this paper can significantly improve the semantic discrimination ability of visual representations, that is, the images after text semantic information completion can capture the core semantic information of the target category more accurately, thus achieving a consistent representation of the image and text bimodality in the semantic space.

[0126] The experimental results show that this method has achieved significantly better performance than the existing baseline methods on six public benchmark datasets. Compared with LABCLIP and CLIPMulti, the average accuracy has increased by 4.5% and 3.7% respectively. In addition, to systematically verify the effectiveness of the designed method in this study, extensive experimental verifications were carried out in this embodiment to prove the effectiveness of the method designed in this application.

[0127] In one embodiment, a zero-shot text classification device based on cross-modal information completion is provided, including: a text and label image acquisition module, a candidate set construction module for the mapping image of the label, an image generation module for cross-modal information completion, and a text classification module, where: The text and label image acquisition module is used to acquire the text sample set and the candidate set of images corresponding to the labels; The candidate set construction module for the mapping image of the label is used to calculate the cosine similarity between the image corresponding to the label and the corresponding text sample and take the average value, and determine the candidate set of the mapping image of each label according to the obtained average cosine similarity; The image generation module for cross-modal information completion is used to select the image with the best performance as the best mapping image of the label according to the performance of the candidate set of the mapping image of each label on the validation set; input the training set samples corresponding to each label into a large-scale pre-trained language model to generate a preset number of candidate keyword phrases to form a keyword candidate set; perform multi-modal matching and selection according to the keyword candidate set, label mapping image, text sample, and label set to obtain the best keywords adapted to the domain; perform image annotation with enhanced text semantics according to the best mapping image of each label and the best keywords adapted to the domain to obtain the image with cross-modal information completion; The text classification module is used to perform zero-shot inference on the test sample set and the images with cross-modal information completion processed by the dual-path enhancement strategy through the CLIP model to obtain the text classification results.

[0128] In one embodiment, the mapping image candidate set construction module for labels is further configured to perform text encoding and image encoding and normalization on the image corresponding to the label and the corresponding text sample respectively by using the image encoder and the text encoder of the CLIP model to obtain the normalized text encoding and the normalized image encoding; calculate the similarity between each candidate image and all text samples by using cosine similarity according to the normalized text encoding and the normalized image encoding, and take the average value to obtain the average cosine similarity of each candidate image; for each label, select the top 5 candidate images from the corresponding image candidate set as the mapping image candidate set of the label according to the average cosine similarity.

[0129] In one embodiment, the image generation module for cross-modal information completion is further configured to input the training set samples corresponding to each label into the GPT-4o large-scale pre-trained language model, and generate a set of relevant keyword phrases according to the user's prompt instructions to form a keyword candidate set.

[0130] In one embodiment, the image generation module for cross-modal information completion is further configured to extract features from the label mapping image, the label, the text sample, and the keywords in the keyword candidate set to obtain the image feature encoding, the label feature encoding, the text sample feature fusion, and the keyword feature encoding; calculate the cosine similarity between the keyword and the label mapping image, the label, and the text sample respectively to obtain the keyword-image semantic similarity, the keyword-label semantic similarity, and the keyword-text sample semantic similarity; perform weighted summation on the keyword-image semantic similarity, the keyword-label semantic similarity, and the keyword-text sample semantic similarity to obtain a comprehensive weighted scoring function; adopt an optimal keyword decision according to the comprehensive weighted scoring function to obtain the best keyword adapted to the domain.

[0131] In one embodiment, the image generation module for cross-modal information completion is further configured to determine the best position mapping of the best keyword adapted to the domain on the corresponding best mapping image according to the adaptive position annotation selection method; perform quantitative evaluation and comparative analysis on the performance of different font sizes on the validation set to determine the optimal font parameter configuration of the best keyword adapted to the domain; complete the image annotation according to the best position mapping and the optimal font parameter configuration of the best keyword adapted to the domain to obtain the image with cross-modal information completion.

[0132] In one embodiment, the specific steps of the adaptive position annotation selection method in the image generation module for cross-modal information completion include: using the CLIP model to perform text and image feature encoding, and performing post-processing of normalization to obtain a text feature matrix and an image feature vector; calculating the cosine similarity between the image and all texts according to the image feature vector and the text feature matrix, and taking the average value to obtain a final evaluation index; selecting the position with the highest semantic consistency between the annotated image and the text sample set as the best position mapping of the text on the image according to the final evaluation index and the annotation effects of different preset positions.

[0133] In one embodiment, the best position mapping is as shown in the expression of the best position mapping described above.

[0134] In one embodiment, the performance of different font sizes is quantitatively evaluated and comparatively analyzed on the validation set to determine the optimal font parameter configuration of the best keywords adapted to the domain, including: designing the objective function and constraint conditions for dynamic font size adjustment as shown in the expressions of the above objective function and constraint conditions.

[0135] When dynamic adjustment is enabled, the most suitable font size is optimally selected based on the constraint conditions as the optimal font parameter configuration of the best keywords adapted to the domain; if the model performance is quantitatively evaluated through the cross-validation set, and if the inference performance of the model still has not been significantly improved after dynamic adjustment, then manual adjustment of the font size is selected to determine the best keywords adapted to the domain.

[0136] For the specific limitations on the zero-shot text classification device based on cross-modal information completion, reference can be made to the limitations on the zero-shot text classification method based on cross-modal information completion in the above text, which will not be elaborated here. Each module in the above zero-shot text classification device based on cross-modal information completion can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or be independent of it, or can be stored in the memory in the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above modules.

[0137] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0138] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.< / label> < / dataset> < / label> < / dataset>

Claims

1. A zero-shot text classification method based on cross-modal information completion, characterized in that: The method comprises: Get the text sample set and the image candidate set corresponding to the label; Calculate the cosine similarity between the image corresponding to the label and the corresponding text sample and take the average. According to the obtained average cosine similarity, determine the candidate set of mapping images for each label; According to the performance of the candidate mapping image set of each label on the validation set, the image with the best performance is selected as the best mapping image of the label; Input the training set samples corresponding to each label into a large-scale pre-trained language model to generate a preset number of candidate keyword phrases to form a keyword candidate set; Perform multimodal matching and selection based on the keyword candidate set, label mapping image, text sample, and label set to obtain the best keyword adapted to the domain; According to the best mapping image of each label and the best keyword adapted to the domain, image annotation with text semantic enhancement is performed to obtain an image with cross-modal information completion; The test sample set and the cross-modal information-completed images are processed using a dual-path enhancement strategy and then subjected to zero-sample reasoning through the CLIP model to obtain text classification results.

2. The zero-shot text classification method based on cross-modal information completion according to claim 1, characterized in that: Calculate the cosine similarity between the image corresponding to the label and the corresponding text sample and take the average. According to the obtained average cosine similarity, determine the candidate set of mapping images for each label, including: The image encoder and text encoder of the CLIP model are used to respectively encode the image corresponding to the label and the corresponding text sample and normalize them to obtain the normalized text encoding and normalized image encoding; Calculating the similarity between each candidate image and all text samples using cosine similarity according to the normalized text code and the normalized image code, and taking an average value to obtain an average cosine similarity of each candidate image; For each label, the top five candidate images are selected from the corresponding image candidate set according to the average cosine similarity as the mapping image candidate set of the label.

3. The zero-shot text classification method based on cross-modal information completion according to claim 1, characterized in that: The training set samples corresponding to each label are input into the large-scale pre-trained language model to generate a preset number of candidate keyword phrases to form a keyword candidate set, including: The training set samples corresponding to each label are input into the GPT-4o large-scale pre-trained language model, and a group of related keyword phrases are generated according to the user's prompt instructions to form a keyword candidate set.

4. The zero-shot text classification method based on cross-modal information completion according to claim 1, characterized in that: Multimodal matching and selection are performed based on the keyword candidate set, label mapping image, text sample, and label set to obtain the best keywords adapted to the domain, including: Extracting features from the label mapping image, the label, the text sample, and the keywords in the keyword candidate set to obtain image feature coding, label feature coding, text sample feature fusion, and keyword feature coding; Calculate the cosine similarity of keyword and label mapping images, labels and text samples respectively to obtain the semantic similarity of keyword images, keyword labels and keyword text samples; The keyword image semantic similarity, keyword label semantic similarity and keyword text sample semantic similarity are weighted and summed to obtain a comprehensive weighted scoring function; The optimal keyword decision is adopted according to the comprehensive weighted scoring function to obtain the best keyword adapted to the field.

5. The zero-shot text classification method based on cross-modal information completion according to claim 1, characterized in that: According to the best mapping image of each label and the best keyword adapted to the field, image annotation with text semantic enhancement is performed to obtain an image with cross-modal information completion, including: Determine the best position mapping of the best keyword adapted to the domain on the corresponding best mapping image according to the adaptive position annotation selection method; Conduct quantitative evaluation and comparative analysis on the performance of different font sizes on the validation set to determine the optimal font parameter configuration for the best keywords adapted to the domain; According to the optimal position mapping of the best keywords adapted to the field and the optimal font parameter configuration, the image is annotated to obtain an image with cross-modal information completion.

6. The zero-shot text classification method based on cross-modal information completion according to claim 5, characterized in that: The specific steps of the adaptive position labeling selection method include: The CLIP model is used to encode text and image features, and normalized post-processing is performed to obtain text feature matrix and image feature vector; According to the image feature vector and the text feature matrix, the cosine similarity between the image and all the texts is calculated, and the average value is taken to obtain the final evaluation index; According to the final evaluation index and the evaluation of the annotation effects of different preset positions, the position that makes the semantic consistency between the annotated image and the text sample set the highest is selected as the optimal position mapping of the text on the image.

7. The zero-shot text classification method based on cross-modal information completion according to claim 5, characterized in that: The optimal position mapping is: in, Represents the optimal position mapping from the category label set to the preset position set, Indicates the label, Indicated in the image Position p Mark keywords process, represents a collection of text samples, Represents the final evaluation indicator.

8. The zero-shot text classification method based on cross-modal information completion according to claim 5, characterized in that: The performance of different font sizes is quantitatively evaluated and compared on the validation set to determine the optimal font parameter configuration for the best keywords adapted to the domain, including: The objective function and constraints for designing dynamic font size adjustment are: in, is the objective function, Indicates the font size. , and Indicates the text width and height, t Indicates the text content to be annotated. W , H Indicates the image size; When dynamic adjustment is enabled, the most appropriate font size is optimized and selected based on the constraints as the optimal font parameter configuration for the best keyword adapted to the domain; if the model performance is quantitatively evaluated through a cross-validation set, if the reasoning performance of the model has not been significantly improved after dynamic adjustment, the font size is manually adjusted to determine the optimal font parameter configuration for the best keyword adapted to the domain.

9. A zero-shot text classification device based on cross-modal information completion, characterized in that: The device comprises: A text and label image acquisition module is used to obtain a text sample set and an image candidate set corresponding to the label; The label mapping image candidate set construction module is used to calculate the cosine similarity between the image corresponding to the label and the corresponding text sample and take the average value, and determine the mapping image candidate set of each label based on the obtained average cosine similarity; The image generation module of cross-modal information completion is used to select the best-performing image as the best mapping image of the label according to the performance of the mapping image candidate set of each label on the verification set; input the training set samples corresponding to each label into the large-scale pre-trained language model to generate a preset number of candidate keyword phrases to form a keyword candidate set; perform multimodal matching and selection based on the keyword candidate set, label mapping image, text sample, and label set to obtain the best keyword adapted to the domain; perform text semantics enhanced image annotation based on the best mapping image of each label and the best keyword adapted to the domain to obtain an image of cross-modal information completion; The text classification module is used to process the test sample set and the image with cross-modal information completion using a dual-path enhancement strategy, and then perform zero-sample reasoning through the CLIP model to obtain a text classification result.

Citation Information

Patent Citations

  • Prompt learning method for modal interaction enhancement of visual language model

    CN116503683A

  • Zero sample text classification method and system based on CLIP and medium

    CN116701637A

  • Zero-sample text classification method based on cross-language integration

    CN118332127A

  • Global visual guidance image description generation method based on cross-modal large model

    CN118378623A

  • Multimodal few-shot learning with frozen language models

    US20240282094A1

Cited By

  • User intention recognition method based on multi-modal model

    CN121479342A

  • Muskmelon powdery mildew scab detection method and system combining text label self-prompting and soft hypergraph reasoning

    CN121686474A

  • A method and system for detecting melon powdery mildew lesion by combining text label self-prompting and soft supergraph reasoning

    CN121686474B

  • Time sequence physiological signal zero sample matching method and system based on cross-domain metric matrix

    CN122220903A

  • Cross-domain metric matrix based zero sample matching method and system for time-series physiological signals

    CN122220903B