Long-tail image recognition method based on multi-modal semantic generation and image-text fusion
By combining the multimodal visual language model and the CLIP model, high-quality tail image samples are generated and screened, which solves the problem of insufficient semantic expression of tail samples under long-tail distribution, and improves the generalization ability and recognition accuracy of the image recognition model.
Patent Information
- Application Number
- CN202510515241.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-08-15
AI Technical Summary
In the image recognition task under long tail distribution, the semantic expression ability and visual diversity of tail samples are insufficient, resulting in insufficient generalization ability of the model in extremely scarce data scenarios.
The multimodal visual language model is used to extract the structured semantic description of tail images, and the image semantic extension description is generated through semantic rewriting and style alignment. The CLIP model is used for semantic consistency and visual quality screening, and an enhanced image set is constructed, and the graphics and text fusion classification model is trained.
It significantly improves the discriminant ability of tail samples and the cross-modal generalization characteristics of the model, ensures the semantic accuracy and visual diversity of the generated samples, reduces the style shift of the training-test domain, and enhances the recognition performance of the model under the long-tail distribution.
Smart Images

Figure CN120495814A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition technology, and in particular to a long-tail image recognition method based on multimodal semantic generation and image-text fusion. Background Art
[0002] In real-world image recognition tasks, data often exhibits a long-tail distribution, meaning a small number of "head" categories have a large number of samples, while a large number of "tail" categories have very few. This data imbalance causes deep learning models to favor the head categories during training, resulting in weak representation of the tail categories and poor classification performance, seriously affecting the overall system's generalization capabilities.
[0003] To alleviate the long-tail distribution problem, existing technologies primarily focus on optimizing model structures and designing loss functions. For example, methods such as class reweighting, resampling, and feature-classifier decoupling have improved tail class performance to a certain extent. However, these methods still struggle to achieve good results in extremely scarce data scenarios.
[0004] In recent years, generative models have brought new directions to the expansion of tail-category samples. Some methods use image-to-image generation to enhance tail-category images. However, due to the lack of semantic guidance, these methods often only produce local color or texture changes, making it difficult to improve structural diversity and semantic expression, resulting in limited practical value of the generated samples. Some research has begun to explore text-to-image generation methods based on text descriptions, generating tail-category images by guiding diffusion models using natural language descriptions, and has made some progress. However, existing methods still have shortcomings in description quality, diversity control, and generated image screening mechanisms, which hinder the performance improvement of tail-category recognition models.
[0005] Therefore, there is an urgent need for a systematic approach that combines semantic control, sample screening, and collaborative modeling of images and text to expand the semantic expression ability and visual diversity of tail class samples from the source and improve their effective utilization in classification models. Summary of the Invention
[0006] In order to overcome the defects and shortcomings of the existing technology, the present invention provides a long-tail image recognition method based on multimodal semantic generation and image-text fusion, which significantly enhances the discrimination ability of tail class samples while maintaining high accuracy and has stronger cross-modal generalization characteristics.
[0007] In order to achieve the above object, the present invention adopts the following technical solutions:
[0008] The present invention provides a long-tail image recognition method based on multimodal semantic generation and image-text fusion, comprising the following steps:
[0009] Obtain tail images, perform semantic modeling on the tail images based on a multimodal visual language model, and extract structured semantic descriptions;
[0010] Taking the structured semantic description as the basic input, the semantics are rewritten and enhanced based on the multimodal visual language model to generate an extended semantic description of the image;
[0011] The sentence vector semantic duplication detection method is used to remove redundant descriptions from the image semantic extended description. The style of the image semantic extended description after removing redundant descriptions is then aligned based on the image style prompt words to obtain the optimized text description set.
[0012] The optimized text description set is input into the text-generated graph model to generate tail-category image samples, which are then screened for semantic and visual quality to construct an enhanced image set for training.
[0013] A training dataset is constructed based on the original long-tail dataset and the enhanced image dataset to train the image-text fusion classification model;
[0014] The image to be identified is input into the trained image-text fusion classification model, and the classification results of all categories are output.
[0015] As a preferred technical solution, semantic modeling of tail images is performed based on a multimodal visual language model to extract structured semantic descriptions, specifically including:
[0016] A structured description template is constructed based on the typical features of tail-category images. The tail-category images are input into a multimodal visual language model. Based on the guidance of the structured description template, semantic description text of the tail-category images is generated, and formatting and grammar checking are performed to obtain a set of structured semantic description texts.
[0017] As a preferred technical solution, redundant descriptions are removed from the image semantic extended description based on the sentence vector semantic duplication detection method, specifically including:
[0018] Based on the sentence vector embedding model, each image semantic extension description is mapped to a high-dimensional semantic space, and vectorized duplication detection is performed on the image semantic extension description;
[0019] For each newly generated image semantic extension description e new , calculate its cosine similarity with all descriptions in the currently retained description set ε, and take the maximum value, which is specifically expressed as:
[0020] sim(e new ,ε)=maxcos(φ(e new ),φ(e i )),e i ∈ε
[0021] Among them, φ() represents the semantic vector representation of the description text, and cos(·,·) represents the cosine similarity function;
[0022] Set the semantic duplication threshold τ, if sim(e new ,ε)>τ, then the description is considered to be semantically duplicated with the existing description and the description is not retained; otherwise, the description is retained.
[0023] As a preferred technical solution, style alignment of the image semantic extended description after removing redundant descriptions is performed based on image style prompt words, specifically including:
[0024] The image encoder based on the CLIP model extracts the tail class image features, performs KMeans clustering to obtain multiple style cluster centers, selects a representative tail class image for each cluster, and inputs it into the multimodal model to generate style description;
[0025] The style description is spliced into the image semantic extended description in the form of phrases after removing redundant descriptions, and style alignment is performed based on the constructed image style clue words.
[0026] As a preferred technical solution, after obtaining the optimized text description set, a category extension description balancing step is further included, specifically including:
[0027] The number of target descriptions for each category is set to a fixed value. If it is determined that the semantic extended description of a certain type of image is less than the number of target descriptions of the category, it is re-expanded. If it is determined that the semantic extended description of a certain type of image exceeds the number of target descriptions of the category, random sampling is performed according to the semantic distance distribution, and finally a set of semantic extended descriptions of each category of images is output.
[0028] As a preferred technical solution, semantic and visual quality screening is performed to construct an enhanced image set for training, specifically including:
[0029] Based on the CLIP model, the image and text are encoded separately and the semantic consistency score is calculated, which is expressed as:
[0030] s a li gn (I i ,T i )=cos(φi mg (I i ),φ text (T i ))
[0031] Among them, I i represents the i-th generated tail image, T i Indicates its corresponding text description, φ img (·) and φ text(·) represents the image encoder and text encoder of the CLIP model, cos(·,·) represents the cosine similarity;
[0032] The visual quality of the generated tail class images is scored based on the pre-trained aesthetic scorer, expressed as:
[0033] s qua li t y (I i )=f aesthetic (I i )
[0034] Among them, I i represents the tail class image generated by the i-th image, f aesthetic represents the pre-trained aesthetic scorer;
[0035] A double-threshold mechanism is used to screen the generated tail class image samples, and the generated tail class images that meet both the semantic consistency score greater than or equal to the score threshold and the visual quality score greater than or equal to the score threshold are retained.
[0036] As a preferred technical solution, during the steps of performing semantic and visual quality screening and constructing an enhanced image set for training, multiple rounds of closed-loop optimization are performed, specifically including:
[0037] Input the tail image samples retained in the previous round of screening into the multimodal visual language model to generate a new image feature description f (t) , combined with the original semantic description e from the previous round (t) With image style prompt words k , build a new prompt template e (t +1) , expressed as:
[0038] e (t+1) =Refine(e (t) ,f (t) ,s k )
[0039] Refine represents the description refinement function, which is used to integrate image content feedback and style diversity information;
[0040] Set the new prompt template (t+1) Input the text graph model to generate image I (t+1) , repeatedly calculate the semantic consistency score and visual quality score, use the double threshold mechanism to filter the generated tail class image samples, and obtain a new round of retained tail class image samples;
[0041] When the set iterative termination condition is met, the closed-loop process ends and the final enhanced image set is output.
[0042] As a preferred technical solution, the set iteration termination conditions include: meeting the set number of iterations, the number of retained images reaching the set number threshold, and the change in the pass rate of two adjacent rounds of screening being less than the set pass rate threshold.
[0043] As a preferred technical solution, a training dataset is constructed based on the original long-tail dataset and the enhanced image set to train the image-text fusion classification model, specifically including:
[0044] Construct a dual-branch network of images and texts. In the image branch, extract features of the input image x based on the image encoder to obtain the image representation vector φ img (x), the description text t of each category based on the text encoder in the text branch c Encode and obtain the category embedding vector φ text (t c ), and normalize the embedding set of all categories;
[0045] The cosine similarity between the image representation vector and each category embedding vector is calculated separately to form the prediction score of the text branch, which is expressed as:
[0046] logits text (x) = cos(φ img (x),φ txt (t c ))
[0047] The image branch passes through the linear classification head W vis Output classification score, expressed as:
[0048] logits vis (x) = W vis (φ img (x))
[0049] The prediction results in the fusion image branch and the text branch are used as the final classification output, which is expressed as:
[0050] logits final (x) = α·logits text (x)+(1-α)·logits vis (x)
[0051] Among them, α represents the learnable fusion weight parameter;
[0052] The original long-tail dataset and the enhanced image set are input into the image-text dual-branch network for training to obtain the trained image-text fusion classification model.
[0053] As a preferred technical solution, the original long-tail dataset and the enhanced image set are input into the image-text dual-branch network for training, and the standard cross entropy loss is used as the training target. Only the linear classification head W is optimized. vis The weight parameter α is fused and the rest of the encoder parameters remain frozen.
[0054] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0055] (1) The present invention introduces a multimodal visual language model (such as Qwen2.5-VL) to perform real perception and semantic extraction of tail class images. The extracted structured text description covers multiple dimensions such as color, texture, posture, and environment. It is fine-grained and visually consistent. Compared with the traditional method of using a pure language model to generate descriptions based on category names, the present invention significantly reduces text hallucination problems and semantic deviations, effectively ensures a high degree of consistency between the generated text and the original image, and improves the semantic accuracy of the tail class extension data from the source.
[0056] (2) The present invention adopts a multi-round Prompt semantic guidance mechanism in the description expansion stage, introducing fine-grained elements such as posture changes within the category, background complexity, and interactive actions, to improve the coverage of the tail category description in the local semantic space; at the same time, combined with the style keywords obtained based on the clustering of the original training images, style alignment descriptions are automatically generated, so that the generated samples are close to the real image distribution in visual style, thereby significantly reducing the style offset between the training and test domains and effectively enhancing the generalization ability of the model.
[0057] (3) To address the semantic ambiguity and sample pattern collapse problems of traditional generative methods, the present invention designs a multi-round "description-generation-screening" closed-loop optimization process: each round uses high-quality generated samples as feedback, dynamically refines the prompt description and adjusts the style information, so that the subsequent generated samples gradually approach the distribution of real tail class images in terms of semantic features and visual details. This closed-loop mechanism can continuously improve the structural diversity and detail realism of the generated images, which is significantly better than the generation effect of a single-round generation method and has stronger sample enhancement capabilities and task adaptability.
[0058] (4) In the sample generation and screening stage, the present invention introduces dual thresholds of semantic consistency scoring based on the CLIP model and image quality scoring based on the aesthetic scoring model, retaining only samples with high semantic accuracy and good visual quality, significantly suppressing the interference of low-quality images on downstream classification training, effectively controlling the noise risk introduced by the generated samples, and ensuring the stability and discriminability of the enhanced data.
[0059] (5) In the tail class recognition stage, the present invention adopts a frozen pre-trained CLIP encoder to extract image and text features respectively, and introduces learnable fusion weights to make collaborative decisions at the feature layer. Only the linear classification head and fusion parameters are trained, which greatly reduces the training overhead and overfitting risk. Compared with the method of directly fine-tuning the visual backbone model, the image-text fusion recognition model constructed by the present invention significantly enhances the discrimination ability of tail class samples while maintaining high accuracy, and has stronger cross-modal generalization characteristics.
[0060] (6) The present invention has carried out systematic collaborative design in all stages from tail class semantic modeling, description diversity control, sample generation quality assurance to image-text fusion classification training. The data flow between each module is closed-loop and the goals are consistent, so that the tail class feature expression, sample distribution adaptability and model recognition performance are improved as a whole. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 Schematic diagram of the process of the long-tail image recognition method based on multimodal semantic generation and image-text fusion of the present invention;
[0062] Figure 2 A schematic diagram of the multi-round image generation optimization process based on image feedback of the present invention;
[0063] Figure 3 This is a schematic diagram of the training of the image-text fusion classification model of the present invention. DETAILED DESCRIPTION
[0064] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0065] Example
[0066] like Figure 1 As shown, this embodiment provides a long-tail image recognition method based on multimodal semantic generation and image-text fusion, which mainly includes four stages: tail image acquisition and semantic modeling, text expansion and screening, image generation and enhancement, and model training and reasoning. Specifically,
[0067] Step S1: Obtain the tail class image and extract the semantic description;
[0068] In this embodiment, we first obtain original tail class image data for training and enhancement. Tail classes refer to a set of categories whose number in training samples is significantly smaller than that of mainstream categories. Tail class images can come from public long-tail datasets (such as ImageNet-LT and iNaturalist) or real data collected from enterprises or specific tasks.
[0069] After collecting the tail class images, a multimodal visual language model with image-text alignment capability (such as Qwen2.5-VL) is used to perform semantic modeling on the images and extract structured semantic descriptions.
[0070] The specific steps include:
[0071] Step S11: constructing a structured Prompt template;
[0072] Based on the typical content characteristics of tail-category images, a structured description template (Prompt template) is designed to guide the multimodal visual language model in generating high-quality semantic descriptions. The Prompt template contains slots for semantic elements such as category name, color, texture, posture, and background. An example template includes: "This image is of [category], with an object of [color] [texture], located in [posture] in [environment]."
[0073] Step S12: input the image and generate an initial semantic description;
[0074] The tail image is input into the multimodal visual language model and guided by the Prompt template in step S11 to generate a semantic description text of the image. The multimodal visual language model realizes the image-to-text conversion through visual perception and language generation capabilities, ensuring that the generated text accurately reflects the image content.
[0075] Step S13: formatting and grammar checking the output statement;
[0076] To ensure the consistency of semantic structure and subsequent processing efficiency, the generated description results are cleaned up and standardized, including standardizing punctuation, removing duplicate words, and merging synonyms. This process can be automated by combining natural language processing toolkits (such as spaCy and jieba).
[0077] The final output is a set of structured semantic description texts that correspond one-to-one with the original tail class images, providing a high-quality semantic foundation for subsequent description expansion, image generation, and model training.
[0078] Step S2: semantic expansion and description screening;
[0079] Based on the structured semantic description obtained in step S1, its semantic diversity is further expanded, and the description quality is optimized through duplicate detection and style alignment mechanisms to ensure that the subsequent image generation task has good semantic support and distribution consistency;
[0080] The specific steps include the following:
[0081] Step S21: multiple rounds of semantic expansion to generate diverse descriptions;
[0082] Using the original description as the basic input, a multimodal visual language model (such as Qwen2.5-VL) is guided to perform semantic rewriting and enhancement. By using a custom prompt template, such as "Please generate five semantically similar but non-repetitive texts for the following description, reflecting different actions, backgrounds, colors, or postures," the model is guided to generate differentiated extended versions. Each extended description maintains the main semantics of the category while enhancing fine-grained feature variations, thereby increasing the diversity and richness of the description.
[0083] Step S22: removing redundant descriptions based on a semantic vector-based duplicate detection mechanism;
[0084] To avoid invalid expansions such as semantic duplication and sentence rewriting during the expansion process, a sentence embedding model (such as Sentence-BERT and all-MiniLM-L6-v2) is used to map each description into a high-dimensional semantic space, and vectorization is performed on the expanded description set to detect duplicates.
[0085] For each newly generated extended description e new , calculate its cosine similarity with all descriptions in the currently retained description set ε, and take the maximum value, which is defined as follows:
[0086] sim(e new ,ε)=maxcos(φ(e new ),φ(e i )),e i ∈ε
[0087] Among them, φ(e) represents the semantic vector representation of text e, and cos(·,·) represents the cosine similarity function. If sim(e new ,ε)>τ, where τ is the set semantic duplication threshold, the description is considered to be semantically duplicated with the existing description and is not retained; otherwise, it is retained. In this embodiment, the threshold τ is set to 0.3.
[0088] The above process can be implemented through Python programming combined with vector libraries (such as Faiss or numpy) to achieve batch calculation and filtering, thereby constructing a semantic description set with good diversity and low information redundancy.
[0089] Step S23: Constructing style prompt words based on the image style clustering results;
[0090] The CLIP model's image encoder extracts representations from the tail-category raw image features in the training set, and performs K-means clustering (e.g., K=10) to obtain several style cluster centers. A representative image is selected for each cluster and fed into a multimodal model to generate a style description (e.g., "low-angle shot," "natural light," "complex background"). This style description is then concatenated into the extended description in the form of a phrase to construct a style prompt with style alignment capabilities, effectively narrowing the style gap between the generated samples and the real images.
[0091] Step S24: Control the number of descriptions in each category to maintain a balance between categories;
[0092] In order to avoid new distribution bias caused by uneven number of descriptions between categories in the subsequent training stage, the number of target descriptions for each category is set to a fixed value T. If a category is not expanded enough, it can be expanded again; if the description exceeds the target value, it is randomly sampled according to the semantic distance distribution, ultimately ensuring that each category has a uniform number of well-distributed high-quality description sets.
[0093] The final output is an extended description set for each category, which has high semantic diversity, low redundancy, style alignment capability, and controlled quantity balance, serving as the direct input of the subsequent image generation module.
[0094] Step S3: text-driven image generation and screening;
[0095] This step is based on the high-quality text description set obtained in step S2. The text-based graph model is used to synthesize tail-category image samples, and a screening mechanism is used to eliminate samples of poor quality or semantic deviation to construct an enhanced image set for training.
[0096] The specific steps include the following:
[0097] Step S31: Generate tail class image samples using the Wensheng graph model;
[0098] Each filtered and style-aligned text description is input into a text-to-image generation model (such as StableDiffusion 2.1) to perform text-guided image generation. To maintain a one-to-one correspondence between descriptions and images, this embodiment sets each text description to generate one image. The generation uses the default sampling configuration, and the image size is uniformly 512×512 pixels. The prompt template example is: "A photo of the class brambling, with brown and white plumage, flying pose, located on a stone wall with moss". After input into the diffusion model, an image with a semantic structure is generated.
[0099] Step S32: Calculate the semantic consistency score;
[0100] To evaluate whether the generated image accurately expresses the corresponding text description, the CLIP model is used to encode the image and text separately and calculate their semantic consistency score, which is defined as follows:
[0101] s a li gn (I i ,T i )=cos(φi mg (I i ),φ text (T i ))
[0102] Among them, I i represents the i-th generated image, t l Indicates its corresponding text description, φ img (·) and φ text (·) represent the image encoder and text encoder of the CLIP model respectively, and cos(·,·) represents the cosine similarity. align (I i ,T i )<τ align , then the image is judged to be inconsistent with the description semantics and will not be retained. In this embodiment, the threshold τ align Set to 0.3.
[0103] Step S33: Evaluate the image visual quality score;
[0104] To ensure the perceptual quality of generated samples, a pre-trained aesthetic scorer is introduced to score the generated images. The score range is [0, 10]. A higher score indicates a higher quality image in terms of clarity, composition, color coordination, etc. The image quality score is defined as:
[0105] s qua lit y (I i )=f aesthetic (I i )
[0106] If s quality (I i )<τ quality , then the image is considered to have poor visual quality and will not be retained. In this embodiment, the threshold τ quality Set to 5.5.
[0107] Step S34: Joint screening to generate image samples;
[0108] Combining the above two scoring results, a dual-threshold mechanism is used to screen high-quality samples and retain generated images that meet the following conditions:
[0109] s align (I i ,T i )≥τ align ands quality (I i )≥τ quality
[0110] Only images that meet both semantic consistency and visual quality requirements are retained as augmented samples. The screening process supports batch automation to ensure stable and consistent quality across the entire augmented sample set.
[0111] Step S35: multiple rounds of description refinement and image generation optimization based on image feedback;
[0112] In order to further improve the structural diversity and visual authenticity of the generated samples, this paper designs a multi-round closed-loop optimization process based on image feedback and semantic refinement on the basis of the initial round of image generation. Figure 2 As shown in the figure, this process guides the generated samples to gradually approach the real image distribution in the semantic and style space through repeated refinement of text descriptions and multiple rounds of generation screening.
[0113] The closed-loop optimization process includes the following three aspects:
[0114] (1) Image feature feedback and description refinement;
[0115] The image samples I retained in the previous round of screening (t) Input into a multimodal visual language model (such as Qwen2.5-VL) to generate a new image feature description f (t) , extract details such as color, posture, background, etc. Combined with the original description of the previous round (t) Style Tips k , build a new refinement Prompte (t+1) :
[0116] e (t+1) =Refine(e (t) ,f (t) ,s k )
[0117] Among them, Refine represents the description refinement function, which is used to integrate image content feedback and style diversity information to generate the next round of description with richer semantics and more visual specificity.
[0118] (2) A new round of image generation and screening;
[0119] e(t+1) As a prompt input, the Wensheng graph model (Stable Diffusion model) generates image I (t+1) , and repeat the semantic consistency calculation and visual quality scoring in steps S32 and S33, and screen according to the double threshold conditions to obtain a new round of retained samples.
[0120] (3) Loop control and termination conditions;
[0121] The closed-loop process can be iteratively executed T max The optimization process terminates prematurely if any of the following conditions are met: the expected number of retained images is reached; the change in the pass rate between two consecutive rounds is less than a set threshold (e.g., 2%); the description content and image style converge, resulting in insufficient diversity gain in the generated samples. Furthermore, if all images generated in the current round fail the semantic consistency or visual quality screening, the description is deemed incapable of generating high-quality samples under the current expansion conditions. The multi-round optimization process for that description is automatically terminated, and the generation process continues with the next description. This strategy avoids idleness, improves generation efficiency, and ultimately ensures the semantic validity and visual quality of the enhanced samples.
[0122] Through this closed-loop mechanism, generated samples can gradually approximate the distribution of real images over multiple rounds of refinement, overcoming the common structural similarity and loss of detail issues in the initial generation round, significantly improving the quality of tail-category sample expansion. This mechanism can be flexibly embedded in the image generation process, providing continuous quality optimization and enhanced style diversity for generated samples.
[0123] Ultimately, the retained image samples and their corresponding text descriptions constitute a structured enhanced sample set, which serves as an important input for the subsequent model training stage.
[0124] Step S4: Image-text fusion classification modeling;
[0125] In this step, based on the constructed enhanced sample set, the long-tail dataset and the enhanced samples are combined to construct a training dataset, and the classifier is trained by fusion modeling of image and text features. Figure 3 As shown in the figure, this stage is based on the frozen multimodal model and uses the collaborative representation of images and text to enhance the ability to distinguish tail class features, thereby improving classification accuracy and model robustness.
[0126] The specific steps include the following:
[0127] Step S41: constructing a graphic-text dual-branch model architecture;
[0128] This embodiment uses a pre-trained multimodal model (such as CLIP) as the basic framework, and uses an image encoder and a text encoder to extract the semantic features of the input image and category description respectively. The image encoder part remains frozen, and only a linear classifier head is added after its output for category prediction.
[0129] The dual-branch structure of pictures and text is as follows:
[0130] Image branch: extract features from the input image x and obtain the image representation vector φ img (x);
[0131] Text branch: description text t for each category c Encode and obtain the category embedding vector φ text (t c ), and normalize the embedding set of all categories.
[0132] Step S42: fusing image and text features for classification prediction;
[0133] During the inference process, the model calculates the cosine similarity between the image features and the text features of each category to form the prediction score of the text branch:
[0134] logits text (x) = cos(φ img (x),φ txt (t c ))
[0135] At the same time, the image branch passes through the linear classification head W vis Output classification score:
[0136] logits vis (x) = W vis (φ img (x))
[0137] Introduce a learnable fusion weight parameter α∈[0,1] to fuse the prediction results of the two branches as the final classification output:
[0138] logits final (x) = α·logits text (x)+(1-α)·logits vis (x)
[0139] Among them, the fusion weight parameter α is automatically optimized through back propagation during the training process to adapt to the degree of dependence of different samples on image and text information.
[0140] Step S43: jointly training the image-text fusion classification model;
[0141] Based on the original long-tail dataset and the enhanced image set, a training dataset is constructed, and a text-image fusion classification model is constructed and trained. The standard cross entropy loss is used as the training target, and only the linear classification head W is optimized. vis and fusion weight parameter α, the rest of the encoder parameters remain frozen to reduce training difficulty and prevent overfitting.
[0142] During the training process, the model automatically learns the optimal fusion method of image and text features to achieve accurate classification of all categories and improve the discrimination ability of tail category features.
[0143] Step S44: model reasoning and test evaluation;
[0144] After model training is complete, the image-text fusion classification model is tested using a test set containing images from all categories. During the testing phase, the test image is input and, through the image-text feature extraction and fusion prediction process, the final category prediction result is output. During the testing phase, the image-text fusion classification model performs a unified recognition task on all image categories, regardless of sample source.
[0145] After completing this step, a picture-text fusion recognition system for all-category classification tasks is formed. It can effectively utilize semantic descriptions and generate sample enhancement information, improve the overall discrimination ability and category balance of the model in long-tail distribution scenarios, and have good generalization performance.
[0146] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered as equivalent replacement methods and are included in the scope of protection of the present invention.
Claims
1. A long-tail image recognition method based on multimodal semantic generation and image-text fusion, characterized in that: The steps include: Obtain tail images, perform semantic modeling on the tail images based on a multimodal visual language model, and extract structured semantic descriptions; Taking the structured semantic description as the basic input, the semantics are rewritten and enhanced based on the multimodal visual language model to generate an extended semantic description of the image; The sentence vector semantic duplication detection method is used to remove redundant descriptions from the image semantic extended description. The style of the image semantic extended description after removing redundant descriptions is then aligned based on the image style prompt words to obtain the optimized text description set. The optimized text description set is input into the text-generated graph model to generate tail-category image samples, which are then screened for semantic and visual quality to construct an enhanced image set for training. A training dataset is constructed based on the original long-tail dataset and the enhanced image dataset to train the image-text fusion classification model; The image to be identified is input into the trained image-text fusion classification model, and the classification results of all categories are output.
2. The long-tail image recognition method based on multimodal semantic generation and image-text fusion according to claim 1 is characterized in that: Based on the multimodal visual language model, semantic modeling is performed on the tail image to extract structured semantic descriptions, including: A structured description template is constructed based on the typical features of tail-category images. The tail-category images are input into a multimodal visual language model. Based on the guidance of the structured description template, semantic description text of the tail-category images is generated, and formatting and grammar checking are performed to obtain a set of structured semantic description texts.
3. The long-tail image recognition method based on multimodal semantic generation and image-text fusion according to claim 1 is characterized in that: The semantic duplication detection method based on sentence vectors is used to remove redundant descriptions of image semantic extension descriptions, including: Based on the sentence vector embedding model, each image semantic extension description is mapped to a high-dimensional semantic space, and vectorized duplication detection is performed on the image semantic extension description; For each newly generated image semantic extension description e new , calculate its cosine similarity with all descriptions in the currently retained description set ε, and take the maximum value, which is specifically expressed as: yes(and new ,ε)=maxcos(φ(e new ),φ(e i )),and i ∈ε Among them, φ(·) represents the semantic vector representation of the description text, and cos(·,·) represents the cosine similarity function; Set the semantic duplication threshold τ, if sim(e new ,ε)>τ, then the description is considered to be semantically duplicated with the existing description and the description is not retained; otherwise, the description is retained.
4. The long-tail image recognition method based on multimodal semantic generation and image-text fusion according to claim 1 is characterized in that: Based on the image style clue words, the style alignment of the image semantic extended description after removing redundant descriptions is performed, specifically including: The image encoder based on the CLIP model extracts the tail class image features, performs KMeans clustering to obtain multiple style cluster centers, selects a representative tail class image for each cluster, and inputs it into the multimodal model to generate style description; The style description is spliced into the image semantic extended description in the form of phrases after removing redundant descriptions, and style alignment is performed based on the constructed image style clue words.
5. The long-tail image recognition method based on multimodal semantic generation and image-text fusion according to claim 1 is characterized in that: After obtaining the optimized text description set, the category extension description balancing step is also included, specifically including: The number of target descriptions for each category is set to a fixed value. If it is determined that the semantic extended description of a certain type of image is less than the number of target descriptions of the category, it is re-expanded. If it is determined that the semantic extended description of a certain type of image exceeds the number of target descriptions of the category, random sampling is performed according to the semantic distance distribution, and finally a set of semantic extended descriptions of each category of images is output.
6. The long-tail image recognition method based on multimodal semantic generation and image-text fusion according to claim 1 is characterized in that: Perform semantic and visual quality screening to build an enhanced image set for training, including: Based on the CLIP model, the image and text are encoded separately and the semantic consistency score is calculated, which is expressed as: s a li gn (I i ,T i )=cos(φi mg (I i ),φ text (T i )) Among them, I i represents the i-th generated tail image, T i Indicates its corresponding text description, φ img (·) and φ text (·) represents the image encoder and text encoder of the CLIP model, cos(·,·) represents the cosine similarity; The visual quality of the generated tail class images is scored based on the pre-trained aesthetic scorer, expressed as: s qua li t y (I i )=f aesthetic (I i ) Among them, I i represents the tail class image generated by the i-th image, f aesthetic represents the pre-trained aesthetic scorer; A double-threshold mechanism is used to screen the generated tail class image samples, and the generated tail class images that meet both the semantic consistency score greater than or equal to the score threshold and the visual quality score greater than or equal to the score threshold are retained.
7. The long-tail image recognition method based on multimodal semantic generation and image-text fusion according to claim 6 is characterized in that: During the steps of semantic and visual quality screening and building an enhanced image set for training, multiple rounds of closed-loop optimization are performed, specifically including: Input the tail image samples retained in the previous round of screening into the multimodal visual language model to generate a new image feature description f (t) , combined with the original semantic description e from the previous round (t) With image style prompt words k , build a new prompt template e (t+1) , expressed as: e (t+1) =Refine(e (t) ,f (t) ,s k ) Refine represents the description refinement function, which is used to integrate image content feedback and style diversity information; Set the new prompt template (t+1) Input the text graph model to generate image I (t+1) , repeatedly calculate the semantic consistency score and visual quality score, use the double threshold mechanism to filter the generated tail class image samples, and obtain a new round of retained tail class image samples; When the set iterative termination condition is met, the closed-loop process ends and the final enhanced image set is output.
8. The long-tail image recognition method based on multimodal semantic generation and image-text fusion according to claim 7 is characterized in that: The set iteration termination conditions include: meeting the set number of iterations, the number of retained images reaching the set number threshold, and the change in the pass rate of two adjacent rounds of screening being less than the set pass rate threshold.
9. The long-tail image recognition method based on multimodal semantic generation and image-text fusion according to claim 1 is characterized in that: A training dataset is constructed based on the original long-tail dataset and the enhanced image dataset to train the image-text fusion classification model, specifically including: Construct a dual-branch network of images and texts. In the image branch, extract features of the input image x based on the image encoder to obtain the image representation vector φ img (x), the description text t of each category based on the text encoder in the text branch c Encode and obtain the category embedding vector φ text (t c ), and normalize the embedding set of all categories; The cosine similarity between the image representation vector and each category embedding vector is calculated separately to form the prediction score of the text branch, which is expressed as: logits text (x)=cos(φ img (x),φ txt (t c )) The image branch passes through the linear classification head W vis Output classification score, expressed as: logits vis (x)=W vis (φ img (x)) The prediction results in the fusion image branch and the text branch are used as the final classification output, which is expressed as: logits final (x)=α·logits text (x)+(1-α)·logits vis (x) Among them, α represents the learnable fusion weight parameter; The original long-tail dataset and the enhanced image set are input into the image-text dual-branch network for training to obtain the trained image-text fusion classification model.
10. The long-tail image recognition method based on multimodal semantic generation and image-text fusion according to claim 9 is characterized in that: The original long-tail dataset and the enhanced image set are input into the image-text dual-branch network for training. The standard cross entropy loss is used as the training target, and only the linear classification head W is optimized. vis The weight parameter α is fused and the rest of the encoder parameters remain frozen.
Citation Information
Patent Citations
Long-tail target detection method based on monitoring scene
CN118097268A
Long tail distribution-oriented vision-language model prompt learning framework
CN118917276A
Fine-grained multi-mode prompt-guided visual relation identification method and device
CN119229204A
Wild animal long-tail data enhancement method combining prior knowledge and diffusion model
CN119360415A
Power transmission line detection method and apparatus, computer device, and storage medium
WO2023174020A1
Cited By
Signal text data generation and bimodal fusion continuous learning method and system under long-tail distribution
CN121145936A
Behavior recognition model training method and device, equipment, storage medium and product
CN121305266A
Behavior recognition model training methods, devices, equipment, storage media and products
CN121305266B
Multi-modal fusion and generative repair-based severe environment three-dimensional modeling method and system
CN121353547A
Visual identification method and system fusing large model and visual model, medium and product
CN121561835A