CLIP-based coarse-to-fine image-text cross-modal semantic conversion method and system
This paper proposes a cross-modal semantic transformation method that combines the CLIP model with a hyperspherical variational encoder and an interpolation module. This method addresses the challenges of fine-grained image-to-text cross-modal semantic transformation using CLIP, and achieves a coarse-to-fine image-to-text cross-modal semantic transformation method for fine-grained text features. It resolves the issues of coarse-to-fine granularity mismatch and feature space offset in existing technologies, realizing a coarse-to-fine image-to-text cross-modal semantic transformation method. This method adapts to the technical challenges of coarse-to-fine transformation in existing technologies, and achieves cross-modal alignment capabilities. It also solves the cross-domain alignment problem in existing cross-modal alignment capabilities. Furthermore, the interpolation module applies coarse-to-fine techniques, achieving the desired technical effect.
Patent Information
- Application Number
- CN202511124626.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-11-25
AI Technical Summary
Existing CLIP models suffer from semantic granularity mismatch and feature space offset issues in downstream tasks. In particular, they cannot effectively capture discriminative features that distinguish closely related subcategories in fine-grained classification tasks, and they lack the ability to refine semantics in cases with few samples.
We adopt a coarse-to-fine image-to-text cross-modal semantic transformation method based on CLIP. Through a cross-modal alignment transfer model, including the CLIP model, a hyperspherical variational encoder, and an interpolation module, we use the hyperspherical variational encoder to estimate the low-dimensional distribution of image and text features, and then use the interpolation module to interpolate in the hypersphere space to estimate the fine-grained text features of the downstream task.
It enables efficient estimation of fine-grained text features from coarse-grained text and unlabeled images when labeled real fine-grained text data is scarce, while maintaining cross-modal alignment capabilities, reducing computational costs and sample requirements, and adapting to semantic modeling for downstream tasks.
Smart Images

Figure CN121010977A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of machine learning technology, specifically to a coarse-to-fine image-text cross-modal semantic conversion method and system based on CLIP. Background Technology
[0002] Contrastive Language-Image Pre-training (CLIP) pioneered a new paradigm in contrastive learning, drawing inspiration from large-scale unsupervised pre-training methods in natural language processing. By optimizing the alignment of contrastive features, these models project image and text embeddings into a unified semantic space, facilitating cross-modal understanding and generation. CLIP avoids the reliance on meticulously crafted annotations found in traditional supervised computer vision methods, exhibiting significant transferability and zero-shot generalization. This innovation allows for direct deployment in various downstream vision tasks, such as image classification, image captioning, and object detection, while minimizing task-specific fine-tuning requirements, thus setting a new benchmark in multimodal representation learning.
[0003] The Visual Language Model (VLM) built on CLIP revolutionizes image representation learning through open-vocabulary semantic alignment. CLIP is trained on network-scale data to extract semantically centered representations, achieving powerful generalization and robustness.
[0004] Despite CLIP's inherent generalization ability and the promising results achieved in numerous studies attempting to apply this pre-trained model to few-shot transfer learning, two key challenges remain in downstream task adaptation: semantic granularity mismatch and feature space shift. The former is primarily evident in fine-grained classification tasks, where CLIP pre-training focuses on coarse-grained alignment between general concepts, failing to capture the discriminative features needed to distinguish closely related subcategories. The latter stems from significant domain differences between the distribution of pre-trained network data and downstream task-specific data; variations in image distribution lead to feature space distortion, thereby weakening the original cross-modal alignment capability.
[0005] To address overfitting in few-shot scenarios where full fine-tuning is not feasible, efficient parameter adaptation techniques for CLIP have emerged. Current methods include cue-based tuning and adapter-based approaches, such as... Figure 1(a) and (b) in the figure illustrate the simplified workflows of cue tuning and adapter methods, respectively. Cue tuning adapts to CLIP through dynamic cues, while the adapter method uses lightweight modules to project features to the target domain. Although these methods achieve significant performance on downstream tasks, they primarily mitigate the feature space shift problem through simplified variants with full parameter fine-tuning. Essentially, they still attempt to retrain on downstream tasks using a small amount of supervised data, which leads to their heavy reliance on variant design to mitigate the impact of very few samples. Furthermore, these methods lack semantic refinement and ignore the inherent interdependencies between text and image embeddings, which could provide valuable guidance in CLIP capability transfer.
[0006] In related technologies, the literature "Hyperspherical Variational Auto-Encoders, Proceedings of the Thirty-Fourth Conference on Uncertainty in Artificial Intelligence, pp. 856-865, 2018" proposes the basic theory of hyperspherical variational autoencoders, but lacks specific problem-oriented and engineering application analysis, mainly remaining at the theoretical level. Patent application CN116244645A utilizes the excellent latent space encoding capability of variational autoencoders, but its focus is mainly on how to learn new knowledge in the latent space, without addressing feature alignment between modalities. Furthermore, this scheme does not escape the trap of retraining with few samples. Patent application CN115563316A proposes a novel modality transformation mechanism that uses a decomposable attention mechanism to re-represent the feature representation of one modality with another modality, but its modal interactions are linearly sequential and do not address how different modalities interact during the encoding process. Summary of the Invention
[0007] The technical problem to be solved by this invention is how to transfer CLIP capabilities to downstream tasks to estimate the true features of downstream text when labeled, real, fine-grained text data is scarce.
[0008] The present invention solves the above-mentioned technical problems through the following technical means:
[0009] A coarse-to-fine image-text cross-modal semantic transformation method based on CLIP is proposed, the method comprising:
[0010] The paired images and coarse-grained text are acquired and input into the cross-modal alignment transfer model, which includes a CLIP model, a hyperspherical variational encoder, and an interpolation module connected in sequence.
[0011] The CLIP model is used to encode paired images and coarse-grained text to obtain downstream image features and pre-trained text features;
[0012] The low-dimensional distributions of downstream image features and pre-trained text features are estimated using a hyperspherical variational encoder, resulting in the hyperspherical latent distributions of downstream image features and pre-trained text features.
[0013] The hypersphere space interpolation is performed using the interpolation module based on the hypersphere latent distribution of downstream image features and the hypersphere latent distribution of pre-trained text features to obtain the estimated distribution of fine-grained text features for the downstream task.
[0014] Furthermore, the hyperspherical variational encoder includes a parallel text hyperspherical variational encoder and a visual hyperspherical variational encoder;
[0015] The text hyperspherical variational encoder consists of a structurally symmetric text encoder and a text decoder;
[0016] The visual hyperspherical variational encoder consists of a structurally symmetrical visual encoder and a visual decoder.
[0017] Furthermore, the interpolation module includes a reparameterized structure and a text decoder, wherein:
[0018] The hypersphere space interpolation of the hypersphere latent distribution of downstream image features and the hypersphere latent distribution of pre-trained text features is performed using the reparameterization structure to obtain the probability density estimate of the interpolated fine-grained text.
[0019] The probability density estimate of the interpolated fine-grained text is decoded using a text decoder to obtain the fine-grained text feature distribution for downstream tasks.
[0020] Furthermore, the hypersphere interpolation process includes:
[0021] The hyperspherical latent distribution of downstream image features and the hyperspherical latent distribution of pre-trained text features are regarded as position points on the hypersphere, and the fine-grained text features to be estimated are modeled as a single-peaked vMF probability density distribution on the hypersphere. The probability density distribution consists of mean and concentration parameters.
[0022] Moving at a constant speed along the arc connecting the two location points to perform geodesic interpolation between the two location points, the probability density of the interpolation point with the smallest difference from the true fine-grained text feature distribution is used as the probability density estimate of the interpolated fine-grained text.
[0023] Further, the step of moving at a constant speed along the arc connecting the two position points to perform geodesic interpolation between the two position points, and obtaining the probability density of the interpolation point with the smallest difference from the true fine-grained text feature distribution as the probability density estimate of the interpolated fine-grained text, includes:
[0024] Spherical linear interpolation is used for the mean, and the formula is expressed as:
[0025]
[0026] In the formula: Let be the mean of the fine-grained text feature distribution to be estimated, θ be the angular distance, and t∈[0,1] be a learnable interpolation coefficient. The mean of the feature distribution of the pre-trained text. This represents the mean of the downstream image feature distribution.
[0027] Logarithmic interpolation of the concentration parameter is expressed by the following formula:
[0028]
[0029] In the formula: The concentration parameter is the fine-grained text feature distribution to be estimated. The concentration parameter of the pre-trained text feature distribution. The concentration parameter is used to represent the downstream image feature distribution.
[0030] The probability density estimate of the interpolated fine-grained text is as follows:
[0031] Furthermore, before acquiring the paired images and coarse-grained text and inputting them into the cross-modal alignment transfer model, the method further includes:
[0032] The cross-modal alignment transfer model is trained in two stages: unsupervised training is used for the hyperspherical variational encoder, and supervised training is used for the interpolation module under the supervision of real fine-grained sample features.
[0033] Furthermore, the total loss function used during unsupervised training of the hyperspherical variational encoder is... for:
[0034]
[0035] In the formula: Let be the loss function of the visual hyperspherical variational encoder. Let be the loss function of the text hyperspherical variational encoder. For the joint reconstruction loss function, f i The downstream image features output by the CLIP model, These are coarse-grained text features after cross-decoding by the image decoder. The pre-trained text features output by the CLIP model. These are the image features after cross-decoding by the text decoder. Let be the mathematical expectation of the spatial distance.
[0036] Furthermore, the loss function used during supervised training of the interpolation module... for:
[0037]
[0038] In the formula: This represents the estimated distribution of fine-grained text features for a downstream task obtained through the text encoder in the hyperspherical variational encoder. This represents the fine-grained text latent space features obtained through interpolation estimation. Represents true, fine-grained text features. This represents the estimated fine-grained textual latent space distribution. This represents the prior distribution of fine-grained text in the latent space. This represents the estimated fine-grained text features obtained after decoding, and KL(||) represents the KL divergence.
[0039] Furthermore, this invention proposes a CLIP-based coarse-to-fine image-text cross-modal semantic conversion system, the system comprising:
[0040] An acquisition module is used to acquire paired images and coarse-grained text and input them into a processing module. The processing module is equipped with a pre-trained cross-modal alignment transfer model, which includes a CLIP model, a hyperspherical variational encoder, and an interpolation module connected in sequence.
[0041] The CLIP model is used to encode paired images and coarse-grained text to obtain downstream image features and pre-trained text features.
[0042] The hyperspherical variational encoder is used to estimate the low-dimensional distribution of downstream image features and pre-trained text features, and obtain the hyperspherical latent distribution of downstream image features and the hyperspherical latent distribution of pre-trained text features.
[0043] The interpolation module is used to perform hypersphere space interpolation based on the hypersphere latent distribution of downstream image features and the hypersphere latent distribution of pre-trained text features to obtain the estimated distribution of fine-grained text features for the downstream task.
[0044] Furthermore, the present invention also proposes a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the CLIP-based coarse-to-fine image-text cross-modal semantic conversion method as described above.
[0045] The advantages of this invention are:
[0046] (1) Unlike traditional methods that directly learn transferability from a small number of downstream samples, this invention does not directly train feature extraction capabilities for downstream tasks. Instead, it focuses on learning the transformation between images and upstream and downstream texts and applying this transformation to downstream samples to achieve transferability. Specifically, it leverages the cross-modal alignment feature already achieved by the CLIP model on massive coarse-grained texts. Through the interpolation module, it connects the pre-trained coarse-grained knowledge of the CLIP model with the downstream fine-grained requirements, and transfers the original feature extraction and modal alignment capabilities of the CLIP model to downstream tasks. It estimates fine-grained text features from coarse-grained texts and unlabeled images, and realizes semantic modeling from coarse to fine granularity. This "holistic learning paradigm" proposed in this invention is more suitable for estimating the real features of downstream texts when labeled real fine-grained text data is scarce.
[0047] (2) The geodesic projection-based interpolation mechanism introduced in this invention can ensure that the interpolation points retain the geometry of the hypersphere and will not collapse into the lower region near the center, thus avoiding the introduction of nonlinear distortion and error accumulation, and ensuring a smooth transition of semantics through continuous parameter evolution.
[0048] (3) By operating under the inherent geometric constraints of the CLIP model, the present invention designs parallel text hyperspherical variational encoders and visual hyperspherical variational encoders, which are better aligned with the original distribution of features extracted by the CLIP model. During model training, a cross-decoding scheme is adopted for the text hyperspherical variational encoder and the visual hyperspherical variational encoder to maintain the alignment between modes, effectively preserving the inherent cross-modal alignment capability of the CLIP model.
[0049] (4) In the training process of the interpolation module, the present invention only needs to optimize a minimum number of learnable parameters, and the computational cost and sample requirements are extremely low.
[0050] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0051] Figure 1This is a schematic diagram comparing the implementation process of the CLIP-based coarse-to-fine image-to-text cross-modal semantic conversion method proposed in this invention with other traditional methods. In this diagram, (a) shows the use of learnable cues to guide the model for cue tuning, (b) shows the use of an adapter-based method to insert an additional adapter after the encoder output to achieve task adaptation, and (c) shows the use of a variational autoencoder to estimate the latent features of the downstream text in the method proposed in this invention.
[0052] Figure 2 This is a schematic diagram comparing the principles of a conventional method and the method proposed in this invention in one embodiment of the invention;
[0053] Figure 3 This is a flowchart illustrating a CLIP-based coarse-to-fine image-text cross-modal semantic conversion method proposed in one embodiment of the present invention.
[0054] Figure 4 This is a schematic diagram of the structure of a cross-modal alignment transfer model in one embodiment of the present invention;
[0055] Figure 5 This is a schematic diagram of the interpolation module in one embodiment of the present invention;
[0056] Figure 6 This is a schematic diagram illustrating the principle of a cross-modal alignment transfer model in one embodiment of the present invention;
[0057] Figure 7 This is a schematic diagram illustrating the principle of interpolation module performing potential hypersphere space interpolation in one embodiment of the present invention;
[0058] Figure 8 This is a schematic diagram of the structure of a CLIP-based coarse-to-fine image-to-text cross-modal semantic conversion system proposed in one embodiment of the present invention;
[0059] Figure 9 This is a schematic diagram comparing the performance of HIVE with zero-sample CLIP in one embodiment of the present invention;
[0060] Figure 10 This is a schematic diagram comparing the accuracy of HIVE with other baseline methods on the AverageAcross All Datasets dataset during a fine-grained image classification experiment in one embodiment of the present invention.
[0061] Figure 11 This is a schematic diagram comparing the accuracy of HIVE with other baseline methods on the OxfordPets dataset during a fine-grained image classification experiment in one embodiment of the present invention.
[0062] Figure 12 This is a schematic diagram comparing the accuracy of HIVE with other baseline methods on the StanfordCars dataset during a fine-grained image classification experiment in one embodiment of the present invention.
[0063] Figure 13 This is a schematic diagram comparing the accuracy of HIVE with other baseline methods on the Flowers102 dataset during a fine-grained image classification experiment in one embodiment of the present invention.
[0064] Figure 14 This is a schematic diagram comparing the accuracy of HIVE with other baseline methods on the Food101 dataset during a fine-grained image classification experiment in one embodiment of the present invention.
[0065] Figure 15 This is a schematic diagram comparing the accuracy of HIVE with other baseline methods on the DTD dataset during a fine-grained image classification experiment in one embodiment of the present invention.
[0066] Figure 16 This is a schematic diagram comparing the accuracy of HIVE with other baseline methods on the EuroSAT dataset during a fine-grained image classification experiment in one embodiment of the present invention.
[0067] Figure 17 This is a schematic diagram comparing the accuracy of HIVE with other baseline methods on the FGCVAircraft dataset during a fine-grained image classification experiment in one embodiment of the present invention.
[0068] Figure 18 This is a schematic diagram comparing the accuracy of HIVE with other baseline methods on the CUB dataset during a fine-grained image classification experiment in one embodiment of the present invention.
[0069] Figure 19 This is a schematic diagram comparing the accuracy of HIVE and baseline methods on three general classification datasets during a non-fine-grained image classification experiment in one embodiment of the present invention.
[0070] Figure 20 This is a schematic diagram comparing the performance of the original HIVE and its variant that replaces SLERP with MLP in one embodiment of the present invention;
[0071] Figure 21 This is a schematic diagram comparing the average accuracy of HIVE and the original CLIP on all eight fine-grained datasets under different visual backbone networks in one embodiment of the present invention. Detailed Implementation
[0072] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0073] It should be noted that, as Figure 2 As shown, to address the lack of semantic refinement in traditional methods, this invention explores the issue from a structural perspective, studying the inherent interdependence between text and image embeddings from a spatial structure standpoint. It assumes a specific transformation relationship exists between the directional distributions of image-text features; more specifically, it proposes a transformation... Connecting three key elements: downstream image features i, pre-trained text features t coarse And fine-grained text features t that exhibit fine-grained semantic alignment with downstream images fine This can be expressed as:
[0074]
[0075] This transformation By using i and t coarse Get t from fine To achieve semantic alignment.
[0076] To further provide an implementable solution, this embodiment will... The goal is to construct a system that is geometrically intuitive and effectively handles directional data. One feasible approach is to use a Variational Autoencoder (VAE), an unsupervised generative model that implements variational inference through an encoder-decoder architecture. Its core objective is to generate new samples by learning the latent distribution of the data. The encoder assumes that the latent variables follow a prior distribution and maps the input data to a variational distribution to approximate the posterior distribution, while the decoder reconstructs the data from the latent variables.
[0077] First, let i and t coarse The high-dimensional representation is reduced to a low-dimensional latent space, then modeled in this compressed domain, and finally the learned representation is reconstructed back to the original high-dimensional space. Therefore, the above transformation can be... Rephrased as:
[0078]
[0079] Where ε and Let t represent the encoder and decoder components of the variational autoencoder, respectively, and h represent the fine-grained text features. fine Potential variables.
[0080] To achieve the above-mentioned inventive concept, such as Figure 3 , Figure 4 , Figure 5 and Figure 6As shown, one embodiment of the present invention proposes a coarse-to-fine image-text cross-modal semantic transfer method based on CLIP, which specifically achieves the above-mentioned transfer by modeling a cross-modal alignment transfer model. The method includes the following steps:
[0081] S10. Obtain the paired image and coarse-grained text and input them into the cross-modal alignment transfer model, wherein the cross-modal alignment transfer model includes a CLIP model, a hyperspherical variational encoder and an interpolation module connected in sequence;
[0082] S20. Use the CLIP model to encode the paired images and coarse-grained text to obtain downstream image features and pre-trained text features;
[0083] It should be noted that the downstream image features mentioned in this embodiment are images that can be used for downstream tasks. The CLIP model used has the characteristic of achieving cross-modal alignment on massive coarse-grained text. The CLIP model is used to align the inherent structure of coarse-grained text features and image features. Moreover, the CLIP model is pre-trained and does not need to be trained together with the hyperspherical variational encoder and interpolation module, thus reducing training costs.
[0084] S30. Use a hyperspherical variational encoder to estimate the low-dimensional distribution of downstream image features and pre-trained text features, and obtain the hyperspherical latent distribution of downstream image features and the hyperspherical latent distribution of pre-trained text features.
[0085] It should be noted that the hyperspherical variational encoder is used to estimate the distribution of downstream image features and pre-trained text features on the hypersphere. In this embodiment, the von Mises-Fisher (vMF) distribution can be specifically selected for variational inference, which matches the original feature characteristics of the CLIP model. Then, the transformation is achieved by geodesic projection based on the hyperspherical distribution estimated by the hyperspherical variational encoder.
[0086] S40. Using the interpolation module, hypersphere space interpolation is performed based on the hypersphere latent distribution of downstream image features and the hypersphere latent distribution of pre-trained text features to obtain the estimated distribution of fine-grained text features for the downstream task.
[0087] It should be noted that in this embodiment, fine-grained text features are represented as spatial midpoints on a hypersphere, located between coarse-grained text features and image features. This transfers the original alignment relationship between image features and coarse-grained text features to the relationship between image features and fine-grained text features. This geometry-based method estimates fine-grained text features from coarse-grained text and unlabeled images by transferring the modal alignment capability of the CLIP model. In this embodiment, an interpolation module is specifically used to find this spatial midpoint.
[0088] This embodiment leverages the CLIP model's cross-modal alignment capability on massive amounts of coarse-grained text. An interpolation module connects the pre-trained coarse-grained knowledge of the CLIP model with downstream fine-grained requirements, learning this structural representation from labeled fine-grained text data. The aim is to naturally transfer the CLIP model's original feature extraction and modal alignment capabilities to downstream target tasks while maintaining structural consistency. Fine-grained text features are estimated from coarse-grained text and unlabeled images, enabling semantic modeling from coarse to fine granularity. This is suitable for estimating the true features of downstream text when labeled, real fine-grained text data is scarce.
[0089] As a further preferred technical solution, in the CLIP model, images and text are respectively processed by an image encoder. and text encoder For processing, these two encoders typically use ResNet or Transformer networks to extract feature representations.
[0090] It should be noted that during the inference process, given a predefined cue template, such as "a photo of [CLS]", all K candidate category labels are replaced with the [CLS] position. These text cuees are then processed by a text encoder. Encode the data to generate feature vectors for a specific category. The probability that an input image x belongs to a specific category k can be obtained as follows:
[0091]
[0092] Among them, t i The text description represents category i, and τ is the temperature parameter.
[0093] It should be noted that in order to use cos(,) to calculate the similarity between text and image features, the features obtained from the encoder during training must be subjected to L... 2 Normalization restricts the coding features of CLIP to a hypersphere. Above, where D represents the embedding dimension.
[0094] To further investigate the distribution characteristics of these eigenvectors on the hypersphere, this embodiment uses a Gaussian-constrained variational autoencoder. and von Mises-Fischer (vMF) constrained variational autoencoders that act explicitly on hyperspherical manifolds An encoder-decoder architecture ε is trained using CLIP encoded vectors. After training, the performance of the two prior-constrained encoders is qualitatively evaluated using reconstruction loss and log-likelihood. Reconstruction loss, one of the optimization metrics during variational autoencoder training, measures the difference between the original and reconstructed features, and its specific form is as follows:
[0095]
[0096] Where h represents the latent variable sampled from the estimated distribution q(h|x) under the constraint of the prior distribution p(x|h). For continuous data, the mean squared error (MSE) is typically used:
[0097]
[0098] in, This refers to the features reconstructed by the decoder.
[0099] Log-likelihood represents the probability of observing a data distribution given a sample of latent variables and their distributions. This example uses the k-sample marginal log-likelihood:
[0100]
[0101] To maintain generality, experiments were conducted on four datasets: Flowers102, StanfordCars, Food101, and Caltech101, to evaluate the encoded representations of images and text. Table 1 presents the results, showing that using... The reconstruction loss is significantly lower than In terms of log-likelihood performance It is significantly superior in low-potential spatial dimensions. Furthermore, when the assumed feature distribution lies on a hypersphere, the reconstruction loss and probability estimation error change more smoothly across different dimensions. Compared to the Gaussian distribution assumption, the hypersphere distribution exhibits better robustness while maintaining consistent performance across different embedding dimensions. Therefore, this embodiment specifically selects the von Mises-Fischer (vMF) distribution for variational inference, which matches the original feature characteristics of the CLIP model.
[0102] Table 1. Reconstruction loss and log-likelihood performance of two variational autoencoders on four selected datasets (average of 10 runs).
[0103]
[0104] As a further preferred technical solution, such as Figure 4 As shown, the hyperspherical variational encoder (HVA) includes a parallel text hyperspherical variational encoder and a visual hyperspherical variational encoder.
[0105] The text hyperspherical variational encoder consists of a structurally symmetric text encoder and a text decoder;
[0106] The visual hyperspherical variational encoder consists of a structurally symmetrical visual encoder and a visual decoder.
[0107] It should be noted that, operating under the inherent geometric constraints of the CLIP model, parallel text hyperspherical variational encoders and visual hyperspherical variational encoders were designed to better align with the original distribution of features extracted by the CLIP model.
[0108] Furthermore, this embodiment employs a hyperspherical variational encoder (HVA) to estimate the low-dimensional distribution of features extracted by the CLIP model, assuming f i and These represent the downstream image features (i.e., the original visual features used in the downstream task) and the pre-trained text features (i.e., the original coarse-grained text features used in the downstream task) encoded by the CLIP model, respectively. The goal of the Hyperspherical Variational Encoder (HVA) is to obtain an approximate latent distribution of the downstream image features. Latent distributions that approximate the features of pre-trained text
[0109] The HVA architecture includes a parallel text hyperspherical variational encoder (text encoder-text decoder). ) and visual hyperspherical variational encoder (visual encoder-visual decoder) Taking image branching as an example, given downstream image features f i Assume its latent distribution is governed by the prior vMF distribution p(f i |h i The constraint is defined by (μ, κ), where μ represents the distribution mean and k represents the distribution concentration. The visual encoder ε... i Composed of multiple noisy linear projection layers, the downstream image features f are first... i Projected into a low-dimensional space. Then, the estimator... and estimator Approximate the posterior vMF distribution q(h) i |f i mean and concentration parameters By reparameterizing the structure, from q(h) i |f i The potential code h is obtained by sampling in ) i As output of the visual encoder. Visual decoder. In terms of architecture, it is similar to the visual encoder ε i Symmetry, from h i Downstream image features f reconstructed in the middle i Reconstruction features
[0110] During the training process of the hyperspherical variational encoder, the variational autoencoder... The Evidence Lower Bound (ELBO) is used as the optimization objective to maximize the marginal likelihood p:
[0111]
[0112] in, For the loss of the visual hyperspherical variational encoder, h i Represents downstream image features f i Potential code, p(h) i ) represents the prior distribution of image features in the latent space; This indicates reconstruction error, prompting the visual decoder Better refactoring f i , while KL(q(h) i |f i )||p(h i )) represents the approximate posterior vMF distribution q(h) i |f i ) and prior p(h i The KL divergence between the two sides ensures the continuity of the latent space and the consistency between the estimated distribution and the prior.
[0113] To ensure that the VAE does not violate the inherent image-text alignment of CLIP when estimating latent variables, the Wasserstein distance between the two modality estimation distributions is optimized to maintain modality alignment. Since the Wasserstein distance between the vMF distributions does not have a closed-form analytical solution, approximating it using the Sinckhorn distance of the sampling distributions increases computational overhead. To address this issue, this embodiment employs a cross-decoding scheme during training to maintain intermodality alignment. Specifically, a visual decoder is used. To reconstruct pre-trained text features Latent space variables Text decoder Reconstruct f i The latent space variable h i The joint reconstruction losses are calculated as follows:
[0114]
[0115] in, For visual decoders to latent code The reconstruction results For text decoders to the potential code h i The reconstruction results Let be the mathematical expectation of the spatial distance.
[0116] By combining ELBO optimization with modal alignment, the final objective function of HVA is expressed as:
[0117]
[0118] in, The loss of the text hyperspherical variational encoder is formally symmetrical to the loss calculation formula of the visual hyperspherical variational encoder as follows:
[0119]
[0120] The first and second terms represent the reconstruction error and KL divergence corresponding to the text branches.
[0121] This embodiment utilizes a trained HVA architecture to ultimately obtain the estimated distribution in a low-dimensional latent space. and Because a cross-decoding scheme is used for the text hyperspherical variational encoder and the visual hyperspherical variational encoder during model training to maintain the alignment between modalities, the inherent cross-modal alignment capability of the CLIP model is effectively preserved.
[0122] As a further preferred technical solution, such as Figure 5 As shown, the interpolation module includes a reparameterized structure and a text decoder, wherein:
[0123] The hypersphere space interpolation of the hypersphere latent distribution of downstream image features and the hypersphere latent distribution of pre-trained text features is performed using the reparameterization structure to obtain the probability density estimate of the interpolated fine-grained text.
[0124] The probability density estimate of the interpolated fine-grained text is decoded using a text decoder to obtain the fine-grained text feature distribution for downstream tasks.
[0125] As a further preferred technical solution, the hypersphere space interpolation process includes:
[0126] The hyperspherical latent distribution of downstream image features and the hyperspherical latent distribution of pre-trained text features are regarded as position points on the hypersphere, and the fine-grained text features to be estimated are modeled as a single-peaked vMF probability density distribution on the hypersphere. The probability density distribution is composed of mean and concentration parameterization.
[0127] Moving at a constant speed along the arc connecting the two location points to perform geodesic interpolation between the two location points, the probability density of the interpolation point with the smallest difference from the true fine-grained text feature distribution is used as the probability density estimate of the interpolated fine-grained text.
[0128] It should be noted that the true fine-grained text feature distribution here refers to the small amount of true labeled fine-grained data f obtained after the text encoder. t fine The probability density.
[0129] Specifically, this embodiment describes the approximate distribution q of given downstream image features and pre-trained text features. i and Fine-grained text features for downstream tasks are derived using hypersphere interpolation. Fitted distribution In n-dimensional Euclidean space In this context, a unit hypersphere is defined as:
[0130]
[0131] Where n represents the variable of the dimension of Euclidean space, Let x represent the hypersphere, and let x represent the image. Let ||||2 represent the set of real numbers with dimension n+1, and ||||2 represent the modulus.
[0132] like Figure 7 As shown, given two vMF distributions and Where μ represents the probability density f(x; μ, k) = C n (κ)exp(κμ T For intuitive visualization, this embodiment uses the mean of two vMF distributions q(x). i and Consider two points A and B on a sphere, and find a point between A and B through interpolation. Figure 7 The blue dot in the image represents the location where the probability density differs least from the actual downstream fine-grained text distribution.
[0133] To achieve this density fit, the interpolation process in this embodiment must satisfy: 1) unimodality, because the subsequent estimation is based on the assumption that both coarse-grained and fine-grained text embeddings are modeled as unimodal vMF distributions; 2) geodesic interpolation, in Figure 7 The movement from point A to point B must proceed along the connected great circle at a constant speed. The former ensures that the interpolation point retains the hyperspherical geometry and does not collapse into the lower region near the center, as this would introduce nonlinear distortion and error accumulation, while the latter guarantees a smooth semantic transition through continuous parameter evolution.
[0134] As a further preferred option, to satisfy the above interpolation conditions, this embodiment uses spherical linear interpolation (SLERP) for the mean to ensure that the interpolation points are strictly located on the hypersphere. The formula is expressed as:
[0135]
[0136] In the formula: Let be the mean of the fine-grained text feature distribution to be estimated, θ be the angular distance, and t∈[0,1] be a learnable interpolation coefficient. The mean of the feature distribution of the pre-trained text. This represents the mean of the downstream image feature distribution.
[0137] Logarithmic interpolation of the concentration parameter is expressed by the following formula:
[0138]
[0139] In the formula: The concentration parameter is the fine-grained text feature distribution to be estimated. The concentration parameter of the pre-trained text feature distribution. The concentration parameter is used to represent the downstream image feature distribution.
[0140] The probability density estimate of the interpolated fine-grained text is as follows:
[0141] As a further preferred technical solution, the interpolation module (Latent Hyperspherical Space Interpolation, LHSI) needs to be trained in a supervised manner. In this embodiment, the fitting relationship is specifically learned by calculating the reconstruction loss and KL divergence, as shown below:
[0142]
[0143] in, This indicates the text encoder ε in HVA. t Fine-grained text of the downstream task obtained The estimated distribution, This represents the estimated fine-grained textual latent variables. This represents the estimated fine-grained distribution of textual latent variables. This represents the prior distribution of fine-grained textual latent variables. This represents the fine-grained text estimation features obtained after decoding by the decoder. express and KL dispersion between; first term The second term is used to measure the difference in distributions between the interpolation distributions q and p. Corresponding to using interpolation distribution pairs f i fine The loss incurred during reconstruction. Here, ε t and It has been pre-trained and already possesses powerful distribution estimation capabilities.
[0144] It should be noted that in the training process of the LHSI interpolation module in this embodiment, only a minimum number of learnable parameters need to be optimized, and the computational cost and sample requirements are extremely low, so that the limited samples in downstream few-sample tasks are sufficient to train the LHSI module.
[0145] Unlike traditional methods that focus on training classifiers using labeled samples, this embodiment does not prioritize the classification performance of the downstream task itself. Instead, it utilizes a large number of coarse-grained unlabeled samples to train the hyperspherical variational encoder, while using a small number of fine-grained labeled samples to train the interpolation module. Subsequently, alignment features from the upstream domain are naturally adapted to the downstream domain. Given the extremely limited sample size, training the low-parameter transformation module is more reasonable than optimizing the high-parameter classification component.
[0146] It should be noted that the fine-grained text features obtained in this embodiment are used for downstream tasks. Semantic alignment from coarse to fine granularity is achieved, thus resolving the granularity mismatch problem. Modal alignment with downstream images is also achieved, guaranteed by HVA's modal alignment preservation capability. Therefore, fine-grained text features... It can be used for various downstream tasks, such as few-shot image classification, semantic segmentation, and action recall.
[0147] The following example uses a few-shot image classification problem, with a support set. and a query set This contains N×K labeled images, where N represents the number of classes and K represents the number of samples in the few-shot learning setting. For the support set... Each sample in the dataset, except for a small number of downstream fine-grained text labels y fine In addition, a coarse-grained label y is assigned to a specific task. coarse This is typically determined by a broader category in the classification task (e.g., "flower" in the Flowers102 dataset). Using y coarse and y fine Replace the [CLS] tag in the prompt template, for example, "a [CLS] photo", and use the pre-trained CLIP model to obtain the corresponding coarse-grained text features. and fine-grained text features For each category of images Image features f are obtained using a pre-trained CLIP model. i . use and f i Based on steps S30 to S40 above, an approximate fine-grained text embedding for category n is derived. f t fineThis serves as a supervisor for the process. Furthermore, downstream image features in HVA reconstruction... Image features f output by the original CLIP model i Apply residual joins between them:
[0148]
[0149] This helps preserve the prior knowledge of the original model and mitigate catastrophic forgetting, where α is an adjustable hyperparameter.
[0150] For query set The samples in the data, with fine-grained labels y fine Not available. This embodiment still follows the same process to obtain f. i and Fit in the same way during training Finally, for each image According to the above formula Get its corresponding Let α represent the target image features, and α represent the residual parameter. Then, its probability is calculated according to the CLIP framework, as follows:
[0151]
[0152] Here, τ is the temperature parameter in the original CLIP model.
[0153] The loss in the classification process can be expressed as:
[0154]
[0155] Where, n gt Let x represent the true label corresponding to the input image x, and B represent the batch size.
[0156] In typical image classification tasks, this embodiment uses a meta-category such as "object" as the y-value. coarse The subsequent steps are consistent with those for fine-grained image classification. One limitation of using meta-categories is that the semantic information conveyed by coarse-grained text descriptions may be insufficient. However, subsequent experiments show that this approach does not significantly affect the performance of the method in this embodiment, which is still able to learn transformations with sufficiently fine-grained semantic representations.
[0157] In addition, such as Figure 8 As shown, another embodiment of the present invention proposes a CLIP-based coarse-to-fine image-to-text cross-modal semantic conversion system, comprising:
[0158] The acquisition module 10 is used to acquire paired images and coarse-grained text and input them into the processing module 20. The processing module 20 is equipped with a pre-trained cross-modal alignment transfer model, which includes a CLIP model, a hyperspherical variational encoder and an interpolation module connected in sequence.
[0159] The CLIP model is used to encode paired images and coarse-grained text to obtain downstream image features and pre-trained text features.
[0160] The hyperspherical variational encoder is used to estimate the low-dimensional distribution of downstream image features and pre-trained text features, and obtain the hyperspherical latent distribution of downstream image features and the hyperspherical latent distribution of pre-trained text features.
[0161] The interpolation module is used to perform hypersphere space interpolation based on the hypersphere latent distribution of downstream image features and the hypersphere latent distribution of pre-trained text features to obtain the estimated distribution of fine-grained text features for the downstream task.
[0162] Furthermore, another embodiment of the present invention proposes a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the CLIP-based coarse-to-fine image-text cross-modal semantic conversion method as described in the first embodiment above.
[0163] It should be noted that other embodiments or specific implementation methods of the CLIP-based coarse-to-fine image-text cross-modal semantic conversion system described in this invention can refer to the above-mentioned method embodiments, and will not be repeated here.
[0164] The following describes the experimental verification of the method proposed in this invention:
[0165] (1) Fine-grained image classification
[0166] Datasets: This embodiment selected eight image classification datasets with different coarse-grained categories, as shown in Table 2, and assigned appropriate coarse-level category labels to each dataset. Statistical information and category label details for these datasets are shown in Table 2. These datasets cover various classification targets, including concrete objects, scenes, actions, satellite images, and textures. Following common practice, 1, 2, 4, 8, and 16 samples were configured for few-shot learning. All training samples were randomly drawn from the original training set; the remaining training samples were not used in the training process to prevent potential distribution leakage. Unlike methods that change the cue template based on the dataset, HIVE uses a uniform cue template: “A photo of [CLS]”. The Hyperspherical Interpolation Variational Encoding (HIVE) consists of a Hyperspherical Variational Adapter (HVA) and a Latent Hyperspherical Spatial Interpolation (LHSI) mechanism.
[0167] Table 2 Fine-grained classification dataset
[0168] Dataset category Label train test Oxford Pets dataset 37 pet 2,944 3,669 StanfordCars dataset 196 car 6,509 8,041 A dataset of 102 flower species (Flowers102) 102 flowers 4,093 2,463 Food101 dataset containing 101 food items. 101 food 50,500 30,300 Texture Dataset (DTD) 47 Texture 2,820 1,692 EuroSAT (European Satellite Imagery Data Set) 10 satellite photos 13,500 8,100 FGVC Aircraft Dataset 100 airplane 3,334 3,333 The University of California, San Diego Bird Dataset (CUB) 200 birds 5,994 5,794
[0169] Baseline Method: This embodiment compares the results of HIVE with CoOp, MapLe, LAMM, and zero-sample CLIP. For fine-tuning methods, parameter settings and cue templates from the prior art are used. For the HIVE method in this embodiment, a fixed cue template is used throughout the experiment.
[0170] Implementation Details: In the variational autoencoder (VAE) architecture of this invention, both the encoder and decoder employ a two-layer multilayer perceptron (MLP) with a latent dimension of 5. The training process consists of two phases: First, the hyperspherical variational encoder is optimized using the AdamW optimizer with a learning rate of 0.0001 and a batch size of 1; subsequently, the complete model is trained with a batch size of 64. For settings with 8 and 16 samples, training is performed for 100 epochs; for other sample sizes, 50 epochs are used with a stochastic gradient descent (SGD) optimizer and a learning rate of 0.002. Image preprocessing strictly follows the standard augmentation protocol of the CLIP model, including random cropping, resizing, and random horizontal flipping. All input images are center-cropped to maintain a short side size of 224 pixels while preserving the original aspect ratio. The entire training process is performed on a single NVIDIA A800 GPU. Except for coarse text label processing, the hyperparameters of all experiments remain consistent.
[0171] Fine-grained classification results: Figure 9This demonstrates the performance improvement of HIVE over zero-sample CLIP results. The proposed method achieves a significant performance improvement using only one sample, with significant gains of 40.5% and 46.2% on the EuroSAT and CUB datasets, respectively. When expanded to a 16-sample setting, significant accuracy improvements were observed on the Stanford Cars, DTD, and FGCVAircraft datasets compared to the one-sample results, indicating that the proposed method can learn more effective semantic representations even with a slight increase in labeled data. Notably, on the Oxford Pets dataset, the original zero-sample baseline already achieved a high accuracy (89.03%), and the relative percentage improvement of HIVE appears less significant. Specifically, regarding baseline comparisons… Figures 10 to 18 The accuracy comparison between the present invention's HIVE method and the baseline method is summarized. For example... Figures 10 to 18 As shown, HIVE achieves superior average performance across all datasets compared to baseline methods. Specifically, it achieves an improvement of approximately 2.4% in the 16-sample setting, which initially demonstrates HIVE's effectiveness in fine-grained classification tasks. Notably, the average accuracy improvement is approximately 3.1% in the 1-sample case, indicating that HIVE can more effectively utilize limited labeled samples for semantic information modeling even in situations with extremely scarce data.
[0172] When examining various datasets, HIVE outperformed other methods on datasets exhibiting a clear gradient from coarse to fine semantics. In the 16-shot setting, compared to MaPLe, HIVE achieved accuracy improvements of +1.50, +4.39, +1.15, and +0.03% on the OxfordPets, StanfordCars, Flowers102, and Food101 datasets, respectively. Compared to LAMM, the corresponding accuracy improvements were +1.81, +1.94, +1.65, and +0.69%, respectively. However, HIVE's performance advantage on other datasets was less pronounced, although it still outperformed other alternatives overall. A notable exception was on the DTD dataset, where, in the 16-shot setting, HIVE showed a 1.60% decrease in accuracy compared to MaPLe and a 0.88% decrease compared to LAMM. This reveals a limitation of our method: since DTD is a textured image classification dataset, HIVE may struggle to learn discriminative semantic representations from such visual patterns effectively, leading to performance degradation.
[0173] In addition, the performance of HIVE with additional visual adapters (VA) and text adapters (TA) was compared with that of HIVE alone (average accuracy across the eight datasets in Table 2).
[0174] HIVE with an additional adapter: The HIVE framework of this invention can be integrated with a two-branch adapter architecture similar to CLIP-Adapter. Specifically, in the downstream image features of HVA reconstruction... and original CLIP image features f i Before applying residual connections, the reconstructed image features are... Fine-grained text features obtained after interpolation The inputs are fed into additional adapter modules. The outputs of these adapters are then combined with the original CLIP output via residual connections to form the final features used for cosine similarity calculation. Experiments were conducted in this embodiment to evaluate the effects of adding additional visual and textual branches, respectively; the experimental results are detailed in Table 3.
[0175] Table 3 Performance Comparison
[0176] Method 1-shot 2-shot 4-shot 8-shot 16-shot HIVE 69.04 72.46 76.26 79.83 82.53 HIVE+ Visual Adapter 68.76 73.11 77.46 81.90 84.24 HIVE+ Text Adapter 67.42 72.38 77.19 79.05 81.87 HIVE+ Visual and Text Adapter 68.53 72.98 77.42 80.73 84.96
[0177] It can be observed that in cases with extremely low data volumes (single samples), the baseline configuration without any additional adapters performs best. As the number of samples increases, the configuration using only the visual adapter outperforms other variants. When the number of samples reaches 16, the dual-adaptor configuration produces the best results. Notably, the configuration using only the text adapter is even worse than the original HIVE baseline. This paper argues that the visual adapter effectively optimizes image feature representations, while text features in HIVE are obtained through geometry-based modeling. In low-sample scenarios, the text adapter often disrupts these established geometric relationships. Only when there are sufficient samples can the text adapter effectively improve performance like the visual adapter. Furthermore, the dual-adaptor implementation requires longer training epochs and introduces approximately 30% more parameters. Considering the small performance improvement, this computational overhead is considered disproportionate to the performance gains.
[0178] (2) Non-fine-grained image classification
[0179] Since the core idea of the HIVE architecture proposed in this embodiment is to model fine-grained text using coarse-grained text and images, this embodiment also studies how the method of the present invention performs when the target downstream task lacks a clear semantic relationship from coarse to fine.
[0180] Datasets and Setup: This embodiment primarily uses Caltech101, UCF101, and ImageNet as datasets. Furthermore, the datasets shown in Table 2 are used for the base-to-new-class generalization task. ImageNetV2, ImageNet-Sketch, ImageNet-A, and ImageNet-R are also used as domain generalization benchmarks. On the Caltech and UCF101 datasets, the same configuration as in (1) the fine-grained image classification experiment is used, while for the large-scale ImageNet dataset containing 1000 classes, the number of training epochs is limited to 50.
[0181] Non-fine-grained classification results: Figure 19 A comparison of HIVE with baseline methods is presented. Due to the lack of explicit semantic associations in downstream tasks, HIVE's performance improvement compared to fine-grained classification tasks is not significant. However, it achieves state-of-the-art performance on ImageNet and UCF101, with average accuracies 0.11%, 0.13%, and 0.07% higher than the second-best method, respectively. On the Caltech101 dataset, the accuracy of this invention is 0.05% lower than MaPLe, ranking second. Although HIVE primarily focuses on fine-grained classification, it also demonstrates competitive performance in downstream tasks that lack this feature.
[0182] Generalization from Base Class to New Class: This invention selects three fine-grained datasets and ImageNet for the base class to new class generalization task. In the base class, training is performed using 16 sample examples, and testing is conducted on both the base class and the new class. Table 4 shows the performance of HIVE and baseline methods on these datasets. On the fine-grained image dataset, the inventive method achieves state-of-the-art performance, outperforming the second-best method by 0.27% and 1.09% on OxfordPets and Food101, respectively. While the accuracy on Flowers102 is slightly lower than MaPLe, the inventive method still demonstrates strong performance and significantly outperforms MaPLe in base class accuracy across multiple benchmarks. For non-fine-grained datasets like ImageNet, although our method performs worse than MaPLe and LAMM, it still shows significant improvement over CoOp and zero-sample CLIP. These results demonstrate that the coarse-to-fine-grained modeling paradigm has consistent transferability when samples share a common coarse-grained class. Therefore, after learning this relationship in the base class, the HIVE framework proposed in this invention can effectively generalize to new classes by adaptively modeling the image features of the new classes. The performance gap observed on ImageNet may stem from a significant paradigm shift in cross-class modeling, which presents a generalization challenge. Nevertheless, the overall generalization ability of this method is still commendable.
[0183] Table 4. Accuracy comparison in base-to-new-class generalization test.
[0184]
[0185]
[0186] Domain Generalization: This invention trained HIVE on ImageNet using a 16-sample setup and evaluated its classification accuracy on multiple out-of-domain benchmarks, including ImageNetV2, ImageNetSketch, ImageNet-A, and ImageNet-R. As detailed in Table 5, the proposed method consistently outperforms zero-sample CLIP across all target domains. The results demonstrate that the proposed method is not domain-specific but rather relies on the semantic relationships between images within that domain, highlighting HIVE's robustness to domain variations.
[0187] Table 5. Accuracy Comparison in Domain Generalization Testing
[0188]
[0189]
[0190] (3) Ablation test
[0191] Impact of Spherical Linear Interpolation (SLERP): This invention proposes a spherical linear interpolation (SLERP) method based on hyperspherical distribution and clarifies its geometric meaning. Figure 7 Here, the present invention replaces the SLERP module with a learnable neural network, aiming to train a two-layer multilayer perceptron (MLP) to achieve equivalent functionality, while evaluating its effectiveness through three fine-grained classification downstream tasks. To investigate the fitting ability of the MLP as an alternative, this embodiment relaxes the few-sample constraint. For the baseline SLERP method, its performance in 1-sample and 16-sample settings is reported.
[0192] Figure 20 The performance comparison between the original HIVE and its variant with SLERP replaced by MLP is shown, where the "all" label indicates fully supervised training. For the variant with SLERP replaced by MLP, an early stopping strategy is used to optimize performance and prevent overfitting to ensure a fair comparison. Overall, removing SLERP results in a significant drop in HIVE performance, with accuracy even lower than CLIP zero-shot performance. Although accuracy gradually improves with increasing sampling count, it only approaches the performance level of the original HIVE with 1 sampling. It is worth noting that even when using the entire training dataset ( Figure 10(Marked with an asterisk in the original text), the MLP-based variant also significantly underperformed compared to HIVE's 16-sample results. These findings indicate that the SLERP module, a core component of HIVE, inherently outperforms traditional network architectures in few-sample scenarios due to its geometrically interpretable design. The module's explicit modeling of hyperspherical interpolation gives it superior generalization ability under limited supervision, a key advantage that simple MLP approximation methods cannot replicate even with increased training data.
[0193] Performance under feature distribution bias: In the original CLIP framework, in order to calculate the cosine similarity between text and image features, the modal features output by the encoder are subjected to L... 2 Normalization. Based on this mathematical property, these modal features are distributed on a hypersphere, and this experiment uses a variational autoencoder to model this distribution. In this experiment, we investigated how the performance of HIVE is affected when these features are not normalized (i.e., not distributed on the hypersphere). Three configurations were evaluated: 1) CLIP output features are not normalized (HIVE without pre-normalization); 2) CLIP output features are normalized, and an additional normalization operation is added before the SLERP module (HIVE with post-normalization); 3) CLIP output features are normalized, and no additional normalization operation is performed before the SLERP module (original HIVE). Table 6 shows the single-shot (1-shot) and 16-shot (16-shot) performance for each configuration.
[0194] Table 6 Performance comparison of the HIVE framework under three normalization settings
[0195] Method Single-sample accuracy (1-shot Acc.) 16-shot accuracy Unprecedented normalization of HIVE 63.27 72.32 HIVE with post-normalization 68.89 80.66 Original HIVE 69.04 82.53
[0196] As shown in Table 6, the original HIVE achieved the best performance. Introducing a normalization operation before the SLERP module appears counterproductive, as it partially compromises inherent semantic information while attempting to enhance performance. This modification resulted in a slight performance drop in the single-sample setting, and the performance degradation became increasingly pronounced with the number of samples. Furthermore, the normalization process that removes CLIP output features violates the fundamental hyperspherical distribution assumption, rendering the distribution fitting process ineffective and leading to a significant performance decrease.
[0197] The impact of latent space dimensionality: When modeling a distribution using a hyperspherical variational encoder, the dimension of the latent space can be freely chosen. This dimension d primarily determines the estimated mean. The dimensionality of the latent space is discussed in Table 1. A preliminary analysis of the impact of dimensionality shows that lower-dimensional latent spaces tend to produce smaller reconstruction losses, while higher-dimensional settings exhibit larger log-likelihood values, indicating better fitting of the latent variable distribution. Therefore, selecting an appropriate latent space dimension to achieve optimal performance remains an important research question. This experiment evaluated the performance of HIVE at a range of latent dimensions from 1 to 256, and the results are shown in Table 7. It can be observed that accuracy does not change significantly with dimensionality at low-dimensional settings. The impact seems more pronounced at the 16-sample setting compared to the single-sample setting. However, performance begins to decline significantly with increasing dimensionality, especially above 50. At dimensions of 256 and above, the variational encoder fails to successfully compute the KL divergence. This is because the vMF distribution estimates the KL divergence using a modified Bessel function, which is numerically unstable and non-convergent at high dimensions. Therefore, it is unclear whether the performance degradation at high dimensions is due to mathematical limitations or the dimensionality itself. Based on the current results, choosing a lower latent space dimension yields better overall performance. In the main experiment, d = 5 was set.
[0198] Table 7 Performance comparison of HIVE (Honeycomb) under different potential spatial dimensions
[0199]
[0200]
[0201] Impact of CLIP Backbone Network: Considering that the original CLIP employed various visual backbone networks, including ResNet-50 (RN50), ResNet-101 (RN101), and visual Transformers with 16 and 32 patches (ViT-B / 16 and ViT-B / 32), in Figure 21 The paper presents the average accuracy of HIVE and the original CLIP across all eight fine-grained datasets on different visual backbone networks. The results demonstrate that the proposed method consistently improves performance across all backbone networks, further validating its strong transferability to downstream tasks.
[0202] It should be noted that this embodiment proposes a novel method from a geometric structure perspective, pioneering a new paradigm for transferring CLIP feature extraction and alignment capabilities through latent spatial structural relationships. Unlike previous methods, we do not use labels for fine-tuning; instead, we leverage them to learn cross-modal relationship structures between modalities for aligning the semantics of the CLIP pre-training domain and downstream task domains. Based on this method, a novel CLIP few-shot fine-tuning method, HIVE, is introduced. The key innovation of this embodiment lies in modeling semantic alignment relationships as distributed structural relationships in the latent space. Experimental results demonstrate the effectiveness of HIVE on established benchmarks, achieving a 46.2% improvement in accuracy for single-sample classification and an 80.0% improvement for 16-sample classification, outperforming the original CLIP. This invention's method performs excellently in fine-grained image classification tasks and is also competitive on non-fine-grained datasets, demonstrating strong transferability and generalization capabilities from upstream to downstream domains. The method of this invention also introduces a new concept: few-shot fine-tuning can be redefined as learning the transformation from upstream to downstream, rather than imitating upstream training on downstream data. When using semantic hierarchy for effective few-shot adaptation, pre-training geometric constraints are preserved, which can make more effective use of the limited samples in few-shot learning.
[0203] It should be noted that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.
[0204] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0205] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0206] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0207] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A CLIP-based coarse-to-fine image-text cross-modal semantic transfer method, characterized in that, include: The paired images and coarse-grained text are acquired and input into the cross-modal alignment transfer model, which includes a CLIP model, a hyperspherical variational encoder, and an interpolation module connected in sequence. The CLIP model is used to encode paired images and coarse-grained text to obtain downstream image features and pre-trained text features; The low-dimensional distributions of downstream image features and pre-trained text features are estimated using a hyperspherical variational encoder, resulting in the hyperspherical latent distributions of downstream image features and pre-trained text features. The hypersphere space interpolation is performed using the interpolation module based on the hypersphere latent distribution of downstream image features and the hypersphere latent distribution of pre-trained text features to obtain the estimated distribution of fine-grained text features for the downstream task.
2. The CLIP-based coarse-to-fine image-text cross-modal semantic conversion method as described in claim 1, characterized in that, The hyperspherical variational encoder includes a parallel text hyperspherical variational encoder and a visual hyperspherical variational encoder. The text hyperspherical variational encoder consists of a structurally symmetric text encoder and a text decoder; The visual hyperspherical variational encoder consists of a structurally symmetrical visual encoder and a visual decoder.
3. The CLIP-based coarse-to-fine image-text cross-modal semantic conversion method as described in claim 1, characterized in that, The interpolation module includes a reparameterized structure and a text decoder, wherein: The hypersphere space interpolation of the hypersphere latent distribution of downstream image features and the hypersphere latent distribution of pre-trained text features is performed using the reparameterization structure to obtain the probability density estimate of the interpolated fine-grained text. The probability density estimate of the interpolated fine-grained text is decoded using a text decoder to obtain the fine-grained text feature distribution for downstream tasks.
4. The CLIP-based coarse-to-fine image-text cross-modal semantic conversion method as described in claim 1 or 3, characterized in that, The hypersphere space interpolation process includes: The hyperspherical latent distribution of downstream image features and the hyperspherical latent distribution of pre-trained text features are regarded as position points on the hypersphere, and the fine-grained text features to be estimated are modeled as a unimodal probability density distribution on the hypersphere. The probability density distribution is parameterized by the mode and concentration. Moving at a constant speed along the arc connecting the two location points to perform geodesic interpolation between the two location points, the probability density of the interpolation point with the smallest difference from the true fine-grained text feature distribution is used as the probability density estimate of the interpolated fine-grained text.
5. The CLIP-based coarse-to-fine image-text cross-modal semantic conversion method as described in claim 4, characterized in that, The process of moving at a constant speed along the arc connecting the two position points to perform geodesic interpolation between the two position points, obtaining the probability density of the interpolation point with the smallest difference from the true fine-grained text feature distribution as the probability density estimate of the interpolated fine-grained text, includes: Spherical linear interpolation is used for the mode, and the formula is expressed as: In the formula: Let θ be the mode of the fine-grained text feature distribution to be estimated, θ be the angular distance, and t∈[0,1] be a learnable interpolation coefficient. The mode of the pre-trained text feature distribution. The mode of the downstream image feature distribution; The formula for logarithmic interpolation of the concentration parameter is as follows: In the formula: Let be the concentration parameter of the fine-grained text feature distribution to be estimated. The concentration parameter of the pre-trained text feature distribution. This is the concentration parameter of the downstream image feature distribution; The probability density estimate of the interpolated fine-grained text is as follows:
6. The CLIP-based coarse-to-fine image-text cross-modal semantic transfer method as described in claim 1, characterized in that, Before acquiring the paired images and coarse-grained text and inputting them into the cross-modal alignment transfer model, the method further includes: The cross-modal alignment transfer model is trained in two stages: unsupervised training is used for the hyperspherical variational encoder, and supervised training is used for the interpolation module under the supervision of real fine-grained sample features.
7. The CLIP-based coarse-to-fine image-text cross-modal semantic conversion method as described in claim 6, characterized in that, The total loss function used during unsupervised training of the hyperspherical variational encoder for: In the formula: For the variational encoder loss of the image branch, The variational encoder loss for the text branch. For the joint reconstruction loss function, f i Encode the downstream image features output by the CLIP model. These are coarse-grained text features obtained after cross-decoding by the image decoder. Encode the pre-trained text features output by the CLIP model. These are the image features after cross-decoding by the text decoder. Let be the mathematical expectation of the spatial distance.
8. The CLIP-based coarse-to-fine image-text cross-modal semantic conversion method as described in claim 6, characterized in that, The loss function used during supervised training of the interpolation module for: In the formula: This represents the estimated distribution of fine-grained text features for a downstream task obtained through the text encoder in the hyperspherical variational encoder. This represents the estimated fine-grained textual latent variables. Represents true fine-grained text features. This represents the estimated fine-grained distribution of textual latent variables. This represents the prior distribution of fine-grained textual latent variables. This represents the estimated fine-grained text features after decoding by the text decoder, and KL(||) represents the KL divergence. Let be the mathematical expectation of the spatial distance.
9. A CLIP-based coarse-to-fine image-text cross-modal semantic conversion system, characterized in that, include: An acquisition module is used to acquire paired images and coarse-grained text and input them into a processing module. The processing module is equipped with a pre-trained cross-modal alignment transfer model, which includes a CLIP model, a hyperspherical variational encoder, and an interpolation module connected in sequence. The CLIP model is used to encode paired images and coarse-grained text to obtain downstream image features and pre-trained text features. The hyperspherical variational encoder is used to estimate the low-dimensional distribution of downstream image features and pre-trained text features, and obtain the hyperspherical latent distribution of downstream image features and the hyperspherical latent distribution of pre-trained text features. The interpolation module is used to perform hypersphere space interpolation based on the hypersphere latent distribution of downstream image features and the hypersphere latent distribution of pre-trained text features to obtain the estimated distribution of fine-grained text features for the downstream task.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-8.
Citation Information
Patent Citations
Cross-modal retrieval method and retrieval system
CN115563316A
Comparative incremental learning-based model training method and malicious traffic classification method and system
CN116244645A