A cross-modal retrieval model modeling method, device, terminal and medium

By using the methods of feature decoupling and dynamic temperature control, a cross-modal retrieval model is constructed, which solves the problem of fine-grained semantic information loss in existing technologies and achieves more accurate cross-modal data matching and improved retrieval results.

CN120492701BActive Publication Date: 2025-09-09GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510969597.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-09-09
Estimated Expiration
2045-07-15

AI Technical Summary

Technical Problem

Existing cross-modal retrieval models lose fine-grained semantic information during the global feature alignment process, making it impossible to accurately match cross-modal data with subtle semantic differences, affecting the accuracy of retrieval results.

Method used

Through feature decoupling processing, the sample data is decomposed into multi-granularity features. Combined with the temperature parameter calculation method and dynamic temperature control, a single-granularity contrast loss function is constructed. Combined with the hierarchical orthogonal constraint and the generative adversarial loss function, the total loss function is constructed. Finally, a cross-modal retrieval model is obtained through iterative training.

Benefits of technology

It effectively retains the rich semantic information within the modal data, improves the cross-modal retrieval model's ability to understand and match complex semantic relationships, and thus improves the accuracy and performance of the retrieval results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492701B_ABST
    Figure CN120492701B_ABST
Patent Text Reader

Abstract

The present application discloses a cross-modal retrieval model modeling method, device, terminal and medium, which relates to the field of cross-modal retrieval technology. The solution provided by the present application first performs feature decoupling on the sample data of different modalities according to different modal types, and then calculates the temperature parameters based on the decoupled features and constructs a cross-modal comparative learning model architecture, and then obtains a cross-modal retrieval model through iterative training of the cross-modal comparative learning model. This solution decomposes the overall features of different modal data into multi-granular features with clear semantic orientation and unified dimensions, retaining the rich semantic information within the modal data, and can provide more discriminative and targeted feature representations for subsequent cross-modal comparative learning, which helps to improve the cross-modal retrieval model's understanding and matching capabilities of complex semantic relationships, thereby improving the overall performance of cross-modal retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of cross-modal retrieval technology, and in particular to a cross-modal retrieval modeling method, device, terminal, and medium. Background Art

[0002] In the field of cross-modal retrieval, internet content is experiencing explosive growth, with a massive influx of multimodal data such as images, text, and audio. Users are increasingly demanding efficient and accurate retrieval of cross-modal information, making cross-modal retrieval technology a research hotspot.

[0003] With the rise of deep learning technology, its powerful automatic feature learning capabilities can extract more representative and semantically rich features from raw data, significantly improving cross-modal retrieval performance. However, cross-modal retrieval models based on deep neural networks often use a global feature alignment strategy. This global feature alignment focuses only on overall similarity, resulting in a significant loss of fine-grained semantic information. This makes it impossible for the model to accurately match cross-modal data with subtle semantic differences, which in turn affects the accuracy of cross-modal retrieval results. Summary of the Invention

[0004] The present application provides a cross-modal retrieval modeling method, device, terminal and medium for achieving the invention purpose of improving the accuracy of cross-modal retrieval.

[0005] To achieve the above-mentioned object of the invention, the first aspect of the present application provides a method for modeling a cross-modal retrieval model, comprising:

[0006] Obtain sample data of different modalities;

[0007] According to the modality type corresponding to the sample data, combined with the preset correspondence between the modality type and the feature decoupling processing method, the sample data is subjected to feature decoupling processing to obtain decoupled granularity features of different granularities, wherein the number of granularities obtained by decoupling the sample data of each modality type is the same;

[0008] According to the decoupled particle size characteristics, a temperature parameter corresponding to the particle size is obtained by temperature parameter calculation;

[0009] Calculating the similarity of decoupled particle size features of different sample data under the same particle size condition, and constructing a single particle size contrast loss function based on the ratio of the decoupled particle size feature similarity to the temperature parameter corresponding to the particle size;

[0010] According to the single-granularity contrast loss function, a total loss function is constructed by combining the layered orthogonal constraint and the generative adversarial loss function;

[0011] According to the total loss function, the initial cross-modal retrieval model is iteratively trained in combination with the back-propagation algorithm. When the output of the total loss function or the number of iterations meets the preset model training termination condition, the cross-modal retrieval model is obtained.

[0012] Preferably, the decoupling granularity features include: entity object granularity features, description attribute granularity features and scene context granularity features.

[0013] Preferably, obtaining the temperature parameter corresponding to the particle size by temperature parameter calculation according to the decoupled particle size feature includes:

[0014] According to the granularity to which the decoupled granularity feature belongs and the sub-granularities contained in the granularity, combined with a preset stratified variance calculation formula, the stratified variance corresponding to each sub-granularity is calculated;

[0015] The layered temperature parameters corresponding to the layered variances are obtained by layered temperature mapping calculation, and then the temperature parameters corresponding to the particle size are obtained according to the weighted sum of the layered temperature parameters.

[0016] Preferably, the layered temperature mapping calculation method is specifically as follows:

[0017]

[0018] Where, is the kth granularity The stratification temperature parameter corresponding to the granularity of the layer, is a learnable parameter that adjusts the mapping relationship between variance and temperature parameters. The kth granularity The layer variance corresponding to the layer sub-granularity, For the The dimension of the layer sub-granularity characteristics, is the initial temperature parameter, and The lower and upper limits of the temperature parameter.

[0019] Preferably, the calculating of the decoupled granularity feature similarity of different sample data under the same granularity condition includes:

[0020] Based on a preset sample database, the decoupling granularity feature similarity between the acquired sample data and the sample data in the sample database under the same granularity condition is calculated.

[0021] Preferably, the single-granularity contrast loss function is specifically:

[0022]

[0023] Where, is the single-granularity contrast loss value of the k-th granularity, is the total size of sample batches, is the feature similarity between sample i and sample j at the kth granularity, is the temperature parameter corresponding to the kth particle size, is the set of negative samples at the kth granularity.

[0024] Preferably, the total loss function is specifically:

[0025]

[0026] Where, is the total loss value, is the single-granularity contrast loss value of the k-th granularity, is the granularity weight of the kth granularity, is the hierarchical orthogonal constraint correlation, is the hierarchical orthogonality constraint correlation weight, is the generator loss value, is the adversarial loss weight.

[0027] A second aspect of the present application provides a cross-modal retrieval model modeling device, comprising:

[0028] A multimodal sample data acquisition unit, used to acquire sample data of different modalities;

[0029] a feature decoupling unit, configured to perform feature decoupling processing on the sample data according to the modality type corresponding to the sample data and in combination with a preset correspondence between the modality type and the feature decoupling processing mode, to obtain decoupled granularity features of different granularities, wherein the number of granularities obtained by decoupling the sample data of each modality type is the same;

[0030] a temperature parameter calculation unit, configured to obtain a temperature parameter corresponding to the particle size by a temperature parameter calculation method according to the decoupled particle size characteristics;

[0031] a single-grain size contrast loss determination unit, configured to calculate the similarity of decoupled grain size features of different sample data under the same grain size condition, and construct a single-grain size contrast loss function based on the ratio of the decoupled grain size feature similarity to the temperature parameter corresponding to the grain size;

[0032] a total loss determination unit, configured to construct a total loss function based on the single-granularity contrast loss function, in combination with a hierarchical orthogonal constraint and a generative adversarial loss function;

[0033] A cross-modal retrieval model training control unit is used to iteratively train the initial cross-modal retrieval model based on the total loss function in combination with the back-propagation algorithm, and obtain the cross-modal retrieval model when the output of the total loss function or the number of iterations meets the preset model training termination condition.

[0034] A third aspect of the present application provides a cross-modal retrieval model modeling terminal, comprising: a memory and a processor;

[0035] The memory is used to store program code, and the program code is used to implement a cross-modal retrieval model modeling method as provided in the first aspect of the present application;

[0036] The processor is configured to read and execute the program code.

[0037] The fourth aspect of the present application provides a computer-readable storage medium, in which program code is stored. The program code is used to be read and executed by a processor to implement a cross-modal retrieval model modeling method as provided in the first aspect of the present application.

[0038] It can be seen from the above technical solutions that this application has the following advantages:

[0039] The solution provided in this application first decouples the features of sample data of different modalities according to different modal types, then calculates temperature parameters based on the decoupled features and constructs a cross-modal comparative learning model architecture, and then obtains a cross-modal retrieval model through iterative training of the cross-modal comparative learning model. This solution decomposes the overall features of different modal data into multi-granular features with clear semantic orientation and unified dimensions, retaining the rich semantic information within the modal data. It can provide more discriminative and targeted feature representations for subsequent cross-modal comparative learning, helping to improve the cross-modal retrieval model's understanding and matching capabilities of complex semantic relationships, thereby improving the overall performance of cross-modal retrieval. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0041] Figure 1 A flowchart of an embodiment of a cross-modal retrieval model modeling method provided in this application.

[0042] Figure 2This is a logical block diagram of an embodiment of a cross-modal retrieval model modeling method provided in this application.

[0043] Figure 3 A schematic structural diagram of an embodiment of a cross-modal retrieval model building device provided in this application.

[0044] Figure 4 A schematic structural diagram of an embodiment of a cross-modal retrieval model modeling terminal provided in this application. DETAILED DESCRIPTION

[0045] Existing cross-modal retrieval techniques face challenges arising from the semantic complexity of multimodal data. With the application of deep learning, existing models, such as the CLIP model, achieve cross-modal matching through global feature alignment. However, global feature fusion results in a loss of fine-grained semantic information. For example, when users need to retrieve images containing specific object attributes and scene context, existing methods struggle to distinguish between data with similar overall features but differing details, resulting in inaccurate retrieval results.

[0046] In view of this, embodiments of the present application provide a cross-modal retrieval model modeling method, device, terminal and medium for achieving the invention objective of improving the accuracy of cross-modal retrieval.

[0047] In order to make the purpose, features, and advantages of the invention of this application more obvious and easy to understand, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the embodiments described below are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0048] First, a detailed description of an embodiment of a cross-modal retrieval modeling method provided by this application is as follows:

[0049] See also Figure 1 and Figure 2 , a cross-modal retrieval model modeling method provided in an embodiment of the present application includes:

[0050] Step 101: Obtain sample data of different modalities;

[0051] Step 102: Based on the modal type corresponding to the sample data and the correspondence between the preset modal type and the feature decoupling processing method, the sample data is subjected to feature decoupling processing to obtain decoupling granularity features of different granularities.

[0052] Among them, the number of granularities obtained by decoupling the sample data of each modality type is the same;

[0053] Step 103: Obtain the temperature parameter corresponding to the particle size by calculating the temperature parameter according to the decoupled particle size characteristics;

[0054] Step 104: Calculate the decoupled particle size feature similarity of different sample data under the same particle size condition, and construct a single particle size comparison loss function based on the ratio of the decoupled particle size feature similarity to the temperature parameter corresponding to the particle size;

[0055] Step 105: Based on the single-granularity contrast loss function, combined with the layered orthogonal constraint and the generative adversarial loss function, a total loss function is constructed;

[0056] Step 106: Based on the total loss function, the initial cross-modal retrieval model is iteratively trained in combination with the back-propagation algorithm. When the output of the total loss function or the number of iterations meets the preset model training termination condition, the cross-modal retrieval model is obtained.

[0057] It should be noted that this application proposes to obtain sample data from different modalities and then perform feature decoupling based on the correspondence between the preset modality type and the decoupling method to obtain a consistent number of decoupled granularity features. The corresponding parameters for each granularity are obtained through temperature parameter calculation, and a single-granularity contrast loss function is constructed. This is combined with hierarchical orthogonal constraints and generative adversarial loss to form a total loss function, and finally iterative training is performed to obtain a cross-modal retrieval model.

[0058] Feature decoupling is the process of decomposing raw data into sub-features with clear semantic meaning. This process ensures that the number of granularities after decomposition of data from different modalities is consistent, establishing a unified foundation for subsequent cross-granularity comparisons. The temperature parameter calculation method dynamically adjusts the intensity of contrastive learning based on feature distribution. The single-granularity contrastive loss function is a contrastive learning objective function designed for features at a specific granularity. It is constructed by the ratio of similarity to a temperature parameter and can adjust the learning intensity based on the characteristics of the feature distribution.

[0059] Specifically, after the sample data undergoes decoupling processing through modal adaptation, multi-granular features such as entity objects, attribute descriptions, and scene contexts with unified dimensions are generated. Each granular feature obtains a corresponding temperature parameter through hierarchical variance calculation, which reflects the degree of discreteness of the feature distribution. When constructing the loss function, the feature similarity of different samples at the same granularity forms a proportional relationship with the temperature parameter, so that features with compact distributions obtain stronger contrast constraints. Hierarchical orthogonal constraints act between features of different granularities to prevent semantic information redundancy, and generative adversarial losses enhance the modality independence of feature distribution. Multi-loss joint optimization enables the model to establish precise cross-modal matching relationships while retaining fine-grained semantics.

[0060] This solution addresses the challenge of fine-grained semantic matching in cross-modal retrieval by preserving hierarchical semantic structures through corresponding multi-granularity decoupling for different types of modal data. This approach, combined with hierarchical orthogonal constraints and contrastive learning, improves feature discrimination. Multi-granularity feature decoupling preserves the data's internal hierarchical semantic information, dynamic temperature parameters optimize the contrastive learning process at different granularities, and hierarchical constraints enhance feature discrimination. Ultimately, this improves the model's ability to understand complex semantic relationships and enables more accurate cross-modal data matching.

[0061] On the basis of the above basic embodiment scheme, the present application further provides decoupling granularity features including entity object granularity features, description attribute granularity features and scene context granularity features.

[0062] Among them, for different types of modal data, the decoupling methods and the three decoupled granularity features obtained will be different. For example, the video image features are decomposed through the gated channel attention mechanism, and the obtained entity object granularity features are the target objects in the image, the descriptive attribute granularity features are the object attributes of the target objects, and the scene context granularity features are the scene environment features; the audio features can be decomposed through the spectrum graph to obtain the entity object granularity features as pitch, the descriptive attribute granularity features as rhythm, and the scene context granularity features as timbre granularity; the text features can be extracted through semantic role labeling to obtain the entity object granularity features as events, the descriptive attribute granularity features as emotions, and the scene context granularity features as context information.

[0063] Entity object granularity features refer to features that represent independent entity elements in the data, such as target objects (objects, people, etc.) identified from video or image modal data, and entity objects described in text modal data. Descriptive attribute granularity features refer to features that describe the state or nature of entity objects, such as the shape and color of the target object in an image, and the decomposition of the mel-spectrogram into low-frequency and high-frequency components that represent rhythmic changes. Scene context granularity features refer to features that reflect the environment or background of the data. Specifically, they can be implemented using, for example, the scene and environment in which the target object is located in an image, and semantic role labeling technology to parse text context features.

[0064] Specifically, in the process of video image processing, the gated channel attention mechanism is used to dynamically adjust the weights of each feature channel, so that the network can focus on the channel containing the main object, thereby separating the physical object features. For audio data, the spectrogram is decomposed into the fundamental frequency component representing the pitch, the low-frequency envelope component reflecting the rhythm change, and the high-frequency harmonic component reflecting the timbre. In text processing, semantic role labeling technology is used to identify the core events, modifying emotional words, and background information such as time and place in the sentence. Through the above method, different modal data are uniformly decoupled into three levels of semantic features, ensuring that subsequent cross-modal comparative learning can be carried out at the same granularity dimension.

[0065] Let’s take visual modality data and text modality data as further examples:

[0066] For the visual modality, the input visual features are In order to achieve accurate feature decoupling, a gated channel attention mechanism is introduced. First, the visual features are processed by the global average pooling operation GAP, and the features are compressed in the spatial dimension to obtain a global feature representation that can reflect the statistical information of the visual features as a whole. Then, the global features are combined with the learnable weights Multiply, then pass The function performs normalization and generates gating weights , and its calculation formula is:

[0067]

[0068] Gating weight matrix The Xavier initialization method is used to ensure stability in the initial stage of training. During the training process, It is jointly optimized with other model parameters and adaptively adjusted through back-propagation to capture the importance differences of features at different granularities.

[0069] Based on the generated gating weights , for the original visual features Perform element-wise multiplication to decouple visual features. Specifically, it is decomposed into three granularities: target object, object attribute, and scene environment. The formula is as follows:

[0070]

[0071]

[0072]

[0073] Where, 、 、 Represent the decoupled objects, attributes and scene features respectively, and their dimensions are This decomposition method based on gated channel attention not only highlights the importance of features at different granularities, but also improves parameter efficiency through learnable weights, enabling the model to learn more representative multi-granularity visual features with fewer parameters.

[0074] In the text mode, the input text features are In order to mine the semantic hierarchy in the text, this module uses dependency parsing technology to determine the position index of entity words, modifiers, and context words in the text, which are recorded as .

[0075] Based on these position indexes, we can use the multi-layer perceptron (MLP) to process and sum the features of the corresponding positions to obtain the entity, modification, and context feature representations of the text. The specific calculation formula is as follows:

[0076]

[0077]

[0078]

[0079] in, 、 、 They are the entity, modification and context features of the text, and the dimensions are Through this semantic role labeling method, the semantic relationship between words in the text is fully considered, making the decoupling of text features more consistent with semantic logic and enhancing the expressive power of text features at different semantic levels.

[0080] Through the above technical solution, this application realizes the fine-grained feature decoupling of cross-modal data in three dimensions: entity objects, object attributes, and scene context, by decomposing the overall features of multimodal data into multi-granular features with clear semantic orientation. The material properties of objects in video images can be cross-modally compared with the rhythmic features of audio signals, and the emotional elements in text descriptions can be associated with scene environment features. This hierarchical decoupling method effectively solves the semantic confusion problem caused by global feature alignment, enabling the model to capture more subtle semantic associations between cross-modal data.

[0081] In some embodiments, the present application further proposes to calculate the stratified variance corresponding to each sub-granularity based on the granularity to which the decoupled granularity feature belongs and the sub-granularity it contains, combined with a preset stratified variance calculation formula; obtain the stratified temperature parameters corresponding to each stratified variance through a stratified temperature mapping calculation method, and then obtain the temperature parameters corresponding to the granularity based on the weighted sum of each stratified temperature parameter.

[0082] Among them, the hierarchical variance calculation formula refers to the mathematical expression used to calculate the difference in sub-granularity feature distribution. Specifically, it can be implemented by the variance statistics of each dimension of each sub-granularity feature vector. This calculation formula can reflect the degree of discreteness of the internal features of the sub-granularity. The hierarchical temperature mapping calculation method refers to converting the hierarchical variance into a functional relationship of the temperature parameter. Specifically, it can be implemented by using a learnable linear transformation combined with a nonlinear activation function. By adjusting the mapping relationship between the variance and the temperature parameter, the temperature parameter can adapt to changes in the feature distribution. The weighted sum refers to the linear combination of temperature parameters at different levels according to preset weights. Specifically, a dynamic weight allocation strategy based on the importance of the sub-granularity can be adopted to ensure that the final temperature parameter can comprehensively reflect the characteristic characteristics of each layer of sub-granularity.

[0083] It should be noted that in cross-modal retrieval tasks, features of different granularities have significant differences in distribution characteristics and semantic complexity, and the traditional fixed temperature parameter contrast learning method cannot adapt to this diversity, which greatly limits the learning effect and performance improvement of the model. In current cross-modal retrieval models, a single temperature parameter is generally used to control the intensity of contrast learning. This approach ignores the inherent differences between features of different granularities. For example, when processing object-level features and scene-level features of an image, object-level features focus on the details of specific objects, and their feature distribution is relatively concentrated; while scene-level features cover a wider range of environmental information and are more dispersed. Using the same temperature parameter for contrast learning cannot effectively optimize the different characteristics of these two features, resulting in the model being difficult to fully explore the semantic relationship between features, affecting retrieval accuracy. Therefore, in terms of temperature parameter calculation, the present application further provides a dynamic hierarchical temperature parameter control logic, which can dynamically adjust the temperature parameters of contrast learning according to the distribution of features of different granularities, and provide the most suitable learning environment for features of each granularity, thereby optimizing the contrast learning process and improving the performance of the model.

[0084] Specifically, the input of dynamic temperature control is the characteristics of each particle size , where B v and B t Represents the sample batch size of the two sets of granularity features, d is the feature dimension. The output is the adaptive temperature parameter , which correspond to the temperature parameters of entity objects, object attributes and scene context granularity features respectively.

[0085] The calculation process is mainly divided into three key steps. The first step is hierarchical variance calculation. By taking the distribution characteristics of different granularity features (objects, attributes, scenes), further subdividing them into sub-granularities (for example, object granularity includes shape, color, etc.), the hierarchical variance is calculated separately. The specific calculation formula is as follows:

[0086]

[0087] In this formula, For the The dimension of the layer sub-granularity characteristics, represents the j-th dimension feature value of the i-th sample in the k-th granularity feature, is the mean value of the j-th dimension of the k-th granularity feature, is the total size of the sample batch, for example, if the input feature is the above and , then B=B v +B t By calculating the variance, we can accurately understand the discreteness of the features in each dimension. The larger the variance, the more dispersed the distribution of the features in that dimension.

[0088] The second step is hierarchical temperature mapping, which generates hierarchical temperature parameters based on sub-granularity variance and obtains the final temperature parameters through weighted fusion:

[0089]

[0090]

[0091] in, is a hierarchical learnable parameter used to adjust the mapping relationship between variance and temperature parameters. The weights are predefined (e.g., shape weight 0.6, color weight 0.4 in object granularity). It is trained jointly with other model parameters, and its gradient is updated through backpropagation of the total loss function. The initial value is set to 1.0 and is dynamically adjusted according to the feature distribution during training to adapt to the optimization requirements of features of different granularity. is the initial temperature value; =0.1 and = 5.0 limits the temperature parameter's range. The Clip function ensures that the generated temperature parameter remains within a reasonable range. When feature variance is large, it indicates a more dispersed feature distribution, and the calculated temperature parameter will be correspondingly larger. In contrastive learning, a higher temperature parameter relaxes the requirement for feature similarity, allowing the model to focus more on rough feature matching and avoid ignoring overall semantic relationships due to excessive focus on detailed differences. Conversely, when variance is small, lowering the temperature parameter increases the stringency of contrast, forcing the model to focus on subtle differences in features and better capture fine-grained semantic distinctions between them.

[0092] Through the variance calculation and temperature mapping process of dynamic temperature control, the temperature parameter is adaptively adjusted according to the distribution of features at different granularities. This dynamic control mechanism provides a more reasonable optimization method for contrastive learning in cross-modal retrieval, enabling the model to better adapt to the learning needs of features at different granularities, effectively improving the model's understanding and learning ability of complex semantic relationships, and thus improving the performance and accuracy of cross-modal retrieval.

[0093] In some embodiments, the present application further proposes calculating the decoupling granularity feature similarity of different sample data under the same granularity conditions, including: based on a preset sample database, calculating the decoupling granularity feature similarity of the acquired sample data and the sample data in the sample database under the same granularity conditions.

[0094] The pre-set sample database refers to a structured dataset that stores multimodal sample data and its decoupled granular features. This can be implemented using a distributed graph database or vector database, providing a reference sample set for cross-modal comparative learning during model training. Decoupled granular feature similarity measures the degree of association between different sample features at the same granularity level. This can be implemented using cosine similarity, Euclidean distance, or a bilinear attention mechanism, to quantify the degree of matching between different modal data at the same semantic granularity.

[0095] It's important to note that in cross-modal retrieval scenarios, features of different granularities vary significantly in semantic richness and distribution. Traditional contrastive learning and negative sampling methods struggle to fully exploit the value of these features, limiting the model's performance when processing complex cross-modal data. To address this challenge, this module builds a semantically aware multi-granularity contrastive optimization module that integrates semantically aware negative sampling and multi-granularity contrastive loss calculation to comprehensively improve the model's performance in cross-modal retrieval.

[0096] Specifically, suppose that in the current cross-modal retrieval task, there is a batch of visual feature sets: , the text feature set is , where B=B v +B t, represents the total batch size, and Represent the dimensions of visual and text features respectively. There is also a momentum memory for storing historical features. The visual feature memory is , the text feature memory is , M is the memory capacity. The capacity M of the visual and textual feature memory is set to 65536, referring to the empirical design from the MOCO series of work. This capacity balances historical feature coverage with computational efficiency, ensuring that the memory stores a sufficient diversity of samples to support difficult negative sample screening.

[0097] Next, perform semantic-aware negative sample screening. Calculate the similarity between the current batch of visual features and the visual feature memory, and between the text features and the text feature memory, using the following formula:

[0098]

[0099]

[0100] The screening threshold is set according to the characteristics of different granularity features. For example, for the visual characteristics of object granularity, the threshold is set to ; For scene-level text features, the threshold is set to (These thresholds will be fine-tuned according to the actual cross-modal retrieval task.) The rules for filtering difficult negative samples are:

[0101]

[0102]

[0103] After the screening is completed, the momentum memory library is updated. The momentum update strategy is adopted, and the specific formula is as follows:

[0104]

[0105]

[0106] Here m is the momentum parameter, which is used to balance the ratio of historical information and current new information in the memory bank, so that the memory bank can effectively adapt to the dynamic changes of data.

[0107] Furthermore, the above basic embodiment mentions the construction of a single-granularity contrast loss function based on the ratio of the decoupled granularity feature similarity to the temperature parameter corresponding to the granularity, and the construction of a total loss function based on the single-granularity contrast loss function, combined with the hierarchical orthogonal constraint and the generative adversarial loss function. The specific expressions and construction process of these two functions can be found in the following examples:

[0108] For features of different granularities, such as object (obj), attribute (attr), and scene (scene) granularity, the contrast loss is calculated separately. The calculation formula for single-granularity contrast loss is:

[0109]

[0110] In this formula, It is used to measure the similarity between the features of sample i and sample j at the kth granularity, where sample j belongs to the negative sample set. The specific similarity calculation method can be flexibly selected according to the actual situation, such as the commonly used cosine similarity. The temperature parameter corresponding to the granularity is used to adjust the strength of the contrast loss and control the level of refinement in model learning. When the feature distribution is relatively dispersed, appropriately increasing the temperature parameter relaxes the strictness of the contrast. Conversely, lowering the temperature parameter increases the strictness of the contrast, thereby guiding the model to more effectively learn the differences between features of different granularities.

[0111] To ensure the independence of features of different granularities and avoid feature redundancy and interference, hierarchical orthogonal constraints are introduced:

[0112]

[0113] in, is the correlation coefficient obtained based on the hierarchical orthogonal constraint formula, For stratification weights (e.g. ), which is used to adjust the orthogonal constraint strength between features of different granularity. Represents the Frobenius norm, which is used to measure the overall "size" of the matrix. Here, it is used to calculate the correlation penalty term between feature matrices of different granularities. and Represent the visual and text features of the k-th granularity, and Representing the granular visual and textual features, k and Both represent granularity indices. The hierarchical orthogonality constraint penalizes the correlation between features of different granularities, prompting the model to learn more discriminative feature representations and improving the model's feature expression capabilities. The superscript "T" represents the transpose of the feature matrix.

[0114] At the same time, the discriminator D is introduced to distinguish whether the features are independent, and the generated adversarial samples are optimized for decoupling:

[0115] Pair Features Add adversarial perturbations to generate adversarial samples:

[0116]

[0117] in, is the perturbation intensity (default ), It is the gradient of the discriminator loss with respect to the feature, which is used to calculate the perturbation direction when generating adversarial samples.

[0118] Discriminator loss:

[0119]

[0120] Among them, B is the total batch size, which refers to the number of samples input each time when the model is trained. is the true feature of the i-th sample at the k-th granularity (such as object, attribute, scene). Adversarial sample features of the i-th sample.

[0121] Generator loss:

[0122]

[0123] in, is the adversarial loss weight.

[0124] Finally, the total loss function is obtained by combining the single-granularity contrast loss and orthogonal constraint terms:

[0125]

[0126] In this total loss function, have , , , They are the losses of object (obj), attribute (attr), and scene (scene) granularity respectively. Weight parameters are used to balance the contribution of different losses to the total loss. By adjusting these weights, the model can more rationally allocate learning resources during training, effectively learn the relationship between various granular features, and achieve comprehensive optimization of model performance.

[0127] Next, during model training, the gradient is calculated using the backpropagation algorithm based on the total loss function, and the model parameters are then updated. Through continuous iterative training, the model can continuously optimize its understanding and ability to distinguish features of different granularities, improving its performance in cross-modal retrieval tasks. The semantic-aware negative sampling strategy ensures the effective use of difficult negative samples, enabling the model to learn the differences between more challenging sample pairs and enhance its understanding of complex semantic relationships. The design of the multi-granularity contrast loss fully considers the differences in features of different granularities. By dynamically adjusting the temperature parameters and introducing orthogonal constraints, the model can better adapt to complex cross-modal data distributions, improving the model's generalization ability and retrieval accuracy.

[0128] The above is a detailed description of the conceptual principles of a cross-modal retrieval model modeling method provided by this application. The following is a detailed description of an example of an implementation application scenario based on the cross-modal retrieval model modeling method provided by this application and the cross-modal retrieval model constructed therefrom, as follows:

[0129] Suppose a large e-commerce platform needs to build a cross-modal retrieval system for product images and product description text to help users quickly find corresponding product images by entering text, or find matching product descriptions by uploading images, thereby improving the efficiency of users' product search.

[0130] First, in the data preparation phase, a large amount of product images and corresponding text descriptions are collected. This data covers a variety of product categories, such as clothing, electronics, and household items, forming a training dataset.

[0131] Next, the collected data is processed using the framework of the present invention. In the feature decoupling module, the product image is input into the visual feature processing part. Through the gated channel attention mechanism, the overall features of the image are decomposed into three granular features: object, attribute, and scene. For example, for a picture of a dress, the object feature focuses on the dress itself, the attribute feature can be the color, pattern, etc. of the dress, and the scene feature is information such as the background environment in the picture. For the product description text, with the help of dependency syntax analysis technology, the position index of the entity words, modifiers, and context words in the text is determined, and then the entity, modification, and context features of the text are obtained respectively using a multi-layer perceptron. Taking the text "Fashionable printed short-sleeved dress, suitable for summer travel" as an example, "dress" is the entity, "fashionable print" and "short sleeve" are modifiers, and "suitable for summer travel" belongs to context information.

[0132] Next, the dynamic temperature controller takes effect. It analyzes the statistical properties of features at different granularities, such as calculating the variance of object, attribute, and scene-level features. For object features focused on the details of the dress, if their distribution is relatively concentrated and their variance is small, the dynamic temperature controller lowers the temperature parameter for contrastive learning, allowing the model to focus more on the subtle differences in these features. For scene features covering a wider range of environmental information, if their distribution is more dispersed and their variance is larger, the temperature parameter is increased, allowing the model to focus more on the rough matching situation, avoiding excessive focus on details while ignoring overall semantic relationships.

[0133] Next, the semantic-aware multi-granularity contrastive optimization module is entered. A momentum memory is first constructed to store historical visual and textual features. During training, the similarity between the current batch of product image features and the visual feature memory, and between the product description text features and the textual feature memory, is calculated. Screening thresholds are set based on the characteristics of features at different granularities to filter out difficult negative samples. For example, for visual features at the object granularity, if a threshold is set, any sample with a calculated similarity greater than the threshold and not a positive sample is considered a difficult negative sample. After screening, the memory is updated according to the momentum update strategy. Next, contrastive losses are calculated for features at the object, attribute, and scene granularities. At the same time, orthogonal constraints are introduced to avoid redundancy and interference between features at different granularities. The total loss function is derived by combining the contrastive losses of each single granularity and the orthogonal constraints. By adjusting the weight parameters in the total loss function, the model can rationally allocate learning resources during training.

[0134] During model training, the gradient is calculated using the backpropagation algorithm based on the total loss function, and the model parameters are continuously updated. After multiple rounds of iterative training, the model's ability to understand and distinguish features of different granularities is continuously optimized. In actual cross-modal retrieval tests, by inputting the product description text "blue short-sleeved T-shirt", the model can accurately retrieve T-shirt images that meet the description from a large number of product images; or by uploading a picture of sports shoes, the model can also accurately find the matching product description text, such as "breathable and lightweight sports shoes, suitable for sports enthusiasts". This shows that the multi-granularity dynamic contrast learning framework of the present invention can effectively improve the accuracy and efficiency of retrieval in actual cross-modal retrieval scenarios, meeting the needs of e-commerce platforms for efficient cross-modal retrieval systems.

[0135] The embodiment of this application proposes a cross-modal retrieval model construction scheme based on multi-granularity dynamic contrastive learning. This scheme uses a multimodal encoder to obtain different modal features, and deeply processes and optimizes the features through a unique module design to achieve more accurate and efficient cross-modal retrieval, improving the performance of the model in complex cross-modal data scenarios. The concept of this scheme can also be extended to multimodal retrieval scenarios such as audio-text and video-text.

[0136] The above is a detailed description of an embodiment of a cross-modal retrieval model modeling method provided by the present application. The following is a detailed description of an embodiment of a cross-modal retrieval model modeling device provided by the present application.

[0137] See also Figure 3 In a second aspect, the present application provides a cross-modal retrieval model modeling device, comprising:

[0138] A multimodal sample data acquisition unit 201 is used to acquire sample data of different modalities;

[0139] The feature decoupling unit 202 is configured to perform feature decoupling processing on the sample data according to the modality type corresponding to the sample data and in combination with a preset correspondence between the modality type and the feature decoupling processing mode, thereby obtaining decoupled granularity features of different granularities, wherein the number of granularities obtained by decoupling the sample data of each modality type is the same;

[0140] The temperature parameter calculation unit 203 is used to obtain the temperature parameter corresponding to the particle size through the temperature parameter calculation method according to the decoupled particle size characteristics;

[0141] The single-grain size contrast loss determination unit 204 is used to calculate the decoupled grain size feature similarity of different sample data under the same grain size condition, and construct a single-grain size contrast loss function based on the ratio of the decoupled grain size feature similarity to the temperature parameter corresponding to the grain size;

[0142] A total loss determination unit 205 is configured to construct a total loss function based on the single-granularity contrast loss function, combined with the layered orthogonal constraint and the generative adversarial loss function;

[0143] The cross-modal retrieval model training control unit 206 is used to iteratively train the initial cross-modal retrieval model based on the total loss function in combination with the back-propagation algorithm. When the output of the total loss function or the number of iterations meets the preset model training termination condition, the cross-modal retrieval model is obtained.

[0144] In addition, the present application also provides a detailed description of a cross-modal retrieval model modeling terminal and a computer-readable storage medium embodiment.

[0145] like Figure 4 As shown, the present application provides a cross-modal retrieval model modeling terminal, which can be implemented as a personal computer, an industrial computer, a server, and an embedded intelligent device. The terminal mainly comprises a memory 33 and a processor 31, and the memory 33 and the processor 31 can be connected via a communication bus 34;

[0146] The memory 33 is used to store program codes, and the program codes are used to implement a cross-modal retrieval model modeling method provided in the above embodiment;

[0147] The processor 31 is used to read and execute program codes.

[0148] Among them, the memory refers to the hardware module used to store program code and training data. Specifically, it can be implemented by a solid-state drive or a mechanical hard drive, and is used to persistently save the algorithm logic and parameter configuration required for cross-modal retrieval model training. The processor refers to the computing unit that executes the program code. Specifically, it can be implemented by a multi-core central processing unit or a graphics processing unit, and accelerates the matrix operations and gradient updates in the model training process through parallel computing. The program code refers to the algorithm set that implements core functions such as feature decoupling, temperature parameter calculation, and contrast loss construction. Specifically, it can include a data preprocessing module, a feature decoupling network layer, and a loss calculation unit. Modular programming is used to ensure the reusability of each functional component.

[0149] Specifically, the terminal solidifies the model training algorithm through the storage medium, and the processor executes a sequence of operations such as loading sample data, feature decoupling processing, and dynamic calculation of temperature parameters. During operation, the memory first loads the preset modality type and decoupling method mapping table, and the processor performs hierarchical feature decomposition on the input multimodal data according to the mapping table. For example, when processing video data, the gated attention mechanism is called to generate three granular features, and the corresponding temperature parameters are calculated based on the feature distribution. The processor iteratively updates the model parameters through the backpropagation algorithm until the loss function converges or reaches the preset iteration threshold, and finally generates a cross-modal retrieval model with multi-granularity feature matching capabilities.

[0150] The present application provides a computer-readable storage medium, in which program code is stored. The program code is used to be read and executed by a processor to implement a cross-modal retrieval model modeling method as provided in the above embodiment.

[0151] A computer-readable storage medium refers to a physical medium capable of persistently storing program code. Specifically, this can be achieved using a solid-state drive, optical disc, or flash memory chip. This achieves data storage through physical structural changes, ensuring that program code can be preserved for a long time even after a power outage. Program code refers to a computer language collection containing executable instructions. Specifically, it can be implemented using Python or C++ programming languages. By defining feature decoupling processing, temperature parameter calculation, loss function construction, and model training processes, it transforms the cross-modal retrieval modeling approach into logical steps that can be executed by a computer.

[0152] Specifically, when the program code is executed by the processor, it first controls the computer to read multimodal sample data from the storage device. It then performs a hierarchical semantic decomposition of the data according to preset feature decoupling rules. For example, a gated channel attention mechanism is implemented on image data to decompose physical object features. Temperature parameters are then dynamically calculated based on the distribution characteristics of the decoupled features. A contrastive loss function with hierarchical orthogonal constraints is constructed, and the model parameters are iteratively optimized using a backpropagation algorithm. The entire training process is fully automated, requiring no manual intervention to adjust the feature alignment strategy.

[0153] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the terminals, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0154] In the several embodiments provided in this application, it should be understood that the disclosed terminals, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.

[0155] The terms "first," "second," "third," "fourth," and the like (if any) in the specification of the present application and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present application described herein, for example, can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such process, method, product, or apparatus.

[0156] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0157] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0158] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0159] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0160] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A cross-modal retrieval model modeling method, characterized in that: include: Obtain sample data of different modalities; According to the modality type corresponding to the sample data, combined with the preset correspondence between the modality type and the feature decoupling processing method, the sample data is subjected to feature decoupling processing to obtain decoupled granularity features of different granularities, wherein the number of granularities obtained by decoupling the sample data of each modality type is the same, and the decoupled granularity features include: entity object granularity features, descriptive attribute granularity features, and scene context granularity features; According to the decoupled particle size characteristics, a temperature parameter corresponding to the particle size is obtained by temperature parameter calculation; Calculating the similarity of decoupled particle size features of different sample data under the same particle size condition, and constructing a single particle size contrast loss function based on the ratio of the decoupled particle size feature similarity to the temperature parameter corresponding to the particle size; According to the single-granularity contrast loss function, a total loss function is constructed by combining the layered orthogonal constraint and the generative adversarial loss function; Iteratively training the initial cross-modal retrieval model according to the total loss function in combination with a back-propagation algorithm, and obtaining a cross-modal retrieval model when the output of the total loss function or the number of iterations meets a preset model training termination condition; According to the modality type corresponding to the sample data and in combination with the preset correspondence between the modality type and the feature decoupling processing mode, the sample data is subjected to feature decoupling processing to obtain decoupling granularity features of different granularities, including: When the modality type corresponding to the sample data is a video image feature, the gated channel attention mechanism is used for decomposition, and the obtained entity object granularity feature is the target object in the image, the descriptive attribute granularity feature is the object attribute of the target object, and the scene context granularity feature is the scene environment feature; When the modality type corresponding to the sample data is an audio feature, the entity object granularity feature obtained by spectrogram decomposition is pitch, the description attribute granularity feature is rhythm, and the scene context granularity feature is timbre granularity; When the modality type corresponding to the sample data is text features, the entity object granularity features extracted through semantic role labeling are events, the description attribute granularity features are emotions, and the scene context granularity features are context information.

2. A cross-modal retrieval model modeling method according to claim 1, characterized in that: The step of obtaining the temperature parameter corresponding to the particle size by calculating the temperature parameter according to the decoupled particle size feature includes: According to the granularity to which the decoupled granularity feature belongs and the sub-granularities contained in the granularity, combined with a preset stratified variance calculation formula, the stratified variance corresponding to each sub-granularity is calculated; The layered temperature parameters corresponding to the layered variances are obtained by layered temperature mapping calculation, and then the temperature parameters corresponding to the particle size are obtained according to the weighted sum of the layered temperature parameters.

3. A cross-modal retrieval modeling method according to claim 2, characterized in that: The layered temperature mapping calculation method is specifically as follows: ; Where, is the kth granularity The stratification temperature parameter corresponding to the granularity of the layer, is a learnable parameter that adjusts the mapping relationship between variance and temperature parameters. The kth granularity The layer variance corresponding to the layer sub-granularity, For the The dimension of the layer sub-granularity characteristics, is the initial temperature parameter, and The lower and upper limits of the temperature parameter.

4. A cross-modal retrieval modeling method according to claim 1, characterized in that: Calculating the similarity of decoupled granularity features of different sample data under the same granularity condition includes: Based on a preset sample database, the decoupling granularity feature similarity between the acquired sample data and the sample data in the sample database under the same granularity condition is calculated.

5. A cross-modal retrieval modeling method according to claim 1, characterized in that: The single-granularity contrast loss function is specifically: ; Where, is the single-granularity contrast loss value of the k-th granularity, is the total size of sample batches, is the feature similarity between sample i and sample j at the kth granularity, is the temperature parameter corresponding to the kth particle size, is the set of negative samples at the kth granularity.

6. A cross-modal retrieval modeling method according to claim 1, characterized in that: The total loss function is specifically: ; Where, is the total loss value, is the single-granularity contrast loss value of the k-th granularity, is the granularity weight of the kth granularity, is the hierarchical orthogonal constraint correlation, is the hierarchical orthogonality constraint correlation weight, is the generator loss value, is the adversarial loss weight.

7. A cross-modal retrieval model building device, characterized in that: include: A multimodal sample data acquisition unit, used to acquire sample data of different modalities; a feature decoupling unit configured to perform feature decoupling processing on the sample data according to the modality type corresponding to the sample data and in combination with a preset correspondence between the modality type and the feature decoupling processing mode, to obtain decoupled granularity features of different granularities, wherein the number of granularities obtained by decoupling the sample data of each modality type is the same, and the decoupled granularity features include: entity object granularity features, descriptive attribute granularity features, and scene context granularity features; a temperature parameter calculation unit, configured to obtain a temperature parameter corresponding to the particle size by a temperature parameter calculation method according to the decoupled particle size characteristics; a single-granularity contrast loss determination unit, configured to calculate the similarity of decoupled granularity features of different sample data under the same granularity condition, and construct a single-granularity contrast loss function based on the ratio of the decoupled granularity feature similarity to the temperature parameter corresponding to the granularity; a total loss determination unit, configured to construct a total loss function based on the single-granularity contrast loss function, in combination with a hierarchical orthogonal constraint and a generative adversarial loss function; A cross-modal retrieval model training control unit, configured to iteratively train the initial cross-modal retrieval model based on the total loss function in combination with a back-propagation algorithm, and obtain a cross-modal retrieval model when the output of the total loss function or the number of iterations meets a preset model training termination condition; The feature decoupling unit is specifically used for: When the modality type corresponding to the sample data is a video image feature, the gated channel attention mechanism is used for decomposition, and the obtained entity object granularity feature is the target object in the image, the descriptive attribute granularity feature is the object attribute of the target object, and the scene context granularity feature is the scene environment feature; When the modality type corresponding to the sample data is an audio feature, the entity object granularity feature obtained by spectrogram decomposition is pitch, the description attribute granularity feature is rhythm, and the scene context granularity feature is timbre granularity; When the modality type corresponding to the sample data is text features, the entity object granularity features extracted through semantic role labeling are events, the description attribute granularity features are emotions, and the scene context granularity features are context information.

8. A cross-modal retrieval model modeling terminal, characterized in that: include: memory and processor; The memory is used to store program code, and the program code is used to implement a cross-modal retrieval model modeling method according to any one of claims 1 to 6; The processor is configured to read and execute the program code.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores program code, which is used to be read and executed by a processor to implement a cross-modal retrieval model modeling method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Cross-modal retrieval method and device based on semantic enhancement, storage medium and terminal

    CN114780777A

  • Cross-modal retrieval method and system based on multi-granularity feature fusion

    CN115391625A