Multimodal data processing method and system based on step-by-step feature enhancement network

By using a step-by-step feature enhancement network and combining external and internal semantic clues, we construct multi-layer co-occurrence relationships, solving the problem of existing technologies failing to fully utilize cross-modal semantic information and improving the accuracy of multimodal data processing and retrieval performance.

CN120104843BActive Publication Date: 2025-09-12SHANDONG HI SPEED COMPANY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510577614.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-09-12
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

Existing multimodal data processing technologies fail to fully utilize the semantic information from corpora of different modal data, especially the external semantic clues of single modal data, which limits the ability to bridge the heterogeneous gap and leads to reduced accuracy of multimodal data processing.

Method used

A step-by-step feature enhancement network is adopted to construct multi-layer co-occurrence relationships through the self-attention mechanism and graph convolutional neural network. External semantic clues and internal semantic clues are combined to achieve cross-modal feature enhancement, establish a semantic propagation path, and flow from the external layer to the internal layer in steps to fully mine cross-modal semantic information.

Benefits of technology

It improves the accuracy of similarity measurement of data in different modalities, enhances the retrieval performance of multimodal data processing, and enhances modal interaction capabilities by comprehensively capturing the semantic clues behind the relationships between different modalities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104843B_ABST
    Figure CN120104843B_ABST
Patent Text Reader

Abstract

The present invention proposes a multimodal data processing method and system based on a step-by-step feature enhancement network, which relates to the technical field of multimodal data processing and targets the following problems: the semantic information of the multimodal corpus is not fully utilized, the potential of cross-modal ESC is not fully explored, and the accuracy of multimodal data processing is low. The method extracts the features of the first modality and the second modality data, and performs self-attention mechanism feature enhancement on them; based on different modality data corpora, semantic concepts of the first modality data and the second modality data are constructed respectively, and multi-layer co-occurrence relationship mining is performed; the features enhanced by the self-attention mechanism are further enhanced by the step-by-step feature enhancement network, and the final data features of the first modality and the second modality are obtained by fusion respectively. The present invention provides a step-by-step feature enhancement network, which effectively improves the multimodal data processing performance by combining cross-modal enhancement of external semantic clues and context mining of internal semantic clues.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of cross-modal image-text mutual retrieval, and in particular relates to a multimodal data processing method and system based on a step-by-step feature enhancement network. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] In today's big data era, multimodal data processing refers to the processing and analysis of information from multiple sensory channels or data forms. These modalities can include text, images, audio, video, and more. Data from different modalities have different characteristics and representations, making it challenging to effectively integrate and process this heterogeneous data. For example, image-text inter-retrieval (ITR) is a fundamental visual-linguistic task that aims to search for images related to a given text query or retrieve text related to a given image query. The key task of ITR is to learn cross-modal similarities between images and text. A major challenge facing ITR is that data from different modalities (images and text) are encoded using different encoders, resulting in differences between these modalities. This encoding difference creates a "heterogeneous gap," making it difficult to directly compare and align features from the two modalities in a unified feature space, thereby increasing the complexity of multimodal data processing.

[0004] There are inherent inconsistencies between the feature representations of multimodal data, namely the "heterogeneous gap". To bridge the heterogeneous gap in multimodal data processing tasks, capturing as many semantic clues as possible is crucial to establishing effective associations between different modal data. These semantic clues can be divided into two categories:

[0005] Internal semantic cues (ISC) refer to the semantic concepts and relationships within data from different modalities. ISC highlights the intrinsic semantic structure between them, including intra-modal correlations within each modality and inter-modal correlations between different modalities.

[0006] External semantic cues (ESC) refer to semantic concepts and their relationships derived from a large-scale corpus of multimodal data, rather than from data from a single modality. ESC captures external semantic patterns, including associations between paired examples in the corpus, such as proximity and co-occurrence frequency. By analyzing global data distribution patterns, ESC provides a global perspective for semantic understanding, complementing the intrinsic semantic information in multimodal data.

[0007] Existing research on multimodal data processing mainly focuses on extracting ISC from data of different modalities by analyzing intra-modal and inter-modal semantics. Current methods are generally divided into three categories: (1) mining ISC in one modality, such as methods such as VSRN and CAMERA; (2) mining ISC across different modalities, such as SCAN, ESL, and MPARN; and (3) hybrid methods that combine (1) and (2), such as HREM and BOOM. However, these methods mainly focus on ISC in single-modal data, while ignoring ESC existing in a wider corpus of data of different modalities, which limits their comprehensive bridging of the heterogeneous gap.

[0008] To address the cognitive limitations of focusing only on ISC in different modal data, recent studies have begun to focus on ESC outside of different modal data. Research on ESC in different modal data processing can be divided into two categories: (1) mining ESC from neighbors, such as DSP, TBNN, NAN, LeaPRR; and (2) mining ESC from common sense knowledge, such as CVSE, VCM, and CSRC. Despite this, these methods still focus on single-modal ESC, that is, extracting ESC from each modality independently, and fail to fully explore the potential of cross-modal ESC. This deficiency limits their ability to bridge the heterogeneous gap and significantly reduces the accuracy of multimodal data processing.

[0009] In summary, the inventors found that existing multimodal data processing technology has not yet fully utilized the semantic information from the corpus of different modal data, limiting its comprehensive bridging of the heterogeneous gap, especially beyond the clues of single modal data, and only independently extracting the ESC of each modality, failing to fully tap the potential of cross-modal ESC, affecting the accuracy of the similarity measurement of different modal data, and thus reducing the accuracy of multimodal data processing. Summary of the Invention

[0010] To overcome the shortcomings of the aforementioned prior art, the present invention provides a multimodal data processing method and system based on a step-by-step feature enhancement network. Through this step-by-step feature enhancement network, a semantic propagation path is established, guiding the flow of semantic information from the external layer to the internal layer in steps, achieving step-by-step feature enhancement from the external layer to the internal layer. By combining cross-modal enhancement of external semantic cues with contextual mining of internal semantic cues, multimodal data processing performance is effectively improved.

[0011] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:

[0012] A first aspect of the present invention provides a multimodal data processing method based on a step-by-step feature enhancement network, comprising:

[0013] Acquiring first modal data and second modal data;

[0014] Extracting first modality data features and second modality data features;

[0015] Performing feature enhancement on the first modality data features and the second modality data features through a self-attention mechanism to obtain enhanced first modality global-level features and enhanced second modality global-level features;

[0016] Based on the first modality data corpus and the second modality data corpus, constructing a first modality semantic concept set and a second modality semantic concept set respectively;

[0017] Perform multi-layer co-occurrence relationship mining on the first modality semantic concept set and the second modality semantic concept set;

[0018] Based on the multi-layer co-occurrence relationship, the enhanced first modality global-level features and the enhanced second modality global-level features are further enhanced through a step-by-step feature enhancement network to obtain multi-layer first modality data features and multi-layer second modality data features, and the final data features of the first modality data and the second modality data are fused respectively.

[0019] As an implementation method, the first modal data feature is extracted, and the specific process is as follows:

[0020] Utilize an object detection model with bottom-up attention to extract the target region;

[0021] A deep convolutional network is used to extract the first modality data features of the target area.

[0022] As an implementation method, the second modal data features are extracted, specifically, by encoding the second modal data using an encoder of a deep neural network to obtain the second modal data features.

[0023] As an implementation method, multi-layer co-occurrence relationship mining is performed on the first modal semantic concept set and the second modal semantic concept set, where the multi-layer co-occurrence relationship includes the fragment layer, the instance layer, and the neighbor layer. The specific process is as follows:

[0024] Perform unimodal co-occurrence relationship mining at the fragment level and generate the first modal initial features and second modal co-occurrence relationship matrix at the fragment level;

[0025] Perform cross-modal co-occurrence relationship mining at the instance level and generate the first modal initial features and second modal co-occurrence relationship matrix at the instance level;

[0026] Perform cross-modal co-occurrence relationship mining at the neighboring layer and generate the first modal initial visual features and the second modal co-occurrence relationship matrix of the neighboring layer;

[0027] Adaptive weak co-occurrence relationship filtering is performed on the first modal initial features and the second modal co-occurrence relationship matrix to obtain a co-occurrence relationship matrix for weak co-occurrence relationship filtering.

[0028] As an implementation method, the enhanced first modality global-level features and the enhanced second modality global-level features are further enhanced by a step-by-step feature enhancement network to obtain multi-layer first modality data features and multi-layer second modality data features, and the final data features of the first modality data and the second modality data are respectively fused. The specific process is as follows:

[0029] Performing external semantic enhancement on the enhanced first modality global-level features and the enhanced second modality global-level features through multi-layer co-occurrence relationships in the external layer;

[0030] The enhanced first modality global-level features and the enhanced second modality global-level features are internally semantically enhanced through the cross-modal context of the internal layer.

[0031] As an implementation method, external semantic enhancement is performed on the enhanced first modality global-level features and the enhanced second modality global-level features through the multi-layer co-occurrence relationship of the external layer, specifically:

[0032] Based on the first modality semantic concept set and the second modality semantic concept set, constructing a first modality and a second modality weighted undirected graph;

[0033] Use graph convolutional neural networks to update the node embeddings of the first modality and second modality weighted undirected graphs layer by layer, and obtain the node embedding feature matrices of the first modality graph and the second modality graph respectively;

[0034] Based on the node embedding feature matrices of the first modal graph and the second modal graph, calculating the first modal feature enhanced by the first modal co-occurrence relationship, the first modal feature enhanced by the second modal co-occurrence relationship, the second modal feature enhanced by the first modal co-occurrence relationship, and the second modal feature enhanced by the second modal co-occurrence relationship;

[0035] Based on the first modality co-occurrence relationship, the first modality feature is enhanced, the second modality feature is enhanced, the first modality feature is enhanced, and the second modality feature is enhanced based on the second modality co-occurrence relationship. The feature enhancement function is used to perform layered first modality feature enhancement and second modality feature enhancement to obtain enhanced first modality features and second modality features in each layer.

[0036] Obtaining enhanced data features of the first modal data based on the first modal data features and the enhanced first modal features of each layer;

[0037] Based on the second modality data features and the enhanced second modality features of each layer, enhanced data features of the second modality data are obtained.

[0038] As an implementation method, internal semantic enhancement is performed on the enhanced first modality global-level features and the enhanced second modality global-level features through the cross-modal context of the internal layer. The specific process is as follows:

[0039] Calculating a similarity matrix between the first modality label and the second modality label based on the augmented data features of the first modality data and the enhanced data features of the second modality data;

[0040] According to the similarity matrix, the first modality attention matrix and the second modality attention matrix are calculated;

[0041] Obtaining an updated feature representation of the first modality data according to the first modality attention matrix and the enhanced data features of the second modality data;

[0042] Obtaining an updated feature representation of the second modality data according to the second modality attention matrix and the enhanced data features of the first modality data;

[0043] The enhanced data features of the first modal data and the updated first modal data feature representation are concatenated column by column, and passed through a fully connected layer to obtain the first modal data features enhanced with cross-modal information;

[0044] The enhanced data features of the second modality data and the updated second modality data feature representation are concatenated column by column and passed through a fully connected layer to obtain the second modality data features enhanced by cross-modal information.

[0045] A second aspect of the present invention provides a multimodal data processing system based on a step-by-step feature enhancement network, comprising:

[0046] A data acquisition module, configured to acquire first modal data and second modal data;

[0047] A feature extraction module, configured to extract features of the first modal data and features of the second modal data;

[0048] An attention enhancement module is used to enhance the features of the first modality data and the second modality data through a self-attention mechanism to obtain enhanced first modality global-level features and enhanced second modality global-level features;

[0049] A semantic construction module, configured to construct a first modality semantic concept set and a second modality semantic concept set based on the first modality data corpus and the second modality data corpus, respectively;

[0050] A multi-layer co-occurrence module is used to mine multi-layer co-occurrence relationships between the first modality semantic concept set and the second modality semantic concept set;

[0051] The feature enhancement module is used to further enhance the enhanced first modality global-level features and the enhanced second modality global-level features based on the multi-layer co-occurrence relationship through a step-by-step feature enhancement network to obtain multi-layer first modality data features and multi-layer second modality data features, and fuse them respectively to obtain the final data features of the first modality data and the second modality data.

[0052] The third aspect provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the program.

[0053] A fourth aspect provides a computer-readable storage medium having a computer program stored thereon, which performs the steps of the above method when executed by a processor.

[0054] One or more of the above technical solutions have the following beneficial effects:

[0055] The step-by-step feature enhancement network SFE in this embodiment establishes a semantic propagation path, guides the flow of semantic information from the external layer to the internal layer in steps, and realizes step-by-step enhancement from the external layer to the internal layer. By combining the cross-modal enhancement of external semantic clues and the context mining of internal semantic clues, the potential of cross-modal ESC is fully explored, the accuracy of the similarity measurement of data in different modalities is improved, and SFE effectively improves the retrieval performance.

[0056] In this embodiment, a multi-level co-occurrence mining method guided by ESC is used for feature enhancement, including co-occurrence relationships at the segment level, instance level, and neighbor level. These co-occurrence relationships significantly strengthen the modal interactions in the outer layer. The cross-modal context mining method guided by ISC further enhances the modal interactions in the inner layer by combining the data features of different modalities with their corresponding contextual features.

[0057] In this embodiment, based on the three-layer co-occurrence relationship, it has a positive impact on multimodal data processing, integrates the co-occurrence relationship of the fragment layer, instance layer and neighbor layer, makes full use of the semantic information from the large-scale modal corpus, and makes full use of their complementary advantages, thereby comprehensively capturing the semantic clues behind different modal relationships.

[0058] In this embodiment, intra-modal and cross-modal feature enhancement are combined to overcome the limitation of intra-modal information lacking inter-modal interaction and make full use of all available semantic clues to improve retrieval performance.

[0059] Through the method of this embodiment, future research in the field of multimodal data processing can further explore the semantic cues between different modalities (such as video and audio). Such work can effectively expand the scope of semantic cues and lead to more effective feature enhancement, ultimately improving multimodal data processing performance.

[0060] Advantages of additional aspects of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0062] Figure 1 This is a schematic diagram of a step-by-step feature enhancement network framework of the first embodiment;

[0063] Figure 2 This is a framework diagram of the external feature enhancement module based on multi-level co-occurrence relationship mining in the first embodiment of the present invention;

[0064] Figure 3 This is a framework diagram of the internal feature enhancement module for cross-modal context mining in the first embodiment of the present invention;

[0065] Figure 4 The visual feature representation results of the baseline using T-SNE on the MSCOCO test set in Example 1 are shown;

[0066] Figure 5 The result of visual feature representation of SFE by T-SNE on the MSCOCO test set in Example 1 is shown;

[0067] Figure 6 This is the visualization text feature representation result of the baseline using T-SNE on the MSCOCO test set in Example 1;

[0068] Figure 7 This is the result of visual text feature representation of SFE by T-SNE on the MSCOCO test set in Example 1. DETAILED DESCRIPTION

[0069] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.

[0070] It should be noted that the terms used herein are for describing particular embodiments only and are not intended to limit the exemplary embodiments according to the present invention.

[0071] In the absence of conflict, the embodiments of the present invention and the features thereof may be combined with each other.

[0072] Example 1

[0073] This embodiment discloses a multimodal data processing method based on a step-by-step feature enhancement network.

[0074] To more clearly illustrate this embodiment, a multimodal data processing implementation process based on a step-by-step feature enhancement network can be specifically described as follows:

[0075] A multimodal data processing method based on a step-by-step feature enhancement network, comprising:

[0076] S1. Acquire first modal data and second modal data;

[0077] S2. Extracting first modality data features and second modality data features;

[0078] S3. Perform feature enhancement on the first modality data features and the second modality data features through a self-attention mechanism to obtain enhanced first modality global-level features and enhanced second modality global-level features;

[0079] S4. Constructing a first modal semantic concept set and a second modal semantic concept set based on the first modal data corpus and the second modal data corpus, respectively;

[0080] S5. Perform multi-layer co-occurrence relationship mining on the first modal semantic concept set and the second modal semantic concept set;

[0081] S6. Based on the multi-layer co-occurrence relationship, the enhanced first modality global-level features and the enhanced second modality global-level features are further enhanced through a step-by-step feature enhancement network to obtain multi-layer first modality data features and multi-layer second modality data features, and the final data features of the first modality data and the second modality data are respectively fused.

[0082] In this embodiment, the multimodal data processing method based on the step-by-step feature enhancement network can be applied to data processing between different modalities, such as image and text mutual retrieval. The following analysis and explanation will take image and text mutual retrieval as an example.

[0083] like Figure 1 As shown, in step S1, first modal data and second modal data are acquired.

[0084] In this embodiment, the target image and target text are obtained from the Visual Genome dataset.

[0085] like Figure 1 As shown, in step S2, the first modal data feature and the second modal data feature are extracted.

[0086] S2-1. Extract first modal data features.

[0087] In this embodiment, the first modal data feature is extracted, and the specific process is as follows:

[0088] (1) Utilize the object detection model with bottom-up attention to extract the target region.

[0089] In this implementation, in the image-text mutual retrieval process, the pre-trained object detection model Faster R-CNN with bottom-up attention is used to extract the top-ranked salient regions with the highest confidence, i.e., visual landmarks, from each image on the Visual Genome dataset.

[0090] (2) Use a deep convolutional network to extract the first modality data features of the target area.

[0091] In this implementation, during the image-text mutual retrieval process, the deep convolutional network ResNet-101 is used to extract the features of the above regions, and the fully connected (FC) layer is used as the visual feature to convert it into a d-dimensional feature space. The formula is:

[0092] (1)

[0093] in, is the coded image area , and are the parameters to be learned.

[0094] The visual features of the image region are represented as

[0095] (2)

[0096] in is the number of image regions, represents the visual features of the first image region, Indicates the Visual features of an image region.

[0097] S2-2. Extract the second modality data features.

[0098] In this embodiment, in the image-text mutual retrieval process, an encoder based on a deep neural network, such as Bi-GRU and Bert, is used to encode text words. , which is converted to Dimension, the formula is:

[0099] (3)

[0100] in, Represents text words To encode, and is a learnable parameter.

[0101] The text features of text words are represented as:

[0102] (4)

[0103] in, is the number of words in the text, The text feature representing the first text word, Indicates the The text features of the text words.

[0104] If data processing is performed on other modal data, modal data features are extracted for specific modal data.

[0105] After the above steps, the initial visual features of the target image and the initial text features of the target text are extracted, providing a data basis for further feature enhancement using the subsequent attention enhancement feature and step-by-step feature enhancement grid.

[0106] like Figure 1 As shown, in step S3, the first modality data features and the second modality data features are feature enhanced through the self-attention mechanism to obtain enhanced first modality global-level features and enhanced second modality global-level features.

[0107] In this embodiment, a self-attention mechanism is used to fully utilize the complementary semantic relationships within each modality to enhance feature representation.

[0108] (1) Obtain enhanced first modal global-level features.

[0109] In this implementation, in the image-text mutual retrieval process, visual markup features As key terms and value terms, the global visual feature vector is the query term of the self-attention mechanism:

[0110]

[0111] (5)

[0112]

[0113] in, , K and Represent the query vector, key vector and value vector in the self-attention mechanism respectively, , and is the weight matrix that needs to be trained.

[0114] The self-attention mechanism calculates the query vector by calculating the dot product similarity and key vector , thereby obtaining enhanced global visual features, the formula is:

[0115] (6)

[0116] in, , represents the enhanced global visual features, Indicates that the initial visual markup feature R learns context information through the self-attention module, Represents the dimension of attention features after dimensionality reduction.

[0117] (2) Obtain enhanced second modality global-level features.

[0118] In this embodiment, in the process of image-text mutual retrieval, the method in (1) is adopted to determine the features of the text and the features of the text mark. As key and value items, the global text feature vector As query items, according to formula (6), they are input into function to obtain enhanced global level text features .

[0119] Among them, if data processing is performed on other modal data, enhanced modal global-level features are obtained for specific modal data.

[0120] After the above steps, the global-level image features and text features enhanced by self-attention are obtained, which achieves preliminary enhancement of image features and text features and improves the accuracy of feature extraction.

[0121] like Figure 1 As shown, in step S4, a first modal semantic concept set and a second modal semantic concept set are constructed based on the first modal data corpus and the second modal data corpus, respectively.

[0122] In order to fully mine the semantic associations in a corpus, it is necessary to fully mine the various semantic concepts contained in the corpus and learn the semantic associations between them. The co-occurrence relationship between semantic concepts can well reflect the semantic associations between concepts.

[0123] In this embodiment, an image-text semantic concept set is constructed for image-text mutual retrieval.

[0124] (1) Construct a set of visual semantic concepts.

[0125] The image is divided into blocks, and each image is divided into 9 blocks on average. Due to the large number of images and the noise they contain, a single image or a single image block cannot directly represent the visual semantic concept.

[0126] For the i image patch, calculate its distance to each centroid, and assign it to the cluster associated with the nearest centroid, the formula is:

[0127] (7)

[0128] in, It is i The cluster index to which the image block belongs; Indicates the i Image blocks visual characteristics; is the square of the Euclidean distance; It is k The characteristics of the centroid k The visual word representation of a centroid is , Through the k The features of all image blocks in a cluster are averaged and the formula is:

[0129] (8)

[0130] in, is the feature set of the image patch belonging to the k-th cluster, and the centroid of the k-th class is defined as the visual word , express The features of , the visual semantic concept set is expressed as: , Represents the total number of visual semantic concepts.

[0131] (2) Construct a set of text semantic concepts.

[0132] All words in the corpus are considered as candidate text semantic concepts. In order to alleviate the impact of word sparsity and irrelevant terms, infrequent words need to be filtered out.

[0133] Specifically, 1) high-frequency words in the concept vocabulary are identified as text semantic concepts and classified into objects, movements, and attributes.

[0134] 2) According to the statistical frequency of concepts in the dataset, keep the concepts (object, motion, attribute) in a balanced distribution ratio of (7:2:1) to obtain the text semantic concept set, which is recorded as ,in Represents the total number of semantic concepts in the text.

[0135] 3) Use word embedding method GloVe to extract text semantic concepts The characteristics of .

[0136] Among them, if data processing is performed on other modal data, a corresponding semantic concept set is constructed for the specific modal data.

[0137] like Figure 1 As shown, in step S5, multi-layer co-occurrence relationship mining is performed on the first modal semantic concept set and the second modal semantic concept set.

[0138] Multi-level co-occurrence relationships include unimodal co-occurrence relationships and cross-modal co-occurrence relationships. These relationships are divided into three layers: fragment layer, instance layer, and neighbor layer. Specifically, the fragment layer corresponds to unimodal co-occurrence relationships, while the instance layer and neighbor layer correspond to cross-modal co-occurrence relationships. For each layer, a semantic concept association graph is constructed, which encapsulates different co-occurrence relationships. In each layer, we construct a co-occurrence matrix based on the corresponding co-occurrence relationship. and To model the association between different semantic concepts.

[0139] S5-1. Perform unimodal co-occurrence relationship mining at the fragment layer, and generate the first modal initial features and the second modal co-occurrence relationship matrix of the fragment layer.

[0140] Fragment Layer ( Unimodal co-occurrence focuses on the relationships within a single modality, i.e., the relationships within image patches or text words. Unimodal concept associations are constructed using visual or textual semantic concepts independently at the segment level.

[0141] In this embodiment, and denote the visual and textual co-occurrence frequency matrices respectively. Specifically, express and The number of times they co-occur in the corpus, and yes The total number of occurrences in the same corpus. Represents a given visual semantic concept In the presence of visual semantic concepts Probability of occurrence.

[0142] Algorithm 1 is used to generate the initial visual and initial text co-occurrence relationship matrices at the segment layer, which are expressed as and .

[0143] Algorithm 1: Generate the single-modal co-occurrence relationship matrix at the segment level. The steps are:

[0144] Input: image-text pair , visual semantic concepts , text semantic concepts .

[0145] Output: Visual and textual co-occurrence matrix of the initial segment layer , .

[0146] 1: Initialize the matrix to 0 and .

[0147] 2: for each image do:

[0148] 3: for each and do:

[0149] 4: if and Also appears in the image Then:

[0150] 5: ;

[0151] 6: Each element of is set to ;

[0152] 7: for each text do:

[0153] 8: for each and do:

[0154] 9: if and Also appears in the text Then:

[0155] 10: ;

[0156] 11: Each element of is set to ;

[0157] 12: return , .

[0158] After the above steps, using data statistical methods, a matrix representing the co-occurrence relationship between visual semantic concepts is generated for the image modality at the fragment level. , generates a matrix representing the co-occurrence relationship between text semantic concepts for the text modality Statistics on the co-occurrence of semantic concepts help to explore the potential semantic associations between different concepts and improve the model's understanding of semantic concepts. In addition, by analyzing co-occurrence patterns, the model can more accurately capture contextual information, improving retrieval and reasoning performance.

[0159] S5-2 performs cross-modal co-occurrence relationship mining at the instance layer and generates the first modal initial features and the second modal co-occurrence relationship matrix at the instance layer.

[0160] At the instance level, the pairwise relationship between image and text is used as a constraint to mine cross-modal co-occurrence relationships. The instance-level cross-modal co-occurrence frequency matrix is ​​denoted as ,in, Indicates the corpus and Frequency of co-occurrence in the same image-text pair.

[0161] Algorithm 2 is used to generate the initial visual and text co-occurrence relationship matrix at the instance layer, which is expressed as , .

[0162] Algorithm 2: Generate instance-level cross-modal co-occurrence relationship matrix. The steps are:

[0163] Input: image-text pair , visual semantic concepts , text semantic concepts .

[0164] Output: Visual and textual co-occurrence matrix of the initial instance layer , .

[0165] 1: Initialize the matrix to zero.

[0166] 2: for each do:

[0167] 3: for each and do:

[0168] 4: if and Appearing at the same time Then:

[0169] 5: ; ;

[0170] 6: , ;

[0171] 7: return , .

[0172] After the above steps, a cross-modal co-occurrence matrix is ​​generated at the instance level using data statistical methods. , and use this matrix to generate a matrix for the image modality to represent the co-occurrence relationship between visual semantic concepts at the instance level , for text modality, generates a matrix representing the co-occurrence relationship between text semantic concepts at the instance level Statistics on the co-occurrence of cross-modal semantic concepts help to explore potential semantic associations between different modalities and enhance the model's ability to understand cross-modal information. In addition, by analyzing cross-modal co-occurrence patterns, semantic alignment can be promoted, allowing the model to more accurately capture cross-modal semantic consistency and complementarity, thereby improving the effectiveness of cross-modal retrieval and reasoning.

[0173] S5-3. Perform cross-modal co-occurrence relationship mining at the neighboring layer, and generate the first modal initial visual features and the second modal co-occurrence relationship matrix of the neighboring layer.

[0174] To further explore cross-modal co-occurrence relationships in a broader semantic space, we extend the cross-modal co-occurrence relationship from the instance level to the neighbor level. This extension aims to reveal richer cross-modal co-occurrence relationships by analyzing the visual neighbors of the target image and the textual neighbors of the target text.

[0175] At the neighbor layer, the pairing relationship between the visual neighbors of the target image and the textual neighbors of the target text is used as a constraint to discover broader cross-modal co-occurrence relationships.

[0176] For an image and a text , using the k-nearest neighbor (KNN) algorithm to identify their first n nearest neighbors, image The nearest neighbors of ,text The nearest neighbors of The cross-modal co-occurrence frequency matrix of the neighboring layer is expressed as ,in express and The co-occurrence frequency of visual neighbors and textual neighbors belonging to the same image-text pair in the corpus.

[0177] Algorithm 3 is used to generate the initial visual features and text co-occurrence relationship matrix of the neighbor layer, which can be expressed as , .

[0178] Algorithm 3 generates the cross-modal co-occurrence relationship matrix at the neighboring layer. The steps are:

[0179] Input: image-text pair , visual semantic concepts , text semantic concepts .

[0180] Output: Visual and textual co-occurrence matrix of the initial nearest neighbor layer , .

[0181] 1: Initialize the matrix to zero.

[0182] 2: for each image-text pair do:

[0183] 3: ; ;

[0184] 4: for each visual neighbor do:

[0185] 5: for each text neighbor do:

[0186] 6: for each and do:

[0187] 7: if and Appearing at the same time Then:

[0188] 8: ; ;

[0189] 9: , ;

[0190] 10: return , .

[0191] S5-4. Perform adaptive weak co-occurrence relationship filtering on the first modal initial features and the second modal co-occurrence relationship matrix to obtain a weak co-occurrence relationship filtered co-occurrence relationship matrix.

[0192] Filtering out weak co-occurrence relationships in the corpus is a key step in improving the performance of the SFE network, which can be summarized into two aspects: on the one hand, weak co-occurrence relationships will cause the model to learn incorrect associations, resulting in overfitting of patterns that lack generalization ability; on the other hand, weak co-occurrence relationships exacerbate the imbalance caused by the long-tail distribution, making the model tend to learn irrelevant patterns.

[0193] Each element Matrix is a random variable ,in . Take less than or equal to The cumulative probability of a value is expressed as ,in, is the probability density function.

[0194] Then, use the inverse function Sure The value of express The inverse function of , the corresponding cumulative distribution function is . The calculation formula is: .

[0195] Since the data distribution patterns of each co-occurrence matrix are different, a pre-defined Represents the proportion of weak co-occurrence relationships. Then, the threshold According to the probability density function For the first matrices Perform adaptive learning. Assume that the co-occurrence matrix Corresponding to the parameter ,by For example, the adaptive weak co-occurrence relationship filtering formula is:

[0196] (9)

[0197] Using the same method as above, the matrix , , , , , Weak co-occurrence thresholds in , , , , , Perform adaptive filtering to generate a co-occurrence relationship matrix filtered by weak co-occurrence relationships , , , , , .

[0198] Among them, if data processing is performed on other modal data, multi-layer co-occurrence relationship mining will be performed on the semantic concept sets of different modalities.

[0199] like Figure 1-3As shown, in step S6, based on the multi-layer co-occurrence relationship, the enhanced first modality global-level features and the enhanced second modality global-level features are further enhanced through a step-by-step feature enhancement network to obtain multi-layer first modality data features and multi-layer second modality data features, and the final data features of the first modality data and the second modality data are respectively fused.

[0200] Semantic concepts and their relationships derived from a large-scale image-text corpus are called ESCs. The ESCs extracted from multi-layer co-occurrences are used to enhance the visual and textual features in the outer layers of the SFE network. The feature enhancement process guided by ESCs is denoted as EFE.

[0201] S6-1. Perform external semantic enhancement on the enhanced first modality global-level features and the enhanced second modality global-level features through multi-layer co-occurrence relationships in the external layer.

[0202] In this embodiment, the specific process is:

[0203] (1) Based on the first modality semantic concept set and the second modality semantic concept set, a first modality and second modality weighted undirected graph is constructed.

[0204] Introducing multi-layer co-occurrence in semantic concept feature representation is helpful for The visual semantic concept of the layer constructs a weighted undirected visual graph, the formula is:

[0205] (10)

[0206] in, Represents a set of nodes, each node corresponds to a visual semantic concept; represents a set of edges, , where the edge Representation node and The relationship between ; the edge weight matrix is ​​expressed as ,in, express and the strength of the relationship between .

[0207] (2) The node embeddings of the first modality and the second modality weighted undirected graphs are updated layer by layer using a graph convolutional neural network to obtain the node embedding feature matrices of the first modality graph and the second modality graph respectively.

[0208] First, a graph convolutional network (GCN) with a total number of layers L is used to update the node embedding of the visual graph layer by layer. . No. l The output of the layer is:

[0209] (11)

[0210] in, For GCN l The node embedding feature matrix of the layer, Is a visual semantic concept The feature representation of represents the edge weight matrix, Indicates that it is located in the GCN l-1 The node embedding feature matrix of the layer, Representation visual graph In GCN l The weight matrix that the layer needs to learn.

[0211] Then, the output of the Lth layer of GCN is used as the final output, that is, the node embedding feature matrix of the visual graph is .

[0212] Similarly, using the above method, we can get the embedding feature matrix of the text graph .

[0213] (3) Based on the node embedding feature matrix of the first modal graph and the second modal graph, the first modal feature enhanced by the first modal co-occurrence relationship, the first modal feature enhanced by the second modal co-occurrence relationship, the second modal feature enhanced by the first modal co-occurrence relationship, and the second modal feature enhanced by the second modal co-occurrence relationship are calculated.

[0214] In this embodiment, visual features and text features are enhanced under the guidance of ESC.

[0215] In order to achieve deep cross-modal interaction, the co-occurrence relationship within the homogeneous modality (same modality) and between heterogeneous modalities (different modalities) is used to enhance features. The calculation formula for enhancing the visual features of the visual co-occurrence relationship of the layer is:

[0216] (12)

[0217] in, express Tier The features of a visual semantic concept, express The feature transformation result is defined as , represents the learnable parameter matrix, is a parameter that controls the fractional intensity, Indicates the number of visual semantic concepts contained in the target image.

[0218] The calculation formula of the text co-occurrence relationship enhanced visual feature is:

[0219] (13)

[0220] in, express Tier The features of a visual semantic concept, express The feature transformation result is defined as , represents the learnable parameter matrix, Indicates the number of text semantic concepts contained in the target text.

[0221] Similarly, according to formulas (12) and (13), the visual co-occurrence relationship is derived to enhance the text feature Enhance text features with text co-occurrence relationships .

[0222] Will and is classified as co-occurrence enhancement of isomorphic modalities (homomodality), while and Co-occurrence enhancement categorized as heterogeneous modalities (different modalities).

[0223] (4) Based on the first modality co-occurrence relationship to enhance the first modality feature, the second modality co-occurrence relationship to enhance the first modality feature, the first modality co-occurrence relationship to enhance the second modality feature, and the second modality co-occurrence relationship to enhance the second modality feature, the feature enhancement function is used to perform layered first modality feature enhancement and second modality feature enhancement to obtain the enhanced first modality feature and second modality feature of each layer.

[0224] In this embodiment, feature enhancement is performed on the visual features and text features at the segment level, instance level, and neighbor level, respectively, based on the ESC-guided Feature Enhancement (EFE) function.

[0225] 1) In the fragment layer, by the parameter Control function The visual features of the fragment layer are enhanced, and the formula is:

[0226] (14)

[0227] (15)

[0228] in, represents the enhanced global visual features, Indicates that visual co-occurrence relationships enhance visual features, Representing text co-occurrence relationships to enhance visual features.

[0229] Similarly, by the parameter The controlled fragment-level text features are enhanced using the formula:

[0230] (16)

[0231] in, represents the enhanced global text features, Indicates that visual co-occurrence relationships enhance text features, Represents text co-occurrence relationships to enhance text features.

[0232] 2) In the parameters and Under the control of , , and .

[0233] (5) Based on the first modal data features and the enhanced first modal features of each layer, the enhanced data features of the first modal data are obtained.

[0234] In this embodiment, for i visual markers , its visual characteristics Through fusion To enhance, the formula is:

[0235] (17)

[0236] Where ⊙ represents element-wise multiplication.

[0237] Then, the enhanced visual features of all regions are concatenated row by row to obtain the enhanced visual features of the target image. The formula is:

[0238] (18)

[0239] in, represents the number of image regions in the target image, Indicates that the target image is The visual features of the first image region after further feature enhancement, Indicates that the target image is After further feature enhancement, the Visual features of an image region.

[0240] (6) Based on the second modality data features and the enhanced second modality features of each layer, the enhanced data features of the second modality data are obtained.

[0241] In this embodiment, for j Text tags , through fusion Enhance its text features , the formula is:

[0242] (19)

[0243] Then, the enhanced visual features of the target text are obtained by connecting the enhanced features of all text tags line by line, as follows:

[0244] (20)

[0245] in, represents the number of text words in the target text, Indicates that the target text is in The text features of the first text word after feature enhancement, Indicates that the target text is in After feature enhancement The text features of the text words.

[0246] Among them, if data processing is performed on data of other modalities, external semantic enhancement is performed based on the multi-layer co-occurrence relationship mining of different modalities.

[0247] After the above steps, mining multi-layer co-occurrence relationships at the external layer effectively overcomes the limitations of semantic associations within a single image-text pair, capturing potential and comprehensive intermodal associations in a higher-level semantic space. Furthermore, mining multi-layer co-occurrence relationships allows for the incorporation of more external knowledge, making the trained model more adaptable to different datasets or tasks.

[0248] S6-2. Perform internal semantic enhancement on the enhanced first modality global-level features and the enhanced second modality global-level features through the cross-modal context of the internal layer.

[0249] like Figure 3 As shown, in this embodiment, visual and textual features are enhanced by mining ESC from multi-level co-occurrence relationships in the outer layer. To effectively complement ISC, valuable semantic information is propagated from the outer layer ESC to the inner layer ISC. Through the IFE module, ISC is combined with cross-modal context mining to further enhance visual and textual features.

[0250] Specifically, (1) based on the augmented data features of the first modality data and the enhanced data features of the second modality data, a similarity matrix between the first modality label and the second modality label is calculated.

[0251] In order to effectively aggregate cross-modal information, we first calculate the similarity matrix between visual and textual tags. S, the formula is:

[0252] (twenty one)

[0253] in, and Respectively The feature vector and A collection of feature vectors of dimensional text tags, each element Quantified visual markers and text tags The similarities between them.

[0254] (2) Calculate the first modality attention matrix and the second modality attention matrix based on the similarity matrix;

[0255] Then, we use S to obtain the attention matrix of the two modalities. For the attention of a specific area, the calculation formula of its visual attention matrix is:

[0256] (twenty two)

[0257] Among them, the softmax function The rows of are normalized to produce the distribution of attention weights across textual tokens for each visual token.

[0258] (3) According to the first modality attention matrix and the enhanced data features of the second modality data, an updated feature representation of the first modality data is obtained.

[0259] These attention weights are used to aggregate the features of the text tags to obtain the updated visual feature representation, which is formulated as:

[0260] (twenty three)

[0261] in, The optimized feature representation representing the visual token associated with each text token, ie, the updated visual feature representation.

[0262] (4) According to the second modality attention matrix and the enhanced visual features of the first modality data, an updated feature representation of the second modality data is obtained.

[0263] Similarly, for the attention of a specific word, the text attention matrix calculation formula is:

[0264] (twenty four)

[0265] Among them, the softmax function The rows of are normalized to produce the attention weights for each word in the context of each visual token.

[0266] These weights are used to aggregate visual features to generate updated text feature representations, as follows:

[0267] (25)

[0268] in, The optimized feature representation of the textual tag corresponding to each visual tag is represented as the updated textual feature representation.

[0269] After the above steps, the transformed feature set and It encapsulates the fine-grained feature representation produced by the cross-modal attention mechanism, thereby effectively capturing the complex relationship between visual and textual modalities.

[0270] (5) The enhanced data features of the first modal data and the updated first modal data feature representation are concatenated column by column and passed through a fully connected layer to obtain the first modal data features enhanced by cross-modal information.

[0271] Visual features With its cross-modal context Concatenate column by column and transfer to the fully connected layer to obtain visual features enhanced by cross-modal information. The formula is:

[0272] FC(concat[ (26)

[0273] (6) The enhanced data features of the second modality data and the updated second modality data feature representation are concatenated column by column and passed through a fully connected layer to obtain the second modality data features enhanced by cross-modal information.

[0274] Text features With its cross-modal context Splice them column by column and transfer them to the fully connected layer to obtain text features enhanced by cross-modal information. The formula is:

[0275] FC(concat (27).

[0276] Among them, if data processing is performed on data of other modalities, internal semantic enhancement is performed based on the multi-layer co-occurrence relationship mining of different modalities.

[0277] S6-3. Fusing the final data features of the first modality data and the final data features of the second modality data to obtain the final data features.

[0278] Enhanced visual features of the fragment layer, instance layer and neighbor layer and enhanced text features The final visual features and text features of the first modal data and the target text are fused respectively, and the formula is:

[0279] (28)

[0280] (29)

[0281] in, and They represent the final visual features of the target image and the final text features of the target text, respectively.

[0282] After the above steps, the enhancement process from the outside to the inside fully compensates for the limitations of the internal semantic relationship of the image-text pair.

[0283] S6-4. Using the final first modal data features and the final second modal data features, perform data processing on the first modal data and the second modal data.

[0284] In this embodiment, the final image visual features and text features are used to perform image-text mutual retrieval.

[0285] For matching image-text pairs , respectively collect and and Mismatched pairs and ,in, and Represents the most difficult negative example. Using the final feature representation and , respectively for the images and text Then, the widely used bidirectional triplet ranking loss is used to align the image and text. The bidirectional triplet ranking loss formula is:

[0286] (30)

[0287] in, is the boundary hyperparameter, sim( ) represents the cosine distance function, , the selection formula for the most difficult negative example is: and .

[0288] In this embodiment, before mutual retrieval of the target image and the target text, the constructed network is first trained and tested.

[0289] Flickr30K is a widely used benchmark dataset containing 31,783 images from Flickr. Each image in Flickr30K is accompanied by five manually annotated sentences, providing rich context. Flickr30K is divided into three subsets: 29,783 images for training, 1,000 images for validation, and 1,000 images for testing.

[0290] The MSCOCO dataset consists of 123,287 images, each annotated with five sentences. We split it into three subsets: 113,287 images for training, 5,000 for validation, and 5,000 for testing. The evaluation involves a challenging setting called MSCOCO (5K), where the proposed method is directly tested on the entire 5,000-image set. To ensure the reliability of the results, five validation runs are performed on the 1,000 test images, and the results are averaged to produce comprehensive performance metrics.

[0291] During training, the proposed model was trained using the PyTorch library on a single NVIDIA GeForce RTX 3090 GPU. The Adam optimizer was used with a mini-batch size of 128 (i.e., B=128), and the model was trained for 30 epochs. The learning rate was set to 0.0002 for the first 15 epochs and then decayed by 10%, i.e., 0.00002, in the remaining epochs. To determine the best model, its performance on the validation set was evaluated at the end of each epoch, and the model with the highest R@sum value was selected. The dimension of the visual features was set to 2048 ( = 2048, the number of visual markers is 36 ( =36). The dimension of text features is set to 300 ( =300). Convert visual and text features to 1024-dimensional feature space ( =1024).

[0292] Meanwhile, in the semantic concept construction, the number of image-text pairs in Flickr30K is 31,783 ( =31,783). The number of image-text pairs in MSCOCO is 123,287 ( =123,287). We set the number of visual concepts and textual concepts to ( =400) and ( = 300. The feature dimension of the visual semantic concept is 1000 ( =1000), the feature dimension of text semantic concepts is 300 ( =300).

[0293] And set the parameters, the self-attention mechanism is set to 1024 ( = 1024). For the nearest neighbor layer, we select the first 10 nearest neighbors (n = 10). In order to adaptively filter weak co-occurrences, the threshold z of the cumulative distribution function is set to 0.7, and the threshold of the co-occurrence matrix is ​​set to , , , , , are adaptively set to 0.0604 ( )、0.0165( )、0.0193( )、0.0189( )、0.0342( )、0.0325( ). The number of layers of GCN is set to 2 (L=2). For visual and text feature enhancement, λ=10 is set in equations (11) and (12). In addition, the function EFE ( ) are set to 0.75 for the fragment layer, instance layer, and neighbor layer, respectively. )、0.65( ) and 0.85 ( ).

[0294] Based on the above conditions, training is carried out, and after the training is completed, it is tested.

[0295] Mutual retrieval tests are conducted on the Flickr30K and MSCOCO test sets. Obviously, the proposed SFE method outperforms the state-of-the-art methods on both datasets:

[0296] On the Flickr30K dataset, SFE achieves significant improvements over the previous best-performing method, achieving 13.1% and 5.2% performance improvements over ESL with Bi-GRU and NUIF with BERT, respectively.

[0297] On the MSCOCO (1K) dataset (consisting of 1,000 images from the MSCOCO benchmark), SFE continues to demonstrate state-of-the-art performance. Compared to the previous best method (NUIF) using Bi-GRU, SFE achieves a 14.0% improvement in overall performance, and compared to the previous best method (NUIF) using BERT, SFE achieves a 5.9% improvement in overall performance. Furthermore, on the MSCOCO (5K) dataset, SFE outperforms the previous best method (IMEB) using Bi-GRU and the previous best method (ESL) using BERT by 6.3% and 2.2%, respectively, further demonstrating its superior performance.

[0298] In summary, SFE achieves state-of-the-art retrieval performance on both benchmark datasets, demonstrating the effectiveness of its step-by-step feature enhancement strategy. The SFE proposed in this example surpasses existing methods by introducing a two-step step-by-step feature enhancement network that propagates semantic information from the outer ESC to the inner ISC. Furthermore, the proposed SFE network introduces cross-modal ESC for the first time, effectively leveraging its potential to achieve significant improvements on the ITR task.

[0299] Unlike DSRAN, VSRN++, and RRTC, which focus on mining unimodal ISC to enhance feature representations by exploring spatial location, contextual relationships, and syntactic connections, SCAN, CAMP, NAAF, and ESL focus on mining cross-modal ISC through more refined cross-modal interactions. In contrast, hybrid unimodal and cross-modal ISC methods such as HREM, SGRAF, and BOOM fully exploit both types of ISC, thereby improving ITR performance. Furthermore, the proposed SFE outperforms the aforementioned three methods by leveraging ESC to reveal external semantic patterns such as inter-pair relationships within the corpus. By analyzing the global data distribution, ESC provides a comprehensive perspective, enriching the intrinsic semantics of individual pairs.

[0300] Compared to CSRC, DRCE, and NUIF, which rely solely on extracting unimodal ESC (e.g., neighbor connections and co-occurrence frequencies), SFE shows superior performance. This advantage stems from two key factors: (1) Unimodal ESC ignores cross-modal correlations, which are crucial for bridging the heterogeneous gap. In contrast, this work pioneers the exploration of cross-modal ESC and fully utilizes its potential to significantly improve ITR accuracy. (2) By fusing the semantic propagation of cross-modal ESC and ISC, SFE significantly improves retrieval performance.

[0301] Furthermore, to validate the adaptability of the SFE approach to pre-trained models, experiments were conducted using the CLIP model (ViT-L / 14) as the encoder. Compared to fine-tuning CLIP, our model consistently outperforms the baseline on both datasets, demonstrating the effectiveness of SFE in leveraging large-scale pre-trained models.

[0302] In this embodiment, an ablation experiment is also performed to verify the retrieval performance.

[0303] This embodiment adopts a step-by-step feature enhancement network. In image-text mutual retrieval, the step-by-step feature enhancement network is divided into two steps to mine internal and external semantic clues respectively. In step S6-1, the focus is on extracting multi-level co-occurrence relationships in the external layer from a large-scale image-text corpus. In contrast, step S6-2 emphasizes mining cross-modal context in the inner layer. From the analysis of the results, we observed that the joint effect of the two steps is the most effective, and step S6-1 alone is better than S6-2 in improving ITR performance. This is because step S6-1 extracts valuable semantic information from the large-scale image-text corpus in the outer layer, enhancing the migration of ESC from the outer layer to the inner layer ISC.

[0304] In this embodiment, multi-level co-occurrence mining is adopted, and the co-occurrence relationship is divided into three layers: fragment layer, instance layer and neighbor layer. Specifically, the fragment layer focuses on unimodal co-occurrence, while the instance layer and neighbor layer focus on cross-modal co-occurrence. Compared with other forms of co-occurrence, it is concluded that the co-occurrence of all three layers, as well as their individual contributions, have a positive impact on the ITR task. Specifically, the neighbor layer has the most significant impact, followed by the instance layer, and the impact of both is greater than that of the fragment layer. This is because the neighbor layer captures a wide range of external semantic clues from a large-scale corpus, which is crucial for bridging the heterogeneous gap. In addition, the simultaneous integration of the co-occurrence of the fragment layer, instance layer and neighbor layer is the best way to fully utilize their complementary advantages and comprehensively capture the semantic clues behind the relationship between image and text.

[0305] In this embodiment, intra-modal enhancement is denoted as "same modality", cross-modal enhancement is denoted as "different modalities", and the combination of the two is denoted as "same modality and different modalities". Compared with other methods, cross-modal enhancement is superior to intra-modal enhancement. This is because cross-modal enhancement directly establishes semantic connections between the visual modality and the textual modality, overcoming the limitation of intra-modal information lacking inter-modal interaction. In addition, the combination of the two enhancement methods in this embodiment achieved the best results. This is due to the simultaneous use of intra-modal and cross-modal enhancement, which fully utilizes all available semantic clues to improve retrieval performance.

[0306] This example also conducts visualization experiments on the MSCOCO test set using the t-SNE tool. Specifically, 10 semantic concepts are selected from the 80 semantic concepts available in MSCOCO to be compared step by step with the visual and text features obtained by the baseline method CVSE. The results are shown in Figure 2. Figure 4 、 Figure 5 、 Figure 6 and Figure 7 As shown, among them, Figure 4 、 Figure 5 Points of the same color represent visual features belonging to the same semantic concept. Figure 6 、 Figure 7Points of the same color represent text feature representations belonging to the same semantic concept. Figure 4 and Figure 5 In the figure, circle ① is the image area about the “cup”; circle ② is the image area about the “ball”; circle ③ is the image area about the “tennis ball”; circle ④ is the image area about the “dog”; circle ⑤ is the image area about the “chair”; circle ⑥ is the image area about the “table”. Figure 6 and Figure 7 In the picture, circle ① Cup: A very small and adorable cat is in a cup. A sandwich and fruit are placed on a plate with tea in a cup next to it. A cookie is placed on a plate with two cups of coffee next to it. Circle ② Ball: A woman with a racket is playing ball. A group of children are holding rackets and tennis balls. A man playing baseball is swinging his arms wildly. Circle ③ Tennis: A boy is swinging a tennis racket. A man is holding a tennis racket on the court. A man is playing tennis on a blue court. Circle ④ Dog: A man is riding a bicycle with a dog. A young woman has a puppy in her purse. A golden retriever is standing in the snow. Circle ⑤ Chair: A group of children are eating in chairs. A family is watching TV in chairs. A book is placed on a chair. Circle ⑥ Table: Some people are eating and drinking next to a red figure. A group of people are eating at a table. A pizza place is served, with several men sitting around pizzas on a folding table.

[0307] By analyzing the visualization results, we can draw the following conclusions:

[0308] Analysis of the visualization results shows that the baseline method cannot effectively distinguish the categories of semantic concepts. Samples from different semantic concepts have overlapping feature distributions. Five overlapping semantic concept pairs are analyzed in detail. For example, Figure 4 and Figure 5 (ball, tennis), (dog, person), and (cup, table), and Figure 6 and Figure 7 This problem arises because the unimodal ESC extracted by the baseline fails to capture cross-modal co-occurrence relationships, limiting its ability to clearly separate semantic concepts.

[0309] Compared with the baseline, the SFE network effectively distinguishes different semantic concepts, and the feature distributions of these concepts have clear boundaries. This improvement is attributed to the introduction of cross-modal ESC, which strengthens the correlation between semantically related concepts while further weakening the correlation between semantically unrelated concepts.

[0310] Compared with the feature segmentation of the visual modality, both the baseline and the proposed method show better semantic concept differentiation in the text modality. This is because visual semantic concepts are more complex and diverse, while feature learning of textual semantic concepts is relatively easy.

[0311] In summary, the multimodal data processing method and system based on the step-by-step feature enhancement network in this embodiment solves the problems existing in existing retrieval technology, can fully utilize the semantic information from a large-scale image-text corpus, effectively bridge the heterogeneous gap, integrate the co-occurrence of the fragment layer, instance layer, and neighbor layer, fully utilize their complementary advantages, and comprehensively capture the semantic clues behind the relationship between images and texts. It can fully tap the potential of cross-modal ESC, improve the accuracy of similarity measurement of data in different modalities, and thus improve the accuracy of ITR.

[0312] Example 2

[0313] The purpose of this embodiment is to provide a multimodal data processing system based on a step-by-step feature enhancement network, including:

[0314] A data acquisition module, configured to acquire first modal data and second modal data;

[0315] A feature extraction module, configured to extract features of the first modal data and features of the second modal data;

[0316] An attention enhancement module is used to enhance the features of the first modality data and the second modality data through a self-attention mechanism to obtain enhanced first modality global-level features and enhanced second modality global-level features;

[0317] A semantic construction module, configured to construct a first modality semantic concept set and a second modality semantic concept set based on the first modality data corpus and the second modality data corpus, respectively;

[0318] A multi-layer co-occurrence module is used to mine multi-layer co-occurrence relationships between the first modality semantic concept set and the second modality semantic concept set;

[0319] The feature enhancement module is used to further enhance the enhanced first modality global-level features and the enhanced second modality global-level features based on the multi-layer co-occurrence relationship through a step-by-step feature enhancement network to obtain multi-layer first modality data features and multi-layer second modality data features, and fuse them respectively to obtain the final data features of the first modality data and the second modality data.

[0320] Based on a provided multimodal data processing system based on a step-by-step feature enhancement network, the method steps in Example 1 are implemented.

[0321] Example 3

[0322] This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the method in the first embodiment are implemented.

[0323] Example 4

[0324] This embodiment provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps of the method in the first embodiment are performed.

[0325] The steps involved in the apparatus of the above embodiment correspond to those of the method embodiment 1. For detailed implementation, please refer to the relevant description of embodiment 1. The term "computer-readable storage medium" should be understood to mean a single medium or multiple media containing one or more instruction sets; it should also be understood to include any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and causing the processor to perform any method of the present invention.

[0326] Those skilled in the art will appreciate that the modules or steps of the present invention described above can be implemented using a general-purpose computer device. Alternatively, they can be implemented using program code executable by a computing device, which can then be stored in a storage device and executed by the computing device. Alternatively, they can be fabricated into separate integrated circuit modules, or multiple modules or steps can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0327] Although the above describes the specific embodiments of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without any creative work are still within the scope of protection of the present invention.

Claims

1. A multimodal data processing method based on a step-by-step feature enhancement network, characterized in that: include: Acquiring first modal data and second modal data; Extracting first modality data features and second modality data features; Performing feature enhancement on the first modality data features and the second modality data features through a self-attention mechanism to obtain enhanced first modality global-level features and enhanced second modality global-level features; Based on the first modality data corpus and the second modality data corpus, constructing a first modality semantic concept set and a second modality semantic concept set respectively; Conduct multi-level co-occurrence relationship mining on the first modality semantic concept set and the second modality semantic concept set; Multi-level co-occurrence mining is used, and the co-occurrence relationship is divided into three layers: fragment layer, instance layer, and neighbor layer. The fragment layer focuses on single-modal co-occurrence, while the instance layer and neighbor layer focus on cross-modal co-occurrence; the co-occurrence of the fragment layer, instance layer, and neighbor layer is integrated; Based on the multi-level co-occurrence relationship, the enhanced first modality global-level features and the enhanced second modality global-level features are further enhanced through a step-by-step feature enhancement network to obtain multi-layer first modality data features and multi-layer second modality data features, and the final data features of the first modality data and the second modality data are respectively fused; The specific process is: Performing external semantic enhancement on the enhanced first modality global-level features and the enhanced second modality global-level features through multi-layer co-occurrence relationships in the external layer; The enhanced first modality global-level features and the enhanced second modality global-level features are internally semantically enhanced through the cross-modal context of the internal layer.

2. The multimodal data processing method based on a step-by-step feature enhancement network according to claim 1, wherein: Extract the first modal data features. The specific process is as follows: Utilize an object detection model with bottom-up attention to extract the target region; A deep convolutional network is used to extract the first modality data features of the target area.

3. The multimodal data processing method based on a step-by-step feature enhancement network according to claim 1, wherein: Extracting the second modal data features, specifically, encoding the second modal data using an encoder of a deep neural network to obtain the second modal data features.

4. The multimodal data processing method based on a step-by-step feature enhancement network according to claim 1, wherein: Multi-layer co-occurrence relationship mining is performed on the first modal semantic concept set and the second modal semantic concept set. The multi-layer co-occurrence relationship includes the fragment layer, instance layer, and neighbor layer. The specific process is as follows: Perform unimodal co-occurrence relationship mining at the fragment level and generate the first modal initial features and second modal co-occurrence relationship matrix at the fragment level; Perform cross-modal co-occurrence relationship mining at the instance level and generate the first modal initial features and second modal co-occurrence relationship matrix at the instance level; Perform cross-modal co-occurrence relationship mining at the neighboring layer and generate the first modal initial visual features and the second modal co-occurrence relationship matrix of the neighboring layer; Adaptive weak co-occurrence relationship filtering is performed on the first modal initial features and the second modal co-occurrence relationship matrix to obtain a co-occurrence relationship matrix for weak co-occurrence relationship filtering.

5. The multimodal data processing method based on a step-by-step feature enhancement network according to claim 1, wherein: The enhanced first modality global-level features and the enhanced second modality global-level features are subjected to external semantic enhancement through the multi-layer co-occurrence relationship of the external layer, specifically: Constructing a first modality and a second modality weighted undirected graph based on the first modality semantic concept set and the second modality semantic concept set; Use graph convolutional neural networks to update the node embeddings of the first modality and second modality weighted undirected graphs layer by layer, and obtain the node embedding feature matrices of the first modality graph and the second modality graph respectively; Based on the node embedding feature matrices of the first modal graph and the second modal graph, calculating the first modal feature enhanced by the first modal co-occurrence relationship, the first modal feature enhanced by the second modal co-occurrence relationship, the second modal feature enhanced by the first modal co-occurrence relationship, and the second modal feature enhanced by the second modal co-occurrence relationship; Based on the first modality co-occurrence relationship, the first modality feature is enhanced, the second modality feature is enhanced, the first modality feature is enhanced, and the second modality feature is enhanced based on the second modality co-occurrence relationship. The feature enhancement function is used to perform layered first modality feature enhancement and second modality feature enhancement to obtain enhanced first modality features and second modality features in each layer. Obtaining enhanced data features of the first modal data based on the first modal data features and the enhanced first modal features of each layer; Based on the second modality data features and the enhanced second modality features of each layer, enhanced data features of the second modality data are obtained.

6. The multimodal data processing method based on a step-by-step feature enhancement network according to claim 1, wherein: The enhanced first modality global-level features and the enhanced second modality global-level features are semantically enhanced through the cross-modal context of the internal layer. The specific process is as follows: Calculating a similarity matrix between the first modality label and the second modality label based on the enhanced data features of the first modality data and the enhanced data features of the second modality data; According to the similarity matrix, the first modality attention matrix and the second modality attention matrix are calculated; Obtaining an updated feature representation of the first modality data according to the first modality attention matrix and the enhanced data features of the second modality data; Obtaining an updated feature representation of the second modality data according to the second modality attention matrix and the enhanced data features of the first modality data; The enhanced data features of the first modal data and the updated first modal data feature representation are concatenated column by column, and passed through a fully connected layer to obtain the first modal data features enhanced with cross-modal information; The enhanced data features of the second modality data and the updated second modality data feature representation are concatenated column by column and passed through a fully connected layer to obtain the second modality data features enhanced by cross-modal information.

7. A multimodal data processing system based on a step-by-step feature enhancement network, characterized in that: Implementing a multimodal data processing method based on a step-by-step feature enhancement network as described in any one of claims 1 to 6, comprising: A data acquisition module, configured to acquire first modal data and second modal data; A feature extraction module, configured to extract features of the first modal data and features of the second modal data; An attention enhancement module is used to enhance the features of the first modality data and the second modality data through a self-attention mechanism to obtain enhanced first modality global-level features and enhanced second modality global-level features; A semantic construction module, configured to construct a first modality semantic concept set and a second modality semantic concept set based on the first modality data corpus and the second modality data corpus, respectively; A multi-layer co-occurrence module is used to mine multi-layer co-occurrence relationships between the first modality semantic concept set and the second modality semantic concept set; The feature enhancement module is used to further enhance the enhanced first modality global-level features and the enhanced second modality global-level features based on the multi-layer co-occurrence relationship through a step-by-step feature enhancement network, obtain multi-layer first modality data features and multi-layer second modality data features, and fuse them respectively to obtain the final data features of the first modality data and the second modality data.

8. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are performed.

Citation Information

Patent Citations

  • Gradual semantic aggregation and structured cognitive enhancement-based image-text matching method

    CN119397048A