Multi-modal data processing method and system based on step-by-step feature enhancement network

By adopting step-by-step feature enhancement network in multimodal data processing, combining external and internal semantic clues, the problem of not being able to effectively cross the heterogeneous gap between modes in the prior art is solved, and more efficient multimodal data processing and image text mutual retrieval performance is achieved.

CN120104843AActive Publication Date: 2025-06-06SHANDONG HI SPEED COMPANY
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510577614.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-06-06
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

Existing multimodal data processing technology fails to fully utilize the semantic information of corpus data of different modalities, resulting in the inability to effectively cross the heterogeneous gap between modalities, affecting the accuracy of multimodal data processing.

Method used

Using a method based on step-by-step feature enhancement network, the semantic information flow is guided to flow from the external layer to the internal layer in step by step, combining the cross-modal enhancement of external semantic cues and the context mining of internal semantic cues to achieve step-by-step feature enhancement of multimodal data features.

Benefits of technology

It effectively improves the multimodal data processing performance, improves the accuracy of data similarity measurements of different modal data, and significantly improves the retrieval performance of image text mutual retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104843A_ABST
    Figure CN120104843A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal data processing method and system based on a step-by-step feature enhancement network, relates to the technical field of multi-modal data processing, and aims to solve the problems that semantic information of a multi-modal corpus is not fully utilized, the potential of a cross-modal ESC is not fully mined, and the accuracy of multi-modal data processing is low. The method comprises the following steps: extracting first modal and second modal data features, and performing self-attention mechanism feature enhancement on the first modal and second modal data features; based on different modal data corpora, respectively constructing semantic concepts of the first modal data and the second modal data, and performing multi-layer co-occurrence relationship mining; and performing further feature enhancement on the features after the self-attention mechanism enhancement through a step-by-step feature enhancement network, and performing fusion to obtain final data features of the first mode and the second mode. According to the method, a step-by-step feature enhancement network is provided, and the multi-modal data processing performance is effectively improved by combining cross-modal enhancement of external semantic clues and context mining of internal semantic clues.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of cross-modal image-text mutual retrieval, and in particular, relates to a multimodal data processing method and system based on a step-by-step feature enhancement network. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] In today's big data era, heterogeneous data processing refers to the processing and analysis of information from multiple perceptual channels or data forms. These modalities can include text, images, audio, video, etc. Data from different modalities have different characteristics and representations. How to effectively fuse and process these heterogeneous data is a challenge. For example, image-text inter-retrieval (ITR) is a basic visual language task that aims to search for images related to a given text query or retrieve text related to a given image query. The key task of ITR is to learn the cross-modal similarity between images and text. A major challenge facing ITR is that data from different modalities (images and text) are encoded using different encoders, which leads to differences between these modalities. This encoding difference triggers a "heterogeneous gap", making it difficult to directly compare and align the features of the two modalities in a unified feature space, thereby increasing the complexity of multimodal data processing.

[0004] There are inherent inconsistencies between the feature representations of multimodal data, namely the "heterogeneous gap". In order to bridge the heterogeneous gap in multimodal data processing tasks, capturing as many semantic clues as possible is crucial to establish effective associations between different modal data. These semantic clues can be divided into two categories: Internal semantic cues (ISC) refer to the semantic concepts and their relationships within different modal data. ISC highlights the intrinsic semantic structure between the two, including the intra-modal correlation within each modality and the inter-modal correlation between different modalities.

[0005] External semantic cues (ESC) refer to semantic concepts and their relations derived from a large-scale corpus of data from different modalities, rather than data from a single modality. ESC captures external semantic patterns, including associations between paired samples in the corpus, such as neighbor relations and co-occurrence frequencies. By analyzing global data distribution patterns, ESC provides a global perspective for semantic understanding and supplements the intrinsic semantic information in different modal data.

[0006] Existing research on multimodal data processing mainly focuses on extracting ISC from data of different modalities by analyzing intra-modal and inter-modal semantics. Current methods are generally divided into three categories: (1) mining ISC in one modality, such as VSRN and CAMERA; (2) mining ISC across different modalities, such as SCAN, ESL and MPARN; (3) hybrid methods that combine (1) and (2), such as HREM and BOOM. However, these methods mainly focus on ISC in single modality data, while ignoring ESC existing in a wider corpus of different modal data, which limits their comprehensive bridging of the heterogeneous gap.

[0007] In order to address the cognitive limitations of focusing only on ISC in different modal data, recent studies have begun to focus on ESC outside of different modal data. Research on ESC in different modal data processing can be divided into two categories: (1) mining ESC from neighbors, such as DSP, TBNN, NAN, LeaPRR; (2) mining ESC from common sense knowledge, such as CVSE, VCM and CSRC. Despite this, these methods are still mainly based on single-modal ESC, that is, extracting ESC from each modality independently, and fail to fully explore the potential of cross-modal ESC. This deficiency limits its ability to cross the heterogeneous gap and greatly reduces the accuracy of multimodal data processing.

[0008] In summary, the inventors found that the existing multimodal data processing technology has not yet fully utilized the semantic information from the corpus of different modal data, which limits its comprehensive crossing of the heterogeneous gap, especially beyond the clues of single modal data, and only independently extracts the ESC of each modality, failing to fully tap the potential of cross-modal ESC, affecting the accuracy of the similarity measurement of different modal data, thereby reducing the accuracy of multimodal data processing. Summary of the invention

[0009] In order to overcome the deficiencies of the above-mentioned prior art, the present invention provides a multimodal data processing method and system based on a step-by-step feature enhancement network. Through the step-by-step feature enhancement network, a semantic propagation path is established to guide the semantic information flow from the external layer to the internal layer in steps, thereby realizing step-by-step feature enhancement from the external layer to the internal layer. By combining the cross-modal enhancement of external semantic clues and the context mining of internal semantic clues, the multimodal data processing performance is effectively improved.

[0010] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions: A first aspect of the present invention provides a multimodal data processing method based on a step-by-step feature enhancement network, comprising: Acquire first modality data and second modality data; Extracting first modality data features and second modality data features; Performing feature enhancement on the first modality data features and the second modality data features through a self-attention mechanism to obtain enhanced first modality global-level features and enhanced second modality global-level features; Based on the first modality data corpus and the second modality data corpus, constructing a first modality semantic concept set and a second modality semantic concept set respectively; Conduct multi-layer co-occurrence relationship mining on the first modality semantic concept set and the second modality semantic concept set; Based on the multi-layer co-occurrence relationship, the enhanced first modality global-level features and the enhanced second modality global-level features are further enhanced through a step-by-step feature enhancement network to obtain multi-layer first modality data features and multi-layer second modality data features, and the final data features of the first modality data and the second modality data are respectively fused.

[0011] As an implementation method, the first modal data feature is extracted, and the specific process is as follows: Utilize an object detection model with bottom-up attention to extract the target region; A deep convolutional network is used to extract the first modality data features of the target area.

[0012] As an implementation method, the second modality data features are extracted by, specifically, encoding the second modality data using an encoder of a deep neural network to obtain the second modality data features.

[0013] As an implementation method, multi-layer co-occurrence relationship mining is performed on the first modal semantic concept set and the second modal semantic concept set, wherein the multi-layer co-occurrence relationship includes a fragment layer, an instance layer, and a neighbor layer. The specific process is as follows: Perform unimodal co-occurrence relationship mining at the fragment layer, and generate the first modal initial features and the second modal co-occurrence relationship matrix at the fragment layer; Perform cross-modal co-occurrence relationship mining at the instance level, and generate the first modal initial features and the second modal co-occurrence relationship matrix at the instance level; Perform cross-modal co-occurrence relationship mining at the neighbor layer, and generate the first modal initial visual features and the second modal co-occurrence relationship matrix of the neighbor layer; Adaptive weak co-occurrence relationship filtering is performed on the first modal initial features and the second modal co-occurrence relationship matrix to obtain a co-occurrence relationship matrix for weak co-occurrence relationship filtering.

[0014] As an implementation method, the enhanced first modality global-level features and the enhanced second modality global-level features are further enhanced through a step-by-step feature enhancement network to obtain multi-layer first modality data features and multi-layer second modality data features, and the final data features of the first modality data and the second modality data are respectively fused. The specific process is: External semantic enhancement is performed on the enhanced first modality global-level features and the enhanced second modality global-level features through multi-layer co-occurrence relationships in the external layer; The enhanced first modality global-level features and the enhanced second modality global-level features are internally semantically enhanced through the cross-modal context of the internal layer.

[0015] As an implementation method, the enhanced first modality global-level features and the enhanced second modality global-level features are externally semantically enhanced through the multi-layer co-occurrence relationship of the external layer, specifically: Based on the first modality semantic concept set and the second modality semantic concept set, construct a first modality and a second modality weighted undirected graph; Using a graph convolutional neural network, the node embeddings of the first modality and the second modality weighted undirected graphs are updated layer by layer to obtain node embedding feature matrices of the first modality graph and the second modality graph respectively; Based on the node embedding feature matrix of the first modal graph and the second modal graph, the first modal feature enhanced by the first modal co-occurrence relationship, the first modal feature enhanced by the second modal co-occurrence relationship, the second modal feature enhanced by the first modal co-occurrence relationship, and the second modal feature enhanced by the second modal co-occurrence relationship are calculated; Based on the first modality co-occurrence relationship, the first modality feature is enhanced, the second modality feature is enhanced, the first modality feature is enhanced, and the second modality feature is enhanced by the first modality co-occurrence relationship, and the second modality feature is enhanced by the second modality co-occurrence relationship, and the feature enhancement function is used to perform layered first modality feature enhancement and second modality feature enhancement to obtain enhanced first modality features and second modality features of each layer; Obtaining enhanced data features of the first modal data based on the first modal data features and the first modal features enhanced at each layer; Based on the second modality data features and the enhanced second modality features of each layer, enhanced data features of the second modality data are obtained.

[0016] As an implementation method, the enhanced first modality global-level features and the enhanced second modality global-level features are internally semantically enhanced through the cross-modal context of the internal layer. The specific process is as follows: Calculating a similarity matrix between the first modality tag and the second modality tag based on the augmented data features of the first modality data and the enhanced data features of the second modality data; According to the similarity matrix, calculating the first modality attention matrix and the second modality attention matrix; Obtaining an updated feature representation of the first modality data according to the first modality attention matrix and the enhanced data features of the second modality data; Obtaining an updated second modality data feature representation according to the second modality attention matrix and the enhanced data features of the first modality data; The enhanced data features of the first modal data and the updated first modal data feature representation are concatenated column by column, and passed through a fully connected layer to obtain the first modal data features enhanced by using cross-modal information; The enhanced data features of the second modality data and the updated second modality data feature representation are concatenated column by column and passed through a fully connected layer to obtain the second modality data features enhanced by cross-modality information.

[0017] A second aspect of the present invention provides a multimodal data processing system based on a step-by-step feature enhancement network, comprising: A data acquisition module, used to acquire first modality data and second modality data; A feature extraction module, used to extract first modality data features and second modality data features; An attention enhancement module, used for performing feature enhancement on the first modality data features and the second modality data features through a self-attention mechanism to obtain enhanced first modality global-level features and enhanced second modality global-level features; A semantic construction module, used to construct a first modality semantic concept set and a second modality semantic concept set based on the first modality data corpus and the second modality data corpus, respectively; A multi-layer co-occurrence module is used to mine multi-layer co-occurrence relationships between the first modality semantic concept set and the second modality semantic concept set; The feature enhancement module is used to further enhance the enhanced first modality global-level features and the enhanced second modality global-level features based on the multi-layer co-occurrence relationship through a step-by-step feature enhancement network to obtain multi-layer first modality data features and multi-layer second modality data features, and fuse them respectively to obtain the final data features of the first modality data and the second modality data.

[0018] The third aspect provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the program.

[0019] A fourth aspect provides a computer-readable storage medium having a computer program stored thereon, which executes the steps of the above method when executed by a processor.

[0020] One or more of the above technical solutions have the following beneficial effects: The step-by-step feature enhancement network SFE in this embodiment establishes a semantic propagation path, guides the semantic information flow from the external layer to the internal layer step by step, and realizes step-by-step enhancement from the external layer to the internal layer. By combining the cross-modal enhancement of external semantic clues and the context mining of internal semantic clues, the potential of cross-modal ESC is fully explored, and the accuracy of the similarity measurement of data of different modalities is improved. SFE effectively improves the retrieval performance.

[0021] In this embodiment, a multi-level co-occurrence relationship mining method guided by ESC is used for feature enhancement, including co-occurrence relationships at the fragment level, instance level, and neighbor level. These co-occurrence relationships significantly enhance the modal interaction in the outer layer, and the cross-modal context mining method guided by ISC further enhances the modal interaction in the inner layer by combining different modal data features with their corresponding context features.

[0022] In this embodiment, based on the three-layer co-occurrence relationship, a positive impact is exerted on multimodal data processing, the co-occurrence relationships of the fragment layer, instance layer and neighbor layer are integrated, the semantic information from the large-scale modal corpus is fully utilized, and their complementary advantages are fully utilized to comprehensively capture the semantic clues behind different modal relationships.

[0023] In this embodiment, intra-modal and cross-modal feature enhancement are combined to overcome the limitation of intra-modal information lacking inter-modal interaction and make full use of all available semantic clues to improve retrieval performance.

[0024] Through the method of this embodiment, future research in the field of multimodal data processing can further explore the semantic clues between different modalities (such as video and audio). These works can effectively expand the scope of semantic clues and lead to more effective feature enhancement, ultimately improving the performance of multimodal data processing.

[0025] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0027] Figure 1 A schematic diagram of a step-by-step feature enhancement network framework of the first embodiment; Figure 2 This is a framework diagram of an external feature enhancement module based on multi-level co-occurrence relationship mining in the first embodiment of the present invention; Figure 3 This is a framework diagram of an internal feature enhancement module for cross-modal context mining in the first embodiment of the present invention; Figure 4 The result of visual feature representation of the baseline by T-SNE in the first embodiment on the MSCOCO test set is shown; Figure 5 The result of visual feature representation of SFE by T-SNE in the first embodiment on the MSCOCO test set is shown; Figure 6 The result of visualizing text features of the baseline using T-SNE on the MSCOCO test set of the first embodiment is shown; Figure 7 The result of visualizing text feature representation of SFE by T-SNE in the first embodiment on the MSCOCO test set is shown. DETAILED DESCRIPTION

[0028] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.

[0029] It should be noted that the terms used herein are for describing specific embodiments only and are not intended to be limiting of exemplary embodiments according to the present invention.

[0030] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.

[0031] Embodiment 1 This embodiment discloses a multimodal data processing method based on a step-by-step feature enhancement network.

[0032] In order to more clearly illustrate the present embodiment, a multimodal data processing implementation process based on a step-by-step feature enhancement network can be specifically described as follows: A multimodal data processing method based on a step-by-step feature enhancement network, comprising: S1, obtaining first modality data and second modality data; S2, extracting first modality data features and second modality data features; S3. Performing feature enhancement on the first modality data features and the second modality data features through a self-attention mechanism to obtain enhanced first modality global-level features and enhanced second modality global-level features; S4. Based on the first modal data corpus and the second modal data corpus, construct a first modal semantic concept set and a second modal semantic concept set respectively; S5, performing multi-layer co-occurrence relationship mining on the first modal semantic concept set and the second modal semantic concept set; S6. Based on the multi-layer co-occurrence relationship, the enhanced first modality global-level features and the enhanced second modality global-level features are further enhanced through a step-by-step feature enhancement network to obtain multi-layer first modality data features and multi-layer second modality data features, and the final data features of the first modality data and the second modality data are respectively fused.

[0033] In this embodiment, the multimodal data processing method based on the step-by-step feature enhancement network can be applied to data processing between different modalities, such as image and text mutual retrieval. The following takes image and text mutual retrieval as an example for analysis and explanation.

[0034] like Figure 1 As shown, in step S1, first modality data and second modality data are acquired.

[0035] In this embodiment, the target image and the target text are obtained from the Visual Genome dataset.

[0036] like Figure 1 As shown, in step S2, the first modality data features and the second modality data features are extracted.

[0037] S2-1. Extract first modal data features.

[0038] In this embodiment, the first modal data feature is extracted, and the specific process is as follows: (1) Utilize an object detection model with bottom-up attention to extract the target region.

[0039] In this implementation, in the image-text mutual retrieval process, the pre-trained object detection model Faster R-CNN with bottom-up attention is used to extract the top-ranked and highest-confidence salient regions, i.e., visual markers, from each image on the Visual Genome dataset.

[0040] (2) Use a deep convolutional network to extract the first modal data features of the target area.

[0041] In this implementation, in the image-text mutual retrieval process, the deep convolutional network ResNet-101 is used to extract the features of the above regions, and the fully connected (FC) layer is used as the visual feature to convert it into a d-dimensional feature space. The formula is: (1) in, is the coded image area , and are the parameters to be learned.

[0042] The visual features of the image region are represented as

[0043] (2) in is the number of image regions, represents the visual features of the first image region, Indicates Visual features of an image region.

[0044] S2-2. Extracting features of the second modality data.

[0045] In this embodiment, in the image-text mutual retrieval process, an encoder based on a deep neural network, such as Bi-GRU and Bert, is used to encode text words. , which is converted to Dimension, the formula is: (3) in, Represents the text word To encode, and are learnable parameters.

[0046] The text features of text words are represented as: (4) in, is the number of words in the text, The text feature representing the first text word, Indicates The text features of the text words.

[0047] If data processing is performed on other modal data, modal data features are extracted for specific modal data.

[0048] After the above steps, the initial visual features of the target image and the initial text features of the target text are extracted, which provides a data basis for the subsequent attention enhancement features and step-by-step feature enhancement grid to further enhance features.

[0049] like Figure 1 As shown, in step S3, the first modality data features and the second modality data features are feature enhanced through a self-attention mechanism to obtain enhanced first modality global-level features and enhanced second modality global-level features.

[0050] In this embodiment, a self-attention mechanism is used to fully utilize the complementary semantic relationships within each modality to enhance feature representation.

[0051] (1) Obtain enhanced first modal global-level features.

[0052] In this implementation, in the image-text mutual retrieval process, the visual tag feature As key terms and value terms, the global visual feature vector is the query term of the self-attention mechanism:

[0053] (5)

[0054] in, , K and Represent the query vector, key vector and value vector in the self-attention mechanism respectively, , and is the weight matrix that needs to be trained.

[0055] The self-attention mechanism calculates the query vector by calculating the dot product similarity and key vector , thereby obtaining enhanced global visual features, the formula is: (6) in, , represents the enhanced global visual features, Indicates that the initial visual tag feature R learns context information through the self-attention module, Represents the dimension of attention features after dimensionality reduction.

[0056] (2) Obtain enhanced second modal global-level features.

[0057] In this embodiment, in the process of image-text mutual retrieval, the method in (1) is adopted to determine the features of the text and the features of the text mark. As key and value items, the global text feature vector As query items, according to formula (6), they are input into function to obtain enhanced global-level text features .

[0058] Among them, if data processing is performed on other modal data, enhanced modal global-level features are obtained for specific modal data.

[0059] After the above steps, the global-level image features and text features enhanced by self-attention are obtained, which achieves preliminary enhancement of image features and text features and improves the accuracy of feature extraction.

[0060] like Figure 1 As shown, in step S4, based on the first modal data corpus and the second modal data corpus, a first modal semantic concept set and a second modal semantic concept set are constructed respectively.

[0061] In order to fully mine the semantic associations in the corpus, it is necessary to fully mine the various semantic concepts contained in the corpus and learn the semantic associations between them. The co-occurrence relationship between semantic concepts can well reflect the semantic associations between concepts.

[0062] In this embodiment, an image-text semantic concept set is constructed for image-text mutual retrieval.

[0063] (1) Construct a set of visual semantic concepts.

[0064] The images are divided into blocks, and each image is divided into 9 blocks on average. Due to the large number of images and the noise they contain, a single image or a single image block cannot directly represent the visual semantic concept.

[0065] For i image patch, calculate its distance to each centroid, and assign it to the cluster associated with the nearest centroid, as follows: (7) in, It is i The cluster index to which the image block belongs; Indicates i Image blocks Visual features; is the square of the Euclidean distance; It is k The characteristics of the centroid k The visual word of the centroid is represented as , Through the k The features of all image blocks in a cluster are averaged and the formula is: (8) in, is the feature set of the image patch belonging to the k-th cluster, and the centroid of the k-th class is defined as the visual word , express The features of the visual semantic concept set are expressed as: , Represents the total number of visual semantic concepts.

[0066] (2) Construct a set of text semantic concepts.

[0067] All words in the corpus are considered as candidate text semantic concepts. In order to alleviate the impact of word sparsity and irrelevant terms, infrequent words need to be filtered.

[0068] Specifically, 1) high-frequency words in the concept vocabulary are identified as text semantic concepts and divided into objects, movements, and attributes.

[0069] 2) According to the statistical frequency of concepts in the data set, keep the concepts (object, motion, attribute) in a balanced distribution ratio of (7:2:1) to obtain the text semantic concept set, denoted as ,in Represents the total number of semantic concepts in the text.

[0070] 3) Use word embedding method GloVe to extract text semantic concepts The characteristics of .

[0071] Among them, if data processing is performed on other modal data, a corresponding semantic concept set is constructed for the specific modal data.

[0072] like Figure 1 As shown, in step S5, multi-layer co-occurrence relationship mining is performed on the first modal semantic concept set and the second modal semantic concept set.

[0073] Multi-level co-occurrence relations include unimodal co-occurrence relations and cross-modal co-occurrence relations. These relations are divided into three layers: fragment layer, instance layer, and neighbor layer. Specifically, the fragment layer corresponds to unimodal co-occurrence relations, while the instance layer and neighbor layer correspond to cross-modal co-occurrence relations. For each layer, a semantic concept association graph is constructed, which encapsulates different co-occurrence relations. In each layer, we construct a co-occurrence matrix according to the corresponding co-occurrence relations and To model the relationship between different semantic concepts.

[0074] S5-1. Perform unimodal co-occurrence relationship mining at the fragment layer, and generate the first modal initial features and the second modal co-occurrence relationship matrix of the fragment layer.

[0075] Fragment Layer( )Unimodal co-occurrence mainly studies the relationship within a single modality, that is, the relationship within an image patch or a text word. Unimodal concept associations are constructed using visual or text semantic concepts independently at the fragment level.

[0076] In this embodiment, and denote the visual and textual co-occurrence frequency matrices respectively. Specifically, express and The number of times they co-occur in the corpus, and yes The total number of occurrences in the same corpus. Represents a given visual semantic concept In the presence of visual semantic concepts Probability of occurrence.

[0077] Algorithm 1 is used to generate the initial visual and initial text co-occurrence relationship matrices at the segment layer, expressed as and .

[0078] Algorithm 1: Generate the single-modal co-occurrence relationship matrix at the fragment level. The steps are: Input: image-text pairs , visual semantic concepts , text semantic concept .

[0079] Output: Visual and textual co-occurrence matrix of the initial segment layer , .

[0080] 1: Initialize the matrix to zero and .

[0081] 2: for each image do: 3: for each and do: 4: if and Also appears in the image Then: 5: ;

[0082] 6: Each element of is set to ; 7: for each text do: 8: for each and do: 9: if and Also appears in the text Then: 10: ;

[0083] 11: Each element of is set to ; 12: return , .

[0084] After the above steps, using data statistics methods, a matrix representing the co-occurrence relationship between visual semantic concepts is generated for the image modality at the fragment level , for the text modality, generates a matrix representing the co-occurrence relationship between text semantic concepts The statistics of the co-occurrence relationship of semantic concepts are conducive to mining the potential semantic associations between different concepts and improving the model's ability to understand semantic concepts. In addition, by analyzing the co-occurrence pattern, the model can more accurately capture contextual information and improve the results of retrieval and reasoning.

[0085] S5-2 performs cross-modal co-occurrence relationship mining at the instance layer, and generates the first modal initial features and the second modal co-occurrence relationship matrix at the instance layer.

[0086] At the instance level, the pairwise relationship between image and text is used as a constraint to mine cross-modal co-occurrence relationships. The instance-level cross-modal co-occurrence frequency matrix is ​​recorded as ,in, Indicates that the corpus and Frequency of co-occurrence in the same image-text pair.

[0087] Algorithm 2 is used to generate the initial visual and text co-occurrence relationship matrix at the instance level, expressed as , .

[0088] Algorithm 2: Generate instance-level cross-modal co-occurrence relationship matrix. The steps are: Input: image-text pair , visual semantic concepts , text semantic concept .

[0089] Output: Visual and textual co-occurrence matrix of the initial instance layer , .

[0090] 1: Initialize the matrix to zero.

[0091] 2: for each do: 3: for each and do: 4: if and Also appeared in Then: 5: ; ; 6: , ; 7: return , .

[0092] After the above steps, a cross-modal co-occurrence matrix is ​​generated at the instance level using data statistics methods. , and use this matrix to generate a matrix for the image modality that represents the co-occurrence relationship between visual semantic concepts at the instance level , for the text modality, generates a matrix that represents the co-occurrence relationship between text semantic concepts at the instance level The statistics of co-occurrence relationships of cross-modal semantic concepts help to explore the potential semantic associations between different modalities and enhance the model's ability to understand cross-modal information. In addition, by analyzing cross-modal co-occurrence patterns, semantic alignment can be promoted, allowing the model to more accurately capture cross-modal semantic consistency and complementarity, thereby improving the effectiveness of cross-modal retrieval and reasoning.

[0093] S5-3. Perform cross-modal co-occurrence relationship mining at the neighbor layer, and generate the first modal initial visual features and the second modal co-occurrence relationship matrix of the neighbor layer.

[0094] In order to further explore cross-modal co-occurrence relations in a broader semantic space, the cross-modal co-occurrence relations are extended from the instance level to the neighbor level. This extension aims to reveal richer cross-modal co-occurrence relations by analyzing the visual neighbors of the target image and the textual neighbors of the target text.

[0095] At the neighbor layer, the pairing relationship between the visual neighbors of the target image and the textual neighbors of the target text is used as a constraint to discover more extensive cross-modal co-occurrence relationships.

[0096] For an image and a text , using the k-nearest neighbor (KNN) algorithm to identify their first n nearest neighbors, image The nearest neighbors of ,text The nearest neighbors of The cross-modal co-occurrence frequency matrix at the neighboring layer is expressed as ,in express and The co-occurrence frequency of visual and textual neighbors belonging to the same image-text pair in the corpus.

[0097] Algorithm 3 is used to generate the initial visual features and text co-occurrence relationship matrix of the neighbor layer, which can be expressed as , .

[0098] Algorithm 3 generates the cross-modal co-occurrence relationship matrix at the neighboring layer. The steps are: Input: image-text pair , visual semantic concepts , text semantic concept .

[0099] Output: Visual and textual co-occurrence matrix of the initial neighbor layer , .

[0100] 1: Initialize the matrix to zero.

[0101] 2: for each image-text pair do: 3: ; ; 4: for each visual neighbor do: 5: for each text neighbor do: 6: for each and do: 7: if and Also appeared in Then: 8: ; ; 9: , ; 10: return , .

[0102] S5-4. Perform adaptive weak co-occurrence relationship filtering on the first modal initial features and the second modal co-occurrence relationship matrix to obtain a co-occurrence relationship matrix filtered by weak co-occurrence relationship.

[0103] Filtering out weak co-occurrence relationships in the corpus is a key step to improve the performance of the SFE network, which can be summarized in two aspects: on the one hand, weak co-occurrence relationships will cause the model to learn incorrect associations, resulting in overfitting of patterns that lack generalization capabilities; on the other hand, weak co-occurrence relationships exacerbate the imbalance caused by the long-tail distribution, making the model tend to learn irrelevant patterns.

[0104] Each element Matrix is a random variable ,in . Take less than or equal to The cumulative probability of a value is expressed as ,in, is the probability density function.

[0105] Then, use the inverse function Sure The value of express The inverse function of , the corresponding cumulative distribution function is . The calculation formula is: .

[0106] Since the data distribution patterns of each co-occurrence matrix are different, a pre-defined represents the proportion of weak co-occurrence relationships. Then, the threshold According to the probability density function For Matrix Perform adaptive learning. Assume that the co-occurrence matrix Corresponding to the parameter ,by Take the example of adaptive weak co-occurrence relationship filtering, the formula is: (9) Using the same method as above, the matrix , , , , , Weak co-occurrence thresholds in , , , , , Perform adaptive filtering to generate a co-occurrence relationship matrix filtered by weak co-occurrence relationships , , , , , .

[0107] Among them, if data processing is performed on other modal data, multi-layer co-occurrence relationship mining will be performed on the semantic concept sets of different modalities.

[0108] like Figure 1-3 As shown, in step S6, based on the multi-layer co-occurrence relationship, the enhanced first modality global-level features and the enhanced second modality global-level features are further enhanced through a step-by-step feature enhancement network to obtain multi-layer first modality data features and multi-layer second modality data features, and the final data features of the first modality data and the second modality data are respectively fused.

[0109] Semantic concepts and their relations derived from a large-scale image-text corpus are called ESC, and the visual and textual features of the outer layer of the SFE network are enhanced using ESC extracted from multi-layer co-occurrences. The feature enhancement process guided by ESC is denoted as EFE.

[0110] S6-1. Perform external semantic enhancement on the enhanced first modality global-level features and the enhanced second modality global-level features through multi-layer co-occurrence relationships in the external layer.

[0111] In this embodiment, the specific process is: (1) Based on the first modality semantic concept set and the second modality semantic concept set, a first modality and second modality weighted undirected graph is constructed.

[0112] Introducing multi-layer co-occurrence in semantic concept feature representation is helpful for The visual semantic concept of the layer constructs a weighted undirected visual graph with the formula: (10) in, Represents a set of nodes, each node corresponds to a visual semantic concept; represents the set of edges, , where edge Representation Node and The relationship between ; the edge weight matrix is ​​expressed as ,in, express and the strength of the relationship between .

[0113] (2) Use a graph convolutional neural network to update the node embeddings of the first modality and the second modality weighted undirected graphs layer by layer to obtain the node embedding feature matrices of the first modality graph and the second modality graph, respectively.

[0114] First, a graph convolutional network (GCN) with a total number of layers L is used to update the node embedding of the visual graph layer by layer. . No. l The output of the layer is: (11) in, For GCN l The node embedding feature matrix of the layer, Is a visual semantic concept The characteristic representation of represents the edge weight matrix, Indicates that it is located in GCN l-1 The node embedding feature matrix of the layer, Visual representation In GCN l The weight matrix that the layer needs to learn.

[0115] Then, the output of the Lth layer of GCN is used as the final output, that is, the node embedding feature matrix of the visual graph is .

[0116] Similarly, using the above method, we get the embedding feature matrix of the text graph .

[0117] (3) Based on the node embedding feature matrix of the first modal graph and the second modal graph, the first modal feature enhanced by the co-occurrence relationship of the first modality, the first modal feature enhanced by the co-occurrence relationship of the second modality, the second modal feature enhanced by the co-occurrence relationship of the first modality, and the second modal feature enhanced by the co-occurrence relationship of the second modality are calculated.

[0118] In this embodiment, visual features and text features are enhanced under the guidance of ESC.

[0119] In order to achieve deep cross-modal interaction, the co-occurrence relationship within isomorphic modalities (same modality) and between heterogeneous modalities (different modalities) is used to enhance features. The calculation formula of the visual co-occurrence relationship of the layer to enhance the visual features is: (12) in, express Tier The features of visual semantic concepts, express The feature transformation result is defined as , represents the learnable parameter matrix, is a parameter controlling the fractional intensity, Indicates the number of visual semantic concepts contained in the target image.

[0120] The calculation formula of the text co-occurrence relationship enhanced visual feature is: (13) in, express Tier The features of visual semantic concepts, express The feature transformation result is defined as , represents the learnable parameter matrix, Indicates the number of text semantic concepts contained in the target text.

[0121] Similarly, according to formulas (12) and (13), the visual co-occurrence relationship is derived to enhance the text feature Enhance text features with text co-occurrence relations .

[0122] Will and The co-occurrence enhancement classified as isomorphic modality (homomodality) and and Co-occurrence enhancement classified as heterogeneous modalities (different modalities).

[0123] (4) Based on the first modality co-occurrence relationship, the first modality feature is enhanced, the second modality feature is enhanced, the first modality feature is enhanced, and the second modality feature is enhanced based on the first modality co-occurrence relationship. The feature enhancement function is used to perform layered first modality feature enhancement and second modality feature enhancement to obtain enhanced first modality features and second modality features of each layer.

[0124] In this embodiment, based on the ESC-guided feature enhancement (EFE) function, feature enhancement is performed on the visual features and text features of the fragment layer, instance layer and neighbor layer respectively.

[0125] 1) In the fragment layer, by the parameter Control Function The visual features of the fragment layer are enhanced, and the formula is: (14) (15) in, represents the enhanced global visual features, Indicates that visual co-occurrence relationships enhance visual features, Representing text co-occurrence relations to enhance visual features.

[0126] Similarly, by the parameter The controlled fragment-level text features are enhanced using the formula: (16) in, represents the enhanced global text features, Indicates that visual co-occurrence relationships enhance text features, Represents text co-occurrence relationships to enhance text features.

[0127] 2) In the parameters and Under the control of , , and .

[0128] (5) Based on the first modal data features and the first modal features enhanced at each layer, an enhanced data feature of the first modal data is obtained.

[0129] In this embodiment, for i Visual markers , its visual features Through fusion To enhance, the formula is: (17) Among them, ⊙ represents element-by-element multiplication.

[0130] Then, the enhanced visual features of all regions are concatenated row by row to obtain the enhanced visual features of the target image. The formula is: (18) in, represents the number of image regions in the target image, Indicates that the target image is The visual features of the first image region after further feature enhancement. Indicates that the target image is After further feature enhancement, the Visual features of an image region.

[0131] (6) Based on the second modality data features and the enhanced second modality features of each layer, an enhanced data feature of the second modality data is obtained.

[0132] In this embodiment, for j Text tags , through the fusion Enhance its text features , the formula is: (19) Then, the enhanced visual features of the target text are obtained by connecting the enhanced features of all text tags line by line, as follows: (20) in, represents the number of text words in the target text, Indicates that the target text is in The text features of the first text word after feature enhancement, Indicates that the target text is in After feature enhancement The text features of the text words.

[0133] Among them, if data processing is performed on data from other modalities, external semantic enhancement is performed based on multi-layer co-occurrence relationship mining of different modalities.

[0134] After the above steps, the mining of multi-layer co-occurrence relationships in the external layer can effectively make up for the limitations of semantic associations in a single image-text pair, and capture the potential and comprehensive associations between modalities in a higher-level semantic space. In addition, by mining multi-layer co-occurrence relationships, more external knowledge can be introduced, making the trained model more adaptable to different data sets or different tasks.

[0135] S6-2. Perform internal semantic enhancement on the enhanced first modality global-level features and the enhanced second modality global-level features through the cross-modal context of the internal layer.

[0136] like Figure 3 As shown, in this embodiment, visual features and text features are enhanced by mining ESC from the multi-level co-occurrence relationship of the external layer. In order to effectively supplement ISC, valuable semantic information is propagated from the ESC of the external layer to the ISC of the internal layer. Through the IFE module, ISC is combined with cross-modal context mining to further enhance visual features and text features.

[0137] Specifically, (1) based on the augmented data features of the first modality data and the enhanced data features of the second modality data, a similarity matrix between the first modality label and the second modality label is calculated.

[0138] In order to effectively aggregate cross-modal information, we first calculate the similarity matrix between visual and textual tags. S , the formula is: (twenty one) in, and Respectively The feature vector and A collection of feature vectors of dimensional text tags, each element Quantified visual markers and text tags The similarities between.

[0139] (2) Calculate the first modality attention matrix and the second modality attention matrix based on the similarity matrix; Then, S is used to obtain the attention matrix of the two modalities. For the attention of a specific area, the calculation formula of its visual attention matrix is: (twenty two) Among them, the softmax function The rows of are normalized to produce the distribution of attention weights across textual tokens for each visual token.

[0140] (3) Obtain an updated feature representation of the first modality data based on the first modality attention matrix and the enhanced data features of the second modality data.

[0141] These attention weights are used to aggregate the features of the text tags to obtain an updated visual feature representation, as follows: (twenty three) in, The optimized feature representation of the visual token associated with each text token, ie, the updated visual feature representation.

[0142] (4) Obtain an updated feature representation of the second modality data based on the second modality attention matrix and the enhanced visual features of the first modality data.

[0143] Similarly, for the attention of a specific word, the text attention matrix calculation formula is: (twenty four) Among them, the softmax function The rows of are normalized to produce the attention weights for each word in the context of each visual token.

[0144] These weights are used to aggregate visual features to generate updated text feature representations, as follows: (25) in, The optimized feature representation of the textual tag corresponding to each visual tag is represented, that is, the updated textual feature representation.

[0145] After the above steps, the transformed feature set and It encapsulates the fine-grained feature representation produced by the cross-modal attention mechanism, thereby effectively capturing the complex relationship between the visual and textual modalities.

[0146] (5) The enhanced data features of the first modal data and the updated first modal data feature representation are concatenated column by column and passed through a fully connected layer to obtain the first modal data features enhanced by cross-modal information.

[0147] The visual features With its cross-modal context Concatenate column by column and transfer to the fully connected layer to obtain visual features enhanced by cross-modal information. The formula is: FC(concat[ (26) (6) The enhanced data features of the second modal data and the updated second modal data feature representation are concatenated column by column and passed through a fully connected layer to obtain the second modal data features enhanced by cross-modal information.

[0148] Text features With its cross-modal context Concatenate column by column and transfer to the fully connected layer to obtain text features enhanced by cross-modal information. The formula is: FC(concat (27).

[0149] Among them, if data processing is performed on data of other modalities, internal semantic enhancement is performed based on the multi-layer co-occurrence relationship mining of different modalities.

[0150] S6-3. Fuse the final data features of the first modality data and the final data features of the second modality data respectively.

[0151] Enhanced visual features of fragment layer, instance layer and neighbor layer and enhanced text features The final visual features and text features of the first modal data and the target text are fused respectively, and the formula is: (28) (29) in, and They represent the final visual features of the target image and the final text features of the target text respectively.

[0152] After the above steps, the enhancement process from outside to inside fully compensates for the limitations of the internal semantic relationship of image-text pairs.

[0153] S6-4. Using the final first modality data features and the final second modality data features, perform data processing on the first modality data and the second modality data.

[0154] In this embodiment, the final image visual features and text features are used to perform image-text mutual retrieval.

[0155] For matched image-text pairs , respectively collect and and Unmatched Pairs and ,in, and Represents the most difficult negative example. Using the final feature representation and , respectively for the images and text Then, the widely used bidirectional triplet ranking loss is used to align the image and text. The bidirectional triplet ranking loss formula is: (30) in, is the boundary hyperparameter, sim( ) represents the cosine distance function, , the selection formula for the most difficult negative example is: and .

[0156] In this embodiment, before the target image and the target text are mutually retrieved, the constructed network is first trained and tested.

[0157] Flickr30K is a widely used benchmark dataset containing 31,783 images from Flickr. Each image in Flickr30K is accompanied by 5 manually annotated sentences, providing rich context. Flickr30K is divided into three subsets: 29,783 images for training, 1,000 images for validation, and another 1,000 images for testing.

[0158] The MSCOCO dataset consists of 123,287 images, each annotated with 5 sentences. We split it into three subsets: 113,287 images for training, 5000 for validation, and another 5000 for testing. The evaluation involves a challenging setting called MSCOCO(5K), where the proposed method is directly tested on the entire 5000 image set. To ensure the reliability of the results, 1000 test images are validated 5 times and the results are averaged to derive comprehensive performance metrics.

[0159] During training, the proposed model was trained using the PyTorch library on a single NVIDIA GeForce RTX 3090 GPU. The mini-batch size using the Adam optimizer was 128 (i.e., B=128), and the model was trained for 30 epochs. For the first 15 epochs, the learning rate was set to 0.0002 and then decayed by 10%, i.e., 0.00002, in the remaining epochs. To determine the best model, its performance on the validation set was evaluated at the end of each epoch, and the model with the highest R@sum value was selected. The dimension of the visual features was set to 2048 ( = 2048, the number of visual markers is 36 ( =36). The dimension of the text feature is set to 300 ( = 300). Convert visual and text features to a 1024-dimensional feature space ( =1024).

[0160] Meanwhile, in the semantic concept construction, the number of image-text pairs in Flickr30K is 31,783 ( =31,783). The number of image-text pairs in MSCOCO is 123,287 ( =123,287). We set the number of visual concepts and textual concepts to ( =400) and ( = 300. The feature dimension of the visual semantic concept is 1000 ( =1000, and the feature dimension of text semantic concepts is 300 ( =300).

[0161] And set the parameters, the self-attention mechanism is set to 1024 ( = 1024). For the neighbor layer, we select the first 10 neighbors (n = 10). In order to adaptively filter weak co-occurrences, the threshold z of the cumulative distribution function is set to 0.7, and the threshold of the co-occurrence matrix is , , , , , are adaptively set to 0.0604 ( )、0.0165( )、0.0193( )、0.0189( )、0.0342( )、0.0325( ). The number of layers of GCN is set to 2 (L=2). For visual and text feature enhancement, λ=10 is set in equations (11) and (12). In addition, the function EFE ( ) are set to 0.75 for the fragment layer, instance layer, and neighbor layer, respectively ( )、0.65( ) and 0.85 ( ).

[0162] Based on the above conditions, training is carried out, and after the training is completed, it is tested.

[0163] Mutual retrieval tests are conducted on the Flickr30K and MSCOCO test sets. Obviously, the proposed SFE method outperforms the state-of-the-art methods on both datasets: On the Flickr30K dataset, SFE achieves significant improvements over the previous best performing methods, achieving 13.1% and 5.2% performance improvements over ESL with Bi-GRU and NUIF with BERT, respectively.

[0164] On the MSCOCO (1K) dataset (which includes 1000 images from the MSCOCO benchmark), SFE continues to show the best performance. Compared with the previous best method (NUIF) using Bi-GRU, SFE improves the overall performance by 14.0% and compared with the previous best method (NUIF) using BERT, it improves the overall performance by 5.9%. In addition, on the MSCOCO (5K) dataset, SFE improves the previous best method (IMEB) using Bi-GRU and the best method (ESL) using BERT by 6.3% and 2.2%, respectively, further demonstrating its superior performance.

[0165] In summary, SFE shows the best retrieval performance on both benchmark datasets, proving the effectiveness of its step-by-step feature enhancement strategy. The SFE proposed in this example surpasses existing methods by introducing a two-step step-by-step feature enhancement network to propagate semantic information from the outer ESC to the inner ISC. In addition, the proposed SFE network introduces cross-modal ESC for the first time, effectively utilizing its potential to achieve significant improvements in the ITR task.

[0166] Different from DSRAN, VSRN++ and RRTC which focus on mining unimodal ISC enhanced feature representation by exploring spatial location, contextual relations and syntactic connections, SCAN, CAMP, NAAF and ESL focus on mining cross-modal ISC through more refined cross-modal interactions. In contrast, hybrid unimodal and cross-modal ISC methods such as HREM, SGRAF and BOOM make full use of both types of ISC, thus improving ITR performance. In addition, the proposed SFE outperforms the above three methods by leveraging ESC to reveal external semantic patterns such as inter-pair relations within the corpus. By analyzing the global data distribution, ESC provides a comprehensive perspective and enriches the intrinsic semantics of a single pair.

[0167] Compared with CSRC, DRCE, and NUIF, which only rely on extracting unimodal ESC (e.g., neighbor connections and co-occurrence frequencies), SFE shows superior performance. This advantage stems from two key factors: (1) Unimodal ESC ignores cross-modal correlations, which are crucial for bridging the heterogeneous gap. In contrast, this work pioneers the exploration of cross-modal ESC and fully exploits its potential to significantly improve ITR accuracy. (2) By fusing the semantic propagation of cross-modal ESC and ISC, SFE significantly improves retrieval performance.

[0168] In addition, to verify the adaptability of the SFE method to pre-trained models, experiments were conducted using the CLIP model (ViT-L / 14) as the encoder. Compared with fine-tuning CLIP, our model consistently outperforms the baseline on both datasets, demonstrating the effectiveness of SFE in leveraging large-scale pre-trained models.

[0169] In this embodiment, in order to verify the retrieval performance, an ablation experiment is also performed.

[0170] This embodiment adopts a step-by-step feature enhancement network. In image-text mutual retrieval, the step-by-step feature enhancement network is divided into two steps to mine internal and external semantic clues respectively. In step S6-1, the focus is on extracting multi-level co-occurrence relationships in the external layer from a large-scale image-text corpus. In contrast, step S6-2 emphasizes mining cross-modal context in the inner layer. From the analysis of the results, we observe that the combined effect of the two steps is the most effective, and step S6-1 alone is superior to S6-2 in improving ITR performance. This is because step S6-1 extracts valuable semantic information from the large-scale image-text corpus in the outer layer, enhancing the migration of ESC from the outer layer to the inner layer ISC.

[0171] In this embodiment, multi-level co-occurrence mining is adopted, and the co-occurrence relationship is divided into three layers: fragment layer, instance layer and neighbor layer. Specifically, the fragment layer focuses on unimodal co-occurrence, while the instance layer and neighbor layer focus on cross-modal co-occurrence. Compared with other forms of co-occurrence, it is concluded that the co-occurrence of all three layers, as well as their individual contributions, have a positive impact on the ITR task. Specifically, the neighbor layer has the most significant impact, followed by the instance layer, and both have a greater impact than the fragment layer. This is because the neighbor layer captures a wide range of external semantic clues from a large-scale corpus, which is crucial to bridging the heterogeneous gap. In addition, simultaneously integrating the co-occurrence of the fragment layer, instance layer and neighbor layer is the best way to fully utilize their complementary advantages and comprehensively capture the semantic clues behind the relationship between images and texts.

[0172] In this embodiment, intra-modal enhancement is recorded as "same modality", cross-modal enhancement is recorded as "different modalities", and the combination of the two is recorded as "same modality and different modalities". Compared with other methods, cross-modal enhancement is better than intra-modal enhancement. This is because cross-modal enhancement directly establishes a semantic connection between the visual modality and the textual modality, overcoming the limitation of the lack of inter-modal interaction of intra-modal information. In addition, the combination of the two enhancement methods in this embodiment achieved the best results. This is due to the simultaneous use of intra-modal and cross-modal enhancements, making full use of all available semantic clues to improve retrieval performance.

[0173] This example also conducts visualization experiments on the MSCOCO test set using the t-SNE tool. Specifically, 10 semantic concepts are selected from the 80 semantic concepts available in MSCOCO to be compared step by step with the visual and text features obtained by the baseline method CVSE. The results are shown in Figure 4 , Figure 5 , Figure 6 and Figure 7 As shown, among which, Figure 4 , Figure 5 Points of the same color represent visual features belonging to the same semantic concept. Figure 6 , Figure 7 Points of the same color represent text feature representations belonging to the same semantic concept. Figure 4 and Figure 5 In the figure, circle ① is the image area about the “cup”; circle ② is the image area about the “ball”; circle ③ is the image area about the “tennis ball”; circle ④ is the image area about the “dog”; circle ⑤ is the image area about the “chair”; circle ⑥ is the image area about the “table”. Figure 6 and Figure 7 In the picture, circle ① cup: a very small and cute cat is in a cup. There is tea in a cup next to a sandwich and fruit on a plate. A cookie is on a plate with two cups of coffee next to it. Circle ② ball: a lady with a racket is playing ball. A group of children with rackets and tennis balls. A man playing baseball swings his arms wildly. Circle ③ tennis: a boy swings with a tennis racket. A man holds a tennis racket on the court. A man plays tennis on a blue court. Circle ④ dog: a man pulls a dog on a bicycle. A young lady has a puppy in her purse. A golden retriever stands in the snow. Circle ⑤ chair: a group of children sit on chairs to eat. A family sits on chairs to watch TV. A book is placed on the chair. Circle ⑥ table: some people eat and drink next to a red body. A group of people sit at a table to eat. Pizza serving place, several men sit around pizza on a folding table.

[0174] By analyzing the visualization results, the following conclusions can be drawn: Analysis of the visualization results shows that the baseline method cannot effectively distinguish the categories of semantic concepts. Samples from different semantic concepts overlap in feature distribution. Five overlapping semantic concept pairs are analyzed in detail. For example, Figure 4 and Figure 5 (ball, tennis ball), (dog, person), and (cup, table), and Figure 6 and Figure 7 This problem arises because the unimodal ESC extracted by the baseline fails to capture cross-modal co-occurrence relations, limiting its ability to clearly separate semantic concepts.

[0175] Compared with the baseline, the SFE network effectively distinguishes different semantic concepts, and the feature distribution of these concepts has clear boundaries. This improvement is attributed to the introduction of cross-modal ESC, which enhances the correlation between semantically related concepts while further weakening the correlation between semantically unrelated concepts.

[0176] Compared with the feature segmentation in visual modality, both the baseline and the proposed method show better semantic concept differentiation in textual modality. This is because visual semantic concepts are more complex and diverse, while feature learning of textual semantic concepts is relatively easy.

[0177] In summary, the multimodal data processing method and system based on the step-by-step feature enhancement network in this embodiment solves the problems existing in the existing retrieval technology, can make full use of the semantic information from the large-scale image-text corpus, effectively bridge the heterogeneous gap, integrate the co-occurrence of the fragment layer, instance layer and neighbor layer, make full use of their complementary advantages, and comprehensively capture the semantic clues behind the relationship between images and texts. It can fully tap the potential of cross-modal ESC, improve the accuracy of the similarity measurement of different modal data, and thus improve the accuracy of ITR.

[0178] Embodiment 2 The purpose of this embodiment is to provide a multimodal data processing system based on a step-by-step feature enhancement network, comprising: A data acquisition module, used to acquire first modality data and second modality data; A feature extraction module, used to extract first modality data features and second modality data features; An attention enhancement module, used for performing feature enhancement on the first modality data features and the second modality data features through a self-attention mechanism to obtain enhanced first modality global-level features and enhanced second modality global-level features; A semantic construction module, used to construct a first modality semantic concept set and a second modality semantic concept set based on the first modality data corpus and the second modality data corpus, respectively; A multi-layer co-occurrence module is used to mine multi-layer co-occurrence relationships between the first modality semantic concept set and the second modality semantic concept set; The feature enhancement module is used to further enhance the enhanced first modality global-level features and the enhanced second modality global-level features based on the multi-layer co-occurrence relationship through a step-by-step feature enhancement network to obtain multi-layer first modality data features and multi-layer second modality data features, and fuse them respectively to obtain the final data features of the first modality data and the second modality data.

[0179] Based on a provided multimodal data processing system based on a step-by-step feature enhancement network, the method steps in the first embodiment are implemented.

[0180] Embodiment 3 This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method in the first embodiment when executing the program.

[0181] Embodiment 4 This embodiment provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps of the method in the first embodiment are performed.

[0182] The steps involved in the apparatus of the above embodiment correspond to the method embodiment 1, and the specific implementation method can refer to the relevant description part of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.

[0183] Those skilled in the art should understand that the modules or steps of the present invention described above can be implemented by a general-purpose computer device, or alternatively, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0184] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.

Claims

1. A multimodal data processing method based on a step-by-step feature enhancement network, characterized in that: include: Acquire first modality data and second modality data; Extracting first modality data features and second modality data features; Performing feature enhancement on the first modality data features and the second modality data features through a self-attention mechanism to obtain enhanced first modality global-level features and enhanced second modality global-level features; Based on the first modality data corpus and the second modality data corpus, constructing a first modality semantic concept set and a second modality semantic concept set respectively; Conduct multi-layer co-occurrence relationship mining on the first modality semantic concept set and the second modality semantic concept set; Based on the multi-layer co-occurrence relationship, the enhanced first modality global-level features and the enhanced second modality global-level features are further enhanced through a step-by-step feature enhancement network to obtain multi-layer first modality data features and multi-layer second modality data features, and the final data features of the first modality data and the second modality data are respectively fused.

2. A multimodal data processing method based on a step-by-step feature enhancement network as claimed in claim 1, characterized in that: Extract the first modal data features. The specific process is as follows: Utilize an object detection model with bottom-up attention to extract the target region; A deep convolutional network is used to extract the first modality data features of the target area.

3. The multimodal data processing method based on a step-by-step feature enhancement network according to claim 1, characterized in that: Extracting the second modal data features, specifically, encoding the second modal data using an encoder of a deep neural network to obtain the second modal data features.

4. The multimodal data processing method based on a step-by-step feature enhancement network according to claim 1, characterized in that: For the first modal semantic concept set and the second modal semantic concept set, multi-layer co-occurrence relationship mining is performed, where the multi-layer co-occurrence relationship includes the fragment layer, instance layer and neighbor layer. The specific process is as follows: Perform unimodal co-occurrence relationship mining at the fragment layer, and generate the first modal initial features and the second modal co-occurrence relationship matrix at the fragment layer; Perform cross-modal co-occurrence relationship mining at the instance level, and generate the first modal initial features and the second modal co-occurrence relationship matrix at the instance level; Perform cross-modal co-occurrence relationship mining at the neighbor layer, and generate the first modal initial visual features and the second modal co-occurrence relationship matrix of the neighbor layer; Adaptive weak co-occurrence relationship filtering is performed on the first modal initial features and the second modal co-occurrence relationship matrix to obtain a co-occurrence relationship matrix for weak co-occurrence relationship filtering.

5. The multimodal data processing method based on a step-by-step feature enhancement network according to claim 1, characterized in that: Based on the multi-layer co-occurrence relationship, the enhanced first modality global-level features and the enhanced second modality global-level features are further enhanced through a step-by-step feature enhancement network to obtain multi-layer first modality data features and multi-layer second modality data features. The specific process is as follows: External semantic enhancement is performed on the enhanced first modality global-level features and the enhanced second modality global-level features through multi-layer co-occurrence relationships in the external layer; The enhanced first modality global-level features and the enhanced second modality global-level features are internally semantically enhanced through the cross-modal context of the internal layer.

6. A multimodal data processing method based on a step-by-step feature enhancement network as claimed in claim 5, characterized in that: The enhanced first modality global-level features and the enhanced second modality global-level features are externally semantically enhanced through the multi-layer co-occurrence relationship of the external layer, specifically: Based on the first modality semantic concept set and the second modality semantic concept set, construct a first modality and a second modality weighted undirected graph; Using a graph convolutional neural network, the node embeddings of the first modality and the second modality weighted undirected graphs are updated layer by layer to obtain node embedding feature matrices of the first modality graph and the second modality graph respectively; Based on the node embedding feature matrix of the first modal graph and the second modal graph, the first modal feature enhanced by the first modal co-occurrence relationship, the first modal feature enhanced by the second modal co-occurrence relationship, the second modal feature enhanced by the first modal co-occurrence relationship, and the second modal feature enhanced by the second modal co-occurrence relationship are calculated; Based on the first modality co-occurrence relationship, the first modality feature is enhanced, the second modality feature is enhanced, the first modality feature is enhanced, and the second modality feature is enhanced by the first modality co-occurrence relationship, and the second modality feature is enhanced by the second modality co-occurrence relationship, and the feature enhancement function is used to perform layered first modality feature enhancement and second modality feature enhancement to obtain enhanced first modality features and second modality features of each layer; Obtaining enhanced data features of the first modal data based on the first modal data features and the first modal features enhanced at each layer; Based on the second modality data features and the enhanced second modality features of each layer, enhanced data features of the second modality data are obtained.

7. The multimodal data processing method based on a step-by-step feature enhancement network according to claim 5, characterized in that: The enhanced first modality global-level features and the enhanced second modality global-level features are internally semantically enhanced through the cross-modal context of the internal layer. The specific process is as follows: Calculating a similarity matrix between the first modality tag and the second modality tag based on the augmented data features of the first modality data and the enhanced data features of the second modality data; According to the similarity matrix, calculating the first modality attention matrix and the second modality attention matrix; Obtaining an updated feature representation of the first modality data according to the first modality attention matrix and the enhanced data features of the second modality data; Obtaining an updated second modality data feature representation according to the second modality attention matrix and the enhanced data features of the first modality data; The enhanced data features of the first modal data and the updated first modal data feature representation are concatenated column by column, and passed through a fully connected layer to obtain the first modal data features enhanced by using cross-modal information; The enhanced data features of the second modality data and the updated second modality data feature representation are concatenated column by column and passed through a fully connected layer to obtain the second modality data features enhanced by cross-modality information.

8. A multimodal data processing system based on a step-by-step feature enhancement network, characterized in that: include: A data acquisition module, used to acquire first modality data and second modality data; A feature extraction module, used to extract first modality data features and second modality data features; An attention enhancement module, used for performing feature enhancement on the first modality data features and the second modality data features through a self-attention mechanism to obtain enhanced first modality global-level features and enhanced second modality global-level features; A semantic construction module, used to construct a first modality semantic concept set and a second modality semantic concept set based on the first modality data corpus and the second modality data corpus, respectively; A multi-layer co-occurrence module is used to mine multi-layer co-occurrence relationships between the first modality semantic concept set and the second modality semantic concept set; The feature enhancement module is used to further enhance the enhanced first modality global-level features and the enhanced second modality global-level features based on the multi-layer co-occurrence relationship through a step-by-step feature enhancement network to obtain multi-layer first modality data features and multi-layer second modality data features, and fuse them respectively to obtain the final data features of the first modality data and the second modality data.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 7 are performed.

Citation Information

Patent Citations

  • Intelligent search method and system based on multi-source heterogeneous data

    CN116049454A

  • Image text retrieval method and system based on context-guided multi-modal association

    CN116737979A

  • Cross-modal image-text retrieval method and system based on local context

    CN116775927A

  • Image text semantic matching method and system applied to image-text retrieval

    CN118132788A

  • Multi-scale visual semantic enhanced multi-modal named entity recognition method and system

    CN118364817A