A Multimodal Semantic Representation Method and System Based on Social-Like Priors

By introducing dynamic structured interaction modules and discriminant pillar verification modules into the self-attention mechanism, the representation of visual and language tasks is optimized, and the rank collapse and representation degradation caused by the traditional self-attention mechanism is solved, and a higher quality multimodal semantic representation is achieved.

CN118194109BActive Publication Date: 2025-05-27SHANDONG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410235890.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-01
Publication Date
2025-05-27
Estimated Expiration
2044-03-01

AI Technical Summary

Technical Problem

Traditional self-attention mechanisms are prone to rank collapse and representational degradation in natural language processing, and it is difficult to generate more expressive visual and linguistic representations.

Method used

A multimodal semantic representation method based on social priors is introduced, and the self-attention mechanism is optimized through the dynamic structured interaction module and the discriminant pillar verification module to avoid redundant information aggregation and global feature homogeneity, and to enhance the discriminantity of visual semantics.

Benefits of technology

It effectively avoids rank collapse and representational degradation in the self-attention mechanism, improves the representation quality of visual and language tasks, and enhances the fusion ability of multimodal features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118194109B_ABST
    Figure CN118194109B_ABST
Patent Text Reader

Abstract

The present invention relates to a multi-modal semantic representation method and system based on social-like priors, including: using a pre-trained visual model to understand a given image and generate corresponding image grid features; using a natural semantic understanding model to effectively represent a given question or query statement and generate statement features; inputting the image grid features and the statement features into a social-like transformer, and through multi-layer encoding, realizing effective fusion of multi-modal features, and finally obtaining high-quality multi-modal semantic representation. The present invention introduces a carefully designed social-like attention mechanism into the traditional transformer architecture to achieve visual structured modeling and discriminative semantic learning, further enhancing the learning of visual context.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a multimodal semantic representation method and system based on quasi-social prior, belonging to the technical field of language processing. Background Art

[0002] Transformer-based methods have been successful in the field of natural language, paving the way for the prosperity of visual reasoning tasks. Many carefully designed Transformer variants have achieved promising performance in various benchmarks. Due to the powerful global modeling ability of the self-attention mechanism in Transformer, these methods not only help learn the internal modal context, but also perform well in cross-modal alignment and complementation. However, as discussed in some existing studies, the traditional self-attention mechanism can easily lead to rank collapse and representation degradation in the absence of feedforward networks and residual connections. Therefore, how to further optimize the effective learning of self-attention and generate more expressive representations for vision and language tasks remains a pressing issue.

[0003] The feature aggregation of different image regions in the self-attention mechanism shares a similar concept with the information transfer in social networks. For the visual self-attention modeling in Transformer, each visual unit, which can be a grid feature (Jiang et al., 2020) or a salient object (Anderson et al., 2018), will aggregate features from other visual units according to the similarity score. This process is considered to be similar to social behavior. In other words, each visual area can be regarded as a social member, and each member tends to make friends with like-minded people. Summary of the invention

[0004] In view of the shortcomings of the prior art, the present invention provides a multimodal semantic representation method based on quasi-social priors;

[0005] The present invention optimizes the learning of the self-attention mechanism from two social perspectives: (1) Inspired by the concept of structural holes mentioned in SNT (Tabassum et al., 2018), the present invention believes that not all pairs of regions tend to establish connections in the self-attention mechanism. Most regions tend to interact with a limited number of regions except for structural holes, which can serve as links to communicate with several isolated regional groups. This mechanism makes the information interaction between different regions more structured and effectively avoids redundant information aggregation and global feature homogenization, demonstrating the dynamic structured interaction module (DSI) of the present invention. (2) After DSI, the mined structural holes act as information hubs responsible for receiving more information sources because they have a greater degree and a greater possibility of connecting with representative regions in different groups. However, how to find these potential structural holes is a question worth exploring. Based on this, the present invention considers the concepts of degree centrality and transitivity in SNT, and further explores the spontaneous formation conditions of pillars: a region is more likely to become a pillar when: 1) it receives more attention from other regions, or, 2) it tends to establish closer connections with other pillars, which naturally leads to discriminative pillar verification (DPV). These two components are sequentially integrated into the basic self-attention mechanism and work systematically as a whole (i.e., SSA) for structured rank optimization and refinement of discriminative representations.

[0006] The present invention also provides a multimodal semantic representation system based on quasi-social priors.

[0007] Terminology explanation:

[0008] Visual Genome Dataset: The Language and Vision Dataset proposed by Stanford in 2016 aims to better understand the world. It annotates a large number of objects and relationships in pictures, as well as question-answer pairs about pictures.

[0009] ResNext model: A ResNet-like network that combines grouped convolution with residual connections to improve visual representation capabilities while reducing parameters.

[0010] Glove: It is a pre-trained vocabulary, in which each vector corresponds to a predefined word.

[0011] LSTM: It is a recurrent neural network used to model sequence relationships and can be used to extract local and global context representations of sequences.

[0012] Traditional cross-attention: In the cross-attention mechanism, the fine-grained information of the two modalities (i.e., each region in the image and each word in the text) is first calculated with a similarity matrix, and the fine-grained information of each modality is enhanced by a linear combination of the information of the other modality, thereby achieving multi-modal representation enhancement and supplementation.

[0013] Multilayer Perceptron Network: The multilayer perceptron network consists of two linear fully connected networks. The first fully connected network achieves dimensionality increase, and the second fully connected network achieves dimensionality reduction.

[0014] The technical solution of the present invention is as follows:

[0015] A multimodal semantic representation method based on quasi-social prior, comprising:

[0016] Use the pre-trained visual model to understand the given image and generate the corresponding image grid features;

[0017] Use the natural semantic understanding model to effectively represent the given question or query sentence and generate sentence features;

[0018] The image grid features and sentence features are input into a social-like transformer together, and multi-modal features are effectively fused through multi-layer encoding to finally obtain high-quality multimodal semantic representation.

[0019] Preferably, according to the present invention, a pre-trained visual model is used to understand a given image and generate corresponding image grid features; comprising: using a pre-trained ResNext model on the Visual Genome dataset to extract visual features from the b-th image, and obtaining a 14x14 feature map as the input of the visual branch, that is, a visual area representation matrix, represented as U b =[u b1 ,u b2 ,…,u bR ], where R is the total number of visual sub-areas.

[0020] Preferably, according to the present invention, a natural semantic understanding model is used to effectively characterize a given question or query statement and generate a statement feature; that is, Glove and LSTM are used to encode a given question or query statement: including:

[0021] First, locate the specific vector representation in Glove according to the word subscript;

[0022] Then, the vector is input into LSTM to encode local and global context information, and further implements the semantic supplementation and enhancement between words based on the multi-head attention mechanism; thus achieving a more comprehensive semantic understanding and generating the corresponding sentence representation set T b =[tb1 ,t b2 ,…,t bW ].

[0023] Preferably, according to the present invention, the social-like transformer is composed of multiple layers of social-like blocks in series, specifically including a social-like self-attention mechanism, a traditional cross attention, a multi-layer perceptron network, a classification head or a regression head;

[0024] A social-like self-attention mechanism enables structured interaction and discriminative visual representation learning;

[0025] Traditional cross-attention integrates textual information with condensed visual information in a fine-grained manner;

[0026] Use a multi-layer perceptron network to complete the dimensionality denoising of features;

[0027] The final multimodal representation is input into the classification head or regression head for downstream tasks of visual question answering and visual positioning.

[0028] Preferably, according to the present invention, a social self-attention mechanism realizes structured interaction and discriminative visual representation learning; including:

[0029] The quasi-social self-attention mechanism includes a dynamic structured interaction module (DSI) and a discriminative pillar verification module (DPV). The dynamic structured interaction module uses the structural hole theory to realize structured interaction in the visual area, and the discriminative pillar verification module reallocates weights for potential structural holes.

[0030] First, the image is visually encoded using the pre-trained ResNext model to obtain the visual area representation matrix U b , where the visual area representation matrix U b Each row represents the representation of a visual area; the representation of the visual area is input into three linear fully connected layers f q ,f k and f v , get the corresponding query Q, key K and value V, and calculate the pairwise interaction score matrix A ir , as shown below:

[0031]

[0032] Among them, u i and u r are the i-th and r-th visual areas, respectively, and q i That is, f q (u i ), is the corresponding query, k r That is, f k (u r), is the corresponding key, and d represents the dimension of the feature;

[0033] Then, two modulation networks are used and (i.e. two linear fully connected networks) using feature u at the same time i and Connection Preference A i,: As input, an effective mask structure that helps form the structural hole is generated in the form of an affine transformation, as shown below:

[0034]

[0035] Where F = [A i,: ;u i ] represents the concatenation of features and connection preferences, R(A i,: ) represents A i,: The mean value, M i,: represents the learned mask;

[0036] Using the obtained mask M i,: Generate hard mask H i,: , as shown below:

[0037] H i,: =d(ReLU(M i,: ))

[0038] Among them, d is the discretization operation;

[0039] Again, the self-attention mechanism is corrected using the learned structured mask as follows:

[0040]

[0041] Among them, H ij is the discretization mask, A ij is the original interaction matrix, is the updated interaction matrix;

[0042] use Update the original visual features;

[0043] The key attribute information (degree information) of the node is used to locate the potential structural holes, as shown below:

[0044] w i =H i,: ⊙(s+A i,: )

[0045] Among them, H i,: ⊙s describes the importance of the node, H i,: ⊙A i,: It describes the transfer relationship of the relationship, so wi The conditions for each node to become a pillar node are described: whether it has a closer connection with more important nodes; at the same time, a more flexible and learnable form is adopted to increase the flexibility of pillar node determination, as shown below:

[0046]

[0047] Among them, s i is the global importance of the i-th node, p i To calculate the final node importance considering node degree information and transitive properties;

[0048] Finally, the original visual features are further corrected using the pillar verification weights as follows:

[0049]

[0050] Among them, the importance of each node is combined to form an importance vector, and Represent the feature representation after pillar verification and the feature representation before respectively. Where p is a vector, each element of which is p i ;

[0051] According to the preferred embodiment of the present invention, the traditional cross attention integrates the text information with the condensed visual information in a fine-grained manner; and uses a multi-layer perceptron network to perform feature denoising; including:

[0052] The feature representation after pillar verification is aligned and interacted with the original text features through the traditional cross-attention mechanism to achieve fine-grained alignment and interaction, and the representation is then input into the multi-layer perceptron network for further denoising and representation optimization.

[0053] Preferably, according to the present invention, the finally obtained multimodal representation is input into a classification head or a regression head for visual question answering and visual positioning downstream tasks; comprising:

[0054] The visual area representation matrix U b With T b After multiple iterations of the social-like transformer, the corrected multimodal representation is obtained

[0055] For the visual problem task, and the sentence representation set T b Input into the classification head for multi-label classification, the training process minimizes the distribution distance between the predicted answer distribution and the one-hot code; for text-based visual localization tasks, The four positions of the predicted detection box are directly obtained by inputting into the regression head. The training process minimizes the distance between the predicted positioning information (i.e., the four position coordinates) and the labeled positioning information.

[0056] Preferably according to the present invention, the optimization function is the adam optimizer optimizer in Pytorch.

[0057] A computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of a multimodal semantic representation method based on quasi-social prior when executing the computer program.

[0058] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a multimodal semantic representation method based on quasi-social priors.

[0059] A multimodal semantic representation system based on quasi-social priors, comprising:

[0060] An image representation module is configured to: use a pre-trained visual model to understand a given image and generate corresponding image grid features;

[0061] The sentence representation module is configured to: effectively represent a given question or query sentence using a natural semantic understanding model to generate sentence features;

[0062] The multimodal semantic representation module is configured to input the image grid features and sentence features into the social-like transformer together, realize the effective fusion of multimodal features through multi-layer encoding, and finally obtain high-quality multimodal semantic representation.

[0063] The beneficial effects of the present invention are:

[0064] 1. Inspired by social networks, the present invention compares the information interaction of the self-attention structure in visual context learning with the information exchange between people in social networks, providing a feasible solution to the feature degradation and rank collapse caused by self-attention.

[0065] 2. The present invention introduces a cleverly designed social attention mechanism into the traditional transformer architecture to achieve visual structured modeling and discriminative semantic learning, further enhancing the learning of visual context.

[0066] 3. The quasi-social self-attention mechanism proposed in the present invention mainly includes two sub-modules, namely dynamic structured interaction and discriminative pillar verification. Dynamic structured modeling avoids redundant interactions between any pairs of nodes by introducing the structural hole theory, making the transmission of information more structured, that is, information within the same semantic group tends to interact only within the semantic group to avoid the introduction of redundant information interference; and structural holes, as the hub connecting multiple semantic groups, play the role of multi-source information link, thereby effectively avoiding the global homogenization of information. After the structural holes are formed, the present invention uses the degree information of the nodes and the connection information of the relationships to dynamically find these structural holes, increase their importance in the global representation, and thus improve the discriminability of visual semantics. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 It is a schematic diagram of the network model architecture proposed by the present invention;

[0068] Figure 2 It is a schematic diagram comparing the impact of the dynamic structured interaction module of the present invention on the interaction modes of different regions;

[0069] Figure 3 It is a schematic diagram comparing the impact of the discriminative pillar verification module of the present invention on multimodal representation distribution learning. DETAILED DESCRIPTION

[0070] The present invention will be further defined below in conjunction with the accompanying drawings and embodiments, but is not limited thereto.

[0071] Example 1

[0072] A multimodal semantic representation method based on quasi-social prior, comprising:

[0073] Use the pre-trained visual model to understand the given image and generate the corresponding image grid features;

[0074] Use the natural semantic understanding model to effectively represent the given question or query sentence and generate sentence features;

[0075] The image grid features and sentence features are input into a social-like transformer together, and multi-modal features are effectively fused through multi-layer encoding to finally obtain high-quality multimodal semantic representation.

[0076] Example 2

[0077] The difference between the multimodal semantic representation method based on quasi-social prior described in Example 1 is that:

[0078] like Figure 1As shown in the figure, by introducing the social-like prior into the traditional transformer, a multimodal collaborative reasoning network model is constructed to achieve visual problem tasks and text-based visual localization tasks. The multimodal collaborative reasoning network model first accurately understands and effectively represents multimodal information (images and texts), then models the visual context information to achieve visual reasoning, aligns the condensed visual information with the text information in a fine-grained manner, enhances each other, and uses the final discriminative representation for classification and regression.

[0079] Use the pre-trained visual model to understand the given image and generate the corresponding image grid features; including: using the pre-trained ResNext model on the Visual Genome dataset to extract visual features from the bth image, and obtain a 14x14 feature map as the input of the visual branch, that is, the visual area representation matrix, represented as U b =[u b1 ,u b2 ,…,u bR ], where R is the total number of visual sub-areas.

[0080] Using the natural semantic understanding model to effectively represent the given question or query statement and generate sentence features; it means: using Glove and LSTM to encode the given question or query statement: including:

[0081] First, locate the specific vector representation in Glove according to the word subscript;

[0082] Then, the vector is input into LSTM to encode local and global context information, and the multi-head attention mechanism is used to further implement the word complement and enhancement between words; specifically, the vector is input into LSTM, and LSTM removes some redundant features in the sentence through the forget gate and the reset gate, and combines the new features with the useful features in the past to achieve more refined text encoding, and further implements the word complement and enhancement between words based on the multi-head attention mechanism, so as to achieve a more comprehensive semantic understanding and generate the corresponding sentence representation set T. b =[t b1 ,t b2 ,…,t bW ].

[0083] The social transformer is composed of multiple layers of social blocks, including social self-attention mechanism, traditional cross attention, multi-layer perceptron network, classification head or regression head.

[0084] A social-like self-attention mechanism enables structured interaction and discriminative visual representation learning;

[0085] Traditional cross-attention integrates textual information with condensed visual information in a fine-grained manner;

[0086] Use a multi-layer perceptron network to complete the dimensionality denoising of features;

[0087] The final multimodal representation is input into the classification head or regression head for downstream tasks of visual question answering and visual positioning.

[0088] A social self-attention mechanism enables structured interaction and discriminative visual representation learning; including:

[0089] The biggest difference between the social transformer and the traditional transformer is the introduction of the social self-attention mechanism: the SSA module. Among them, SSA replaces the traditional self-attention mechanism Self-Attention. The social self-attention mechanism includes the dynamic structured interaction module (DSI) and the discriminative pillar verification module (DPV). The dynamic structured interaction module uses the structural hole theory to realize the structured interaction of the visual area and avoid the transmission of redundant information in the interaction; the discriminative pillar verification module reallocates the weights for potential structural holes, thereby enhancing the discriminative representation.

[0090] Improve the traditional visual self-attention mechanism based on quasi-social prior.

[0091] First, the image is visually encoded using the pre-trained ResNext model to obtain the visual area representation matrix U b , where the visual area representation matrix U b Each row of represents the representation of a visual area; as in the traditional self-attention mechanism modeling network, the representation of the visual area is input into three linear fully connected layers f q ,f k and f v , get the corresponding query Q, key K and value V, and calculate the pairwise interaction score matrix A ir , as shown below:

[0092]

[0093] Among them, u i and u r are the i-th and r-th visual areas, respectively, and q i That is, f q (u i ), is the corresponding query, k r That is, f k (u r ), is the corresponding key, and d represents the dimension of the feature;

[0094] Then, in the dynamic structured interaction module (DSI), according to the structural hole theory, whether a node becomes a structural hole depends on the relative position context and its own characteristics; therefore, two modulation networks are used and (i.e. two linear fully connected networks) using feature u at the same time i and Connection Preference A i,: As input, an effective mask structure that helps form the structural hole is generated in the form of an affine transformation, as shown below:

[0095]

[0096] Where F = [A i,: ;u i ] represents the concatenation of features and connection preferences, R(A i,: ) represents A i,: The mean value, M i,: represents the learned mask;

[0097] Using the obtained mask M i,: Generate hard mask H i,: , thereby enhancing the removal of redundant relations, as shown below:

[0098] H i,: =d(ReLU(M i,: ))

[0099] Among them, d is the discretization operation;

[0100] Figure 2 This is a comparative diagram of the impact of dynamic structured interaction modules on interaction patterns in different regions; structural holes, as hubs connecting multiple semantic groups, are associated with more nodes, while non-structural hole nodes only interact with neighboring nodes. The existence of structural holes is conducive to rank learning of the interaction matrix (i.e., more diversified interactions) and avoids global homogeneity of information.

[0101] Again, the learned structured mask is used to correct the self-attention mechanism as follows:

[0102]

[0103] Among them, H ij is the discretization mask, A ij is the original interaction matrix, is the updated interaction matrix;

[0104] use The original visual features are updated, and the structural holes make the information interaction more structured, which is reflected in the fact that the learned semantic representation within the group has better continuity and roughly approximates the semantic distribution of the original image.

[0105] In the discriminative pillar verification module, the key attribute information (degree information) of the node is used to locate the potential structural holes, as shown below:

[0106] w i =H i,: ⊙(s+A i,: )

[0107] Among them, H i,: ⊙s describes the importance of the node, H i,: ⊙A i,: Draw the transfer relationship of the relationship, so w i The conditions for each node to become a pillar node are described: whether it has a closer connection with more important nodes; at the same time, a more flexible and learnable form is adopted to increase the flexibility of pillar node determination, as shown below:

[0108]

[0109] Among them, s i is the global importance of the i-th node, p i The final node importance is considered taking into account the node degree information and the transitive property; the transitive property can be understood as: when an important interaction occurs with an important node, this node is considered to be important as well.

[0110] Finally, the original visual features are further corrected using the pillar verification weights as follows:

[0111]

[0112] Among them, the importance of each node is combined to form an importance vector, and Represent the feature representation after pillar verification and the feature representation before respectively. Where p is a vector, each element of which is p i ; Figure 3 The figure is a comparative diagram of the impact of the discriminative pillar verification module on the multimodal representation distribution learning. (a) shows the multimodal representation distribution learning without the discriminative pillar verification module, and (b) shows the multimodal representation distribution learning with the discriminative pillar verification module. The visual representation after pillar verification has better discriminability.

[0113] Traditional cross attention integrates text information with condensed visual information in a fine-grained manner; multi-layer perceptron network is used to denoise features; including:

[0114] The feature representation after pillar verification is aligned and interacted with the original text features through the traditional cross-attention mechanism to achieve fine-grained alignment and interaction, thereby enhancing the multimodal semantic representation of vision and text respectively. The representation is then input into the multi-layer perceptron network for further denoising and representation optimization.

[0115] The final multimodal representation is input into the classification head or regression head for downstream tasks of visual question answering and visual positioning, including:

[0116] The visual area representation matrix U b With T b After multiple iterations of the social-like transformer, the corrected multimodal representation is obtained

[0117] For the visual problem task, and the sentence representation set T b Input into the classification head for multi-label classification, the training process minimizes the distribution distance between the predicted answer distribution and the one-hot code; for text-based visual localization tasks, The four positions of the predicted detection box are directly obtained by inputting into the regression head. The training process minimizes the distance between the predicted positioning information (i.e., the four position coordinates) and the labeled positioning information.

[0118] The optimization function is used to solve the parameters of the entire network model in an end-to-end manner.

[0119] The optimization function is the adam optimizer in Pytorch.

[0120] Three visual localization datasets, RefCOCO, RefCOCO+, and RefCOCOg, are used to verify the gain of the social prior proposed in this embodiment on the basic transformer. Table 1 is the result data table;

[0121] Table 1

[0122]

[0123] Table 2 is an introduction table of the sources of the comparison models in Table 1;

[0124] Table 2

[0125]

[0126] As can be seen from Table 1, the method of this embodiment achieves further gains on the two latest transformer-based multimodal modeling methods TransVG and VLVTG, achieving the best performance.

[0127] Example 3

[0128] A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the multimodal semantic representation method based on quasi-social prior described in Example 1 or 2 are implemented.

[0129] Example 4

[0130] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the multimodal semantic representation method based on quasi-social priors described in Example 1 or 2.

[0131] Example 5

[0132] A multimodal semantic representation system based on quasi-social priors, comprising:

[0133] An image representation module is configured to: use a pre-trained visual model to understand a given image and generate corresponding image grid features;

[0134] The sentence representation module is configured to: effectively represent a given question or query sentence using a natural semantic understanding model to generate sentence features;

[0135] The multimodal semantic representation module is configured to input the image grid features and sentence features into the social-like transformer together, realize the effective fusion of multimodal features through multi-layer encoding, and finally obtain high-quality multimodal semantic representation.

Claims

1. A multimodal semantic representation method based on quasi-social prior, characterized in that: include: Use the pre-trained visual model to understand the given image and generate the corresponding image grid features; Use the natural semantic understanding model to effectively represent the given question or query sentence and generate sentence features; The image grid features and sentence features are input into the social-like transformer together, and multi-modal features are effectively fused through multi-layer encoding to finally obtain high-quality multi-modal semantic representation. The social transformer is composed of multiple layers of social blocks, including social self-attention mechanism, traditional cross attention, multi-layer perceptron network, classification head or regression head. A social-like self-attention mechanism enables structured interaction and discriminative visual representation learning; Traditional cross-attention integrates textual information with condensed visual information in a fine-grained manner; Use a multi-layer perceptron network to complete the dimensionality denoising of features; The final multimodal representation is input into the classification head or regression head for downstream tasks of visual question answering and visual localization. A social-like self-attention mechanism enables structured interaction and discriminative visual representation learning; include: The quasi-social self-attention mechanism includes a dynamic structured interaction module and a discriminative pillar verification module. The dynamic structured interaction module uses the structural hole theory to realize the structured interaction of the visual area, and the discriminative pillar verification module redistributes weights for potential structural holes. First, the image is visually encoded using the pre-trained ResNext model to obtain the visual area representation matrix U b , where the visual area representation matrix U b Each row represents the representation of a visual area; the representation of the visual area is input into three linear fully connected layers f q ,f k and f v , get the corresponding query Q, key K and value V, and calculate the pairwise interaction score matrix A ir , as shown below: Among them, u i and u r are the i-th and r-th visual areas, respectively, and q i That is, f q (u i ), is the corresponding query, k r That is, f k (u r ), is the corresponding key, and d represents the dimension of the feature; Then, two modulation networks are used and Using feature u i and Connection Preference A i,: As input, an effective mask structure that helps form the structural hole is generated in the form of an affine transformation, as shown below: Where F = [A i,: ;u i ] represents the concatenation of features and connection preferences, R(A i,: ) represents A i,: The mean value, M i,: represents the learned mask; Using the obtained mask M i,: Generate hard mask H i,: , as shown below: H i,: =d(ReLU(M i,: )) Among them, d is the discretization operation; Again, the learned structured mask is used to correct the self-attention mechanism as follows: Among them, H ij is the discretization mask, A ij is the original interaction matrix, is the updated interaction matrix; use Update the original visual features; The key attribute information of the node is used to locate the potential structural holes, as shown below: w i =H i,: ⊙(s+A i,: ) Among them, H i,: ⊙s describes the importance of the node, H i,: ⊙A i,: It describes the transfer relationship of the relationship, so w i The conditions for each node to become a pillar node are described: whether it has a closer connection with more important nodes; at the same time, a more flexible and learnable form is adopted to increase the flexibility of pillar node determination, as shown below: Among them, s i is the global importance of the i-th node, p i To calculate the final node importance considering node degree information and transitive properties; Finally, the original visual features are further corrected using the pillar verification weights as follows: Among them, the importance of each node is combined to form an importance vector, and They represent the feature representation before and after pillar verification respectively; p is a vector, each element of which is p i .

2. According to claim 1, a multimodal semantic representation method based on quasi-social prior is characterized in that: Use the pre-trained visual model to understand the given image and generate the corresponding image grid features; including: using the pre-trained ResNext model on the VisualGenome dataset to extract visual features from the bth image, and the obtained feature map is used as the input of the visual branch, that is, the visual area representation matrix, represented as U b =[u b1 ,u b2 ,…,u bR ], where R is the total number of visual sub-areas.

3. The multimodal semantic representation method based on quasi-social prior according to claim 1, characterized in that: Use the natural semantic understanding model to effectively represent the given question or query sentence and generate sentence features; It means: using Glove and LSTM to encode a given question or query statement: including: First, locate the specific vector representation in Glove according to the word subscript; Then, the vector is input into LSTM to encode local and global context information, and further implements the semantic supplementation and enhancement between words based on the multi-head attention mechanism; thus achieving a more comprehensive semantic understanding and generating the corresponding sentence representation set T b =[t b1 ,t b2 ,…,t bW ].

4. The multimodal semantic representation method based on quasi-social prior according to claim 1, characterized in that: Traditional cross attention integrates text information with condensed visual information in a fine-grained manner; multi-layer perceptron network is used to denoise features; including: The feature representation after pillar verification is aligned and interacted with the original text features through the traditional cross-attention mechanism to achieve fine-grained alignment and interaction, and the representation is then input into the multi-layer perceptron network for further denoising and representation optimization.

5. The multimodal semantic representation method based on quasi-social prior according to claim 1, characterized in that: The final multimodal representation is input into the classification head or regression head for downstream tasks of visual question answering and visual localization. include: The visual area representation matrix U b and the sentence representation set T b After multiple iterations of the social-like transformer, the corrected multimodal representation is obtained For the visual problem task, and the sentence representation set T b Input into the classification head for multi-label classification, the training process minimizes the distribution distance between the predicted answer distribution and the one-hot code; for text-based visual localization tasks, The four positions of the predicted detection box are directly obtained by inputting into the regression head. The training process minimizes the distance between the predicted positioning information and the labeled positioning information.

6. The multimodal semantic representation method based on quasi-social prior according to claim 5, characterized in that: The optimization function is the adam optimizer in Pytorch.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the multimodal semantic representation method based on quasi-social priors described in any one of claims 1-6 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the multimodal semantic representation method based on quasi-social priors described in any one of claims 1 to 6 are implemented.

9. A multimodal semantic representation system based on quasi-social priors, characterized in that: include: An image representation module is configured to: use a pre-trained visual model to understand a given image and generate corresponding image grid features; The sentence representation module is configured to: effectively represent a given question or query sentence using a natural semantic understanding model to generate sentence features; The multimodal semantic representation module is configured to: input the image grid features and sentence features into the social-like transformer together, realize the effective fusion of multimodal features through multi-layer encoding, and finally obtain high-quality multimodal semantic representation; The social transformer is composed of multiple layers of social blocks, including social self-attention mechanism, traditional cross attention, multi-layer perceptron network, classification head or regression head. A social-like self-attention mechanism enables structured interaction and discriminative visual representation learning; Traditional cross-attention integrates textual information with condensed visual information in a fine-grained manner; Use a multi-layer perceptron network to complete the dimensionality denoising of features; The final multimodal representation is input into the classification head or regression head for downstream tasks of visual question answering and visual localization. A social-like self-attention mechanism enables structured interaction and discriminative visual representation learning; include: The quasi-social self-attention mechanism includes a dynamic structured interaction module and a discriminative pillar verification module. The dynamic structured interaction module uses the structural hole theory to realize the structured interaction of the visual area, and the discriminative pillar verification module redistributes weights for potential structural holes. First, the image is visually encoded using the pre-trained ResNext model to obtain the visual area representation matrix U b , where the visual area representation matrix U b Each row represents the representation of a visual area; the representation of the visual area is input into three linear fully connected layers f q ,f k and f v , get the corresponding query Q, key K and value V, and calculate the pairwise interaction score matrix A ir , as shown below: Among them, u i and u r are the i-th and r-th visual areas, respectively, and q i That is, f q (u i ), is the corresponding query, k r That is, f k (u r ), is the corresponding key, and d represents the dimension of the feature; Then, two modulation networks are used and Using feature u i and Connection Preference A i,: As input, an effective mask structure that helps form the structural hole is generated in the form of an affine transformation, as shown below: Where F = [A i,: ;u i ] represents the concatenation of features and connection preferences, R(A i,: ) represents A i,: The mean value, M i,: represents the learned mask; Using the obtained mask M i,: Generate hard mask H i,: , as shown below: H i,: =d(ReLU(M i,: )) Among them, d is the discretization operation; Again, the learned structured mask is used to correct the self-attention mechanism as follows: Among them, H ij is the discretization mask, A ij is the original interaction matrix, is the updated interaction matrix; use Update the original visual features; The key attribute information of the node is used to locate the potential structural holes, as shown below: w i =H i,: ⊙(s+A i,: ) Among them, H i,: ⊙s describes the importance of the node, H i,: ⊙A i,: It describes the transfer relationship of the relationship, so w i The conditions for each node to become a pillar node are described: whether it has a closer connection with more important nodes; at the same time, a more flexible and learnable form is adopted to increase the flexibility of pillar node determination, as shown below: Among them, s i is the global importance of the i-th node, p i To calculate the final node importance considering node degree information and transitive properties; Finally, the original visual features are further corrected using the pillar verification weights as follows: Among them, the importance of each node is combined to form an importance vector, and Represent the feature representation after pillar verification and the feature representation before respectively; p is a vector, each element of which is p i .

Citation Information

Patent Citations

  • Visual question and answer oriented method of context awareness based on multi-modal interaction

    CN114970517A