Composite graph construction method, system, device and storage medium for multimodal data
By constructing a composite graph in a social network and using the channel dynamic switching mechanism to perform modal information interaction, the problem of incomplete information in the multimodal data graph is solved, and more comprehensive information representation and improvement of deep learning tasks is achieved.
Patent Information
- Application Number
- CN202210438076.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-25
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-04-25
AI Technical Summary
When the prior art builds multimodal data graphs in social networks, it is difficult to effectively integrate and extract missing or incomplete information of modality, resulting in poor deep abstract semantic mining of information.
A composite graph construction method for multimodal data is proposed. By acquiring user data and resource data, using the channel dynamic switching mechanism to perform information interaction between modes, generating a more comprehensive second modal feature embedding matrix, and constructing a composite graph based on user features and resource relationships.
It realizes a more comprehensive representation of the information of users and resources in social networks, improves the training effect of deep learning tasks such as semantic analysis and sentiment analysis, and enhances the ability of information integration and extraction.
Smart Images

Figure CN114911979B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a method, system, device and storage medium for constructing a composite graph of multimodal data. Background Art
[0002] As the functions of social platforms become more and more complete, social networks have become an important part of people's daily lives. When users are active on social networks, they will have a series of interactive behaviors with texts, pictures, videos and other modalities on the network, thus generating trillions of bytes of social data every day. These different modal data not only differ in form and data volume, but also have complex and diverse data structures. The intricate social data increases the difficulty of information integration and extraction.
[0003] Generally speaking, the more modalities in the same resource, the larger the amount of data it contains, and the better it can reflect the real situation of the resource. However, due to external factors such as network speed, resource content updates, insufficient memory, or hardware damage, modal information may be lost during resource use in social networks, resulting in incomplete modal data. It is crucial to integrate and extract valid information as much as possible from the missing modal information.
[0004] Relevant research shows that the basis for analyzing social networks is graph theory. Graph data contains very rich relational information, and can be used for inference and learning from unstructured data such as text, pictures, and videos. The way the graph is constructed has an important impact on subsequent task analysis. Using a graph structure to represent information can achieve the conversion of unstructured data to structured data. Graph neural networks can be used to extract and integrate information from graph structures. For example, in semantic analysis tasks, inputting the information contained in the graph structure into the graph neural network for training can make it easier to learn the deep semantics between users and resources, thereby improving the comprehensiveness and accuracy of semantic analysis tasks.
[0005] However, the information contained in the graph structure obtained by most current graph neural network models during graph construction is incomplete information with missing modalities, which cannot effectively mine the deep abstract semantics of the information, bringing negative impacts to subsequent tasks. Summary of the invention
[0006] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, the present invention proposes a composite graph construction method, system, device and storage medium of multimodal data, and the constructed composite graph more comprehensively represents various information of users and social resources in social networks.
[0007] On the one hand, an embodiment of the present invention provides a method for constructing a composite graph of multimodal data, comprising the following steps:
[0008] Acquire a user data set, a resource data set, and a relationship between a user and a resource, wherein the resource data set includes a plurality of resource data, and the resource data includes a plurality of modal data;
[0009] Determine a plurality of user feature embedding matrices according to the user data set;
[0010] Processing the plurality of modal data respectively to obtain a plurality of first modal feature embedding matrices;
[0011] Based on the channel dynamic exchange mechanism, multiple second modal feature embedding matrices are obtained by performing inter-modal information interaction according to multiple first modal feature embedding matrices of each resource data;
[0012] A composite graph is constructed according to the user feature embedding matrix, multiple second modality feature embedding matrices of resource data, and the relationship between the user and resources.
[0013] According to some embodiments of the present invention, the method for constructing a composite graph of multimodal data further includes the following steps:
[0014] Determine the similarity between users according to the plurality of user feature embedding matrices, and obtain the relationship between users according to the similarity between users;
[0015] updating the user feature embedding matrix according to the relationship between users, the relationship between user resources and a plurality of the second modality feature embedding matrices;
[0016] Determine the similarity between resources according to the plurality of second modality feature embedding matrices, and obtain the relationship between resources according to the similarity between resources;
[0017] Determine a complementary information feature embedding matrix between resources based on the relationship between the resources and multiple second modality feature embedding matrices of resource data;
[0018] The composite graph is constructed according to the updated user feature embedding matrix and the complementary information feature embedding matrix.
[0019] According to some embodiments of the present invention, determining a plurality of user feature embedding matrices according to the user data set comprises the following steps:
[0020] Granulating the user data set according to the information metric function to obtain an information granule set, wherein the information granule set includes a plurality of user information granules;
[0021] The information particle set is integrated into a multi-head attention mechanism and a common common sense knowledge base mechanism for abstract feature representation to obtain the user feature embedding matrix.
[0022] According to some embodiments of the present invention, the processing of the plurality of modal data to obtain a plurality of first modal feature embedding matrices comprises the following steps:
[0023] For text modal data, a Text-Transformer encoder and a common knowledge base fusion mechanism are used to perform feature representation to obtain a first modal feature embedding matrix of the text modal data, wherein the text modal data is represented as a word embedding sequence, absolute position information is integrated into the word embedding, and the word embedding is integrated into a first common knowledge base feature embedding matrix corresponding to the absolute position information according to the absolute position information, and the relative position information is integrated into the word embedding according to the relative distance between the absolute position information and the word embedding, to obtain the first modal feature embedding matrix of the text modal data;
[0024] For image modality data, a CNN+Visual-Transformer encoder and a common knowledge base fusion mechanism are used for feature representation to obtain a first modality feature embedding matrix of the image modality data, wherein the image modality data is subjected to a convolution pooling operation to obtain multiple first image feature embedding matrices, and a position code and a second common knowledge knowledge base feature embedding matrix are integrated into each first image feature embedding matrix to obtain a second image feature embedding matrix, and the second image feature embedding matrix is input into a Visual-Transformer encoder based on a multi-head attention mechanism for encoding to obtain a first modality feature embedding matrix of the image modality data;
[0025] For video modality data, a Video-Transformer encoder is used in combination with a common sense knowledge base fusion mechanism for feature representation to obtain a first modality feature embedding matrix of the video modality data, wherein the video modality data is input into a first layer encoder based on a standard multi-head self-attention mechanism for encoding to obtain a video feature embedding matrix, and the video feature embedding matrix is input into a second layer encoder based on a multi-head attention mechanism with a local sensing bias block for encoding to obtain a first modality feature embedding matrix of the video modality data, wherein the multi-head attention mechanism with a local sensing bias block is determined by the local sensing bias block position information and the second common sense knowledge base feature embedding matrix.
[0026] According to some embodiments of the present invention, the channel dynamic exchange mechanism is based on each resource data multiple first modal feature embedding matrix to perform inter-modal information interaction to obtain multiple second modal feature embedding matrices, including the following steps:
[0027] Allocating values at different positions in each of the first modal features embedding matrix into different channels;
[0028] The second modal feature embedding matrix is obtained by performing channel exchange on the weight values of the modal fusion results according to the numerical values in the channels within each of the first modal feature embedding matrices, wherein, if the current first modal feature embedding matrix contains a numerical value whose weight value is less than the fusion threshold, the channel corresponding to the numerical value is determined as the channel to be mapped, and the average value of the channels in other first modal feature embedding matrices with the same position as the channel to be mapped is substituted into the channel to be mapped.
[0029] According to some embodiments of the present invention, determining the similarity between resources according to the plurality of second modality feature embedding matrices, and obtaining the relationship between resources according to the similarity between resources comprises the following steps:
[0030] Transversely splicing the multiple second modal feature embedding matrices of each resource data to obtain a multimodal representation matrix;
[0031] Extracting the maximum value of each row in the multimodal representation matrix to obtain a first vector, extracting the minimum value of each row in the multimodal representation matrix to obtain a second vector, and extracting the average of each row in the multimodal representation matrix to obtain a third vector;
[0032] Concatenating the first vector, the second vector, and the third vector to obtain a multimodal feature embedding matrix of each resource data;
[0033] The similarity between resources is determined according to the multimodal feature embedding matrix of different resource data, and the relationship between resources is obtained according to the similarity between resources.
[0034] According to some embodiments of the present invention, the feature embedding of the composite graph is represented as:
[0035] Message=(e u , e u,f , e f,m , e f1,f2 |u∈U, f, f1, f2∈F, m∈N + );
[0036] Among them, e u ∈R |U|×d Used to characterize user data, d is the d-dimensional feature of user data, e u,f ∈R |F|×d The feature embedding matrix used to characterize user u under resource f, e f,m ∈R |F|×m It is used to characterize each modal data in resource f, m is the m-dimensional feature of the modal data in resource f, and e f1,f2 ∈R |F|×|F| Used to represent the complementary information between resource f1 and resource f2.
[0037] On the other hand, an embodiment of the present invention further provides a composite graph construction system for multimodal data, including:
[0038] The first module is used to obtain a user data set, a resource data set, and a relationship between a user and a resource, wherein the resource data set includes a plurality of resource data, and the resource data includes a plurality of modal data;
[0039] A second module is used to determine a plurality of user feature embedding matrices according to the user data set;
[0040] A third module is used to process the multiple modal data respectively to obtain multiple first modal feature embedding matrices;
[0041] A fourth module is used to obtain multiple second modal feature embedding matrices by performing inter-modal information interaction based on multiple first modal feature embedding matrices of each resource data based on a channel dynamic exchange mechanism;
[0042] The fifth module is used to construct a composite graph based on the user feature embedding matrix, multiple second modality feature embedding matrices of resource data and the relationship between the user and resources.
[0043] On the other hand, an embodiment of the present invention further provides a composite graph construction device for multimodal data, comprising:
[0044] at least one processor;
[0045] at least one memory for storing at least one program;
[0046] When the at least one program is executed by the at least one processor, the at least one processor implements the composite graph construction method of multimodal data as described above.
[0047] On the other hand, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the method for constructing a composite graph of multimodal data as described above.
[0048] The above-mentioned technical scheme of the present invention has at least one of the following advantages or beneficial effects: the present application obtains various user data and resource data on the social network, and then exchanges information between different modal data in the resource data based on the channel dynamic exchange mechanism, thereby improving the information in each modal data, and constructs a composite graph by improving the resource data, user data and the relationship between user resources after the modal data is improved, so that the composite graph more comprehensively represents various information of users and social resources in the social network, and is more likely to converge in deep learning tasks such as semantic analysis and sentiment analysis based on the composite graph, thereby improving the training effects of various task models. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 is a flow chart of a method for constructing a composite graph of multimodal data provided by an embodiment of the present invention;
[0050] Figure 2 is a schematic diagram of a composite graph construction system for multimodal data provided by an embodiment of the present invention;
[0051] Figure 3 is a schematic diagram of a composite graph construction device for multimodal data provided by an embodiment of the present invention;
[0052] Figure 4 is a schematic diagram of a heterogeneous composite graph provided by an embodiment of the present invention;
[0053] Figure 5 It is a schematic diagram of the user data processing and resource data processing process provided by an embodiment of the present invention;
[0054] Figure 6 It is a schematic diagram of the process of constructing a composite graph of multimodal data provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0055] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar components or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.
[0056] In the description of the present invention, it should be understood that descriptions involving orientation, such as up, down, left, right, etc., the orientations or positional relationships indicated are based on the orientations or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present invention.
[0057] In the description of the present invention, if there is a description of first, second, etc., it is only for the purpose of distinguishing technical features, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the sequence of the indicated technical features.
[0058] The embodiment of the present invention provides a method for constructing a composite graph of multimodal data, referring to Figure 1 The method for constructing a composite graph of multimodal data of an embodiment of the present invention includes but is not limited to step S110, step S120, step S130, step S140 and step S150.
[0059] Step S110, obtaining a user data set, a resource data set, and a relationship between a user and a resource, wherein the resource data set includes a plurality of resource data, and the resource data includes a plurality of modal data;
[0060] Step S120, determining a plurality of user feature embedding matrices according to the user data set;
[0061] Step S130, processing the multiple modal data respectively to obtain multiple first modal feature embedding matrices;
[0062] Step S140, based on the channel dynamic exchange mechanism, performing inter-modal information interaction according to the multiple first modal feature embedding matrices of each resource data to obtain multiple second modal feature embedding matrices;
[0063] Step S150, constructing a composite graph according to the user feature embedding matrix, multiple second modality feature embedding matrices of resource data, and the relationship between users and resources.
[0064] In this embodiment, various user data and resource data on the social network are obtained, and then information is interacted between different modal data in the resource data based on the channel dynamic exchange mechanism, so as to improve the information in each modal data. By improving the resource data, user data and the relationship between user resources after the modal data is improved, a composite graph is constructed with users as vertices, resources as super-vertices, and the relationship between users and resources as edges, so that the composite graph more comprehensively represents various information of users and social resources in the social network. Based on the composite graph, it is easier to converge in deep learning tasks such as semantic analysis and sentiment analysis, thereby improving the training effect of various task models.
[0065] According to some specific embodiments of the present invention, the method for constructing a composite graph of multimodal data further includes but is not limited to the following steps:
[0066] Step S210, determining the similarity between users according to multiple user feature embedding matrices, and obtaining the relationship between users according to the similarity between users;
[0067] Step S220, updating the user feature embedding matrix according to the relationship between users, the relationship between user resources and multiple second modality feature embedding matrices;
[0068] Step S230, determining the similarity between resources according to the plurality of second modal feature embedding matrices, and obtaining the relationship between resources according to the similarity between resources;
[0069] Step S240, determining a complementary information feature embedding matrix between resources according to the relationship between resources and multiple second modality feature embedding matrices of resource data;
[0070] Step S250, constructing the composite graph according to the updated user feature embedding matrix and the complementary information feature embedding matrix.
[0071] Specifically, in social networks, a large amount of interactive information is generated between users and resources, such as user browsing, searching and other operations. The same resource usually contains information of different modalities. For example, in the use of an online course, users may see a large number of picture modalities during the search process, and there are video modalities during the viewing process. Auxiliary learning methods will be exposed to modal information such as text and voice. These interactive data are usually described as user-resource bipartite graphs, and user-resource bipartite graphs only consider the information between users and resources, which will cause the loss of information between users or between resources. Therefore, the embodiment of the present invention constructs a heterogeneous composite graph of various elements in social data to represent user data and resource data, as well as the relationship between data.
[0072] Reference Figure 4 , the heterogeneous composite graph is represented as:
[0073] G = (v, E);
[0074] Where ν = U∪F, which represents the point set containing user vertices and resource vertices. Where U = {u1, u2, ...u K}, F = {f1, f2, ... f M} represent the user vertex set and resource vertex set respectively, K represents the total number of users, M represents the total number of resources, and s i ={m1, m2, ..., m j |i∈M,m∈N +} indicates that the i-th resource is internally composed of m j E = E1 ∪ E2 ∪ E3 represents different edge relationships, E1 = {(u1, u2) | u1∈U, u2∈U} represents the implicit relationship between users, E2 = {(u, f) | u∈U, f∈F} represents the interaction information between users and resources, and E3 = {(f1, f2) | f1∈F, f2∈F} represents the hidden relationship information between resources.
[0075] After extracting features from user data, we get the user feature embedding matrix. After extracting features from resource data and complementing the data between modalities based on the dynamic channel exchange mechanism, we get the second modality feature embedding matrix. Then, we update the user feature embedding matrix by comprehensively considering the explicit edge relationship E2 and the implicit edge relationship E1. We get the complementary information feature embedding matrix between resources by considering the implicit edge relationship E3. Then, we construct a heterogeneous composite graph based on various embedding matrices. The feature embedding of the composite graph is expressed as:
[0076] Message=(e u , e u,f , e f,m , e f1,f2|u∈U, f, f1, f2∈F, m∈N + );
[0077] Among them, e u ∈R |U|×d is the user feature embedding matrix, which is used to represent user data. d is the d-dimensional feature of user data. Data is updated through implicit edge relationship E1. u,f ∈R |F|×d The feature embedding matrix used to characterize user u under resource f can be directly obtained through the explicit edge relationship E2; f,m ∈R |F|×m is the second modal feature embedding matrix, which is used to characterize each modal data in resource f, and m is the m-dimensional feature of the modal data in resource f; e f1,f2 ∈R |F|×|F| It is a complementary information feature embedding matrix, which is used to represent the complementary information between resources f1 and f2, and to update data through implicit edge relationship E3.
[0078] In this embodiment, different edge relationships are obtained by similarity calculation based on user characteristics and resource characteristics, and the missing information and unimportant information of the resource modality are further integrated in combination with the channel dynamic exchange mechanism, and all the obtained information is integrated to construct a heterogeneous composite graph. The heterogeneous composite graph contains three types of implicit relationships, namely, similarity relationships between users, complementary relationships between resource internal modalities, and hidden relationships between resources, as well as one type of explicit relationship, namely, interactive information between users and resources. Compared with the graph structure data in the related art, the heterogeneous composite graph of the embodiment of the present invention more fully considers the coordination relationship between various resources inside and outside, and integrates the incomplete modal data in the process of constructing the graph, so that the data contained in the obtained heterogeneous composite graph is more complete.
[0079] According to some specific embodiments of the present invention, step S120 includes but is not limited to the following steps:
[0080] Granulating the user data set according to the information metric function to obtain an information granule set, wherein the information granule set includes a plurality of user information granules;
[0081] The information particle set is integrated into the multi-head attention mechanism and the public common sense knowledge base mechanism to perform abstract feature representation and obtain the user feature embedding matrix.
[0082] Specifically, refer to Figure 5 ,When learning the deep abstract semantics of user data sets, we first design a ,user feature granulation strategy based on the information metric function to ,realize the merging and granulation of user data record level, ,simplify attributes and eliminate redundant information.
[0083] The granulation process uses a more classic neighborhood rough set, which can flexibly use the radius information of the neighborhood to make similarity judgments on user samples in the task space. In order to speed up the information granulation process, a more efficient design method is adopted, that is, assuming that the decision system is DS=<U,AT∪{d}> , the neighborhood radius is δ, x j ∈U, if So that abs(a(x i )-a(x j ))>δ, then where U = {x1, x2...x n} is a non-empty finite set consisting of n users, x n represents the nth user, AT is a non-empty finite set of attributes, d represents the decision attribute, A={a1,a2...,a m} represents a reduction that satisfies the constraints, and a(x) represents the value of user sample x on a.
[0084] The granulation strategy of the embodiment of the present invention can reduce the number of users that need to be traversed in the traditional granulation strategy, directly filter out users that do not necessarily belong to the neighborhood space, thereby reducing the sample space that needs to be considered in the granulation process, achieving the effect of quickly realizing the granulation of neighborhood information, and realizing dimensionality reduction processing on user data. The information granule set G can be expressed as G = {δ(x1), δ(x2), ..., δ(x n )}.
[0085] User data can be recorded in text mode. In order to mine the deep abstract semantics of different user data, the information granule set G is integrated with multi-head attention and public common sense knowledge base mechanism to perform abstract feature representation. The integration of public common sense knowledge base with real semantic information can make the deep abstract semantic learning process more stable and can also automatically mine the text features required for specific tasks.
[0086] First, each type of user data information, such as age, gender, login time, social time, region, etc., is regarded as a node in the graph. Then, by initializing the weight matrix, different weight values are assigned to each node. During the training process, the multi-head attention mechanism aggregates the information of the target node and the neighboring node according to the weight value of the assigned node when updating the node hidden layer, and weights the information contained in the public common sense knowledge base, focusing on semantic related information. The calculation process is shown in formula (1):
[0087]
[0088] Among them, W is a shared parameter, I *represents the d-dimensional public common sense knowledge base information matrix in the specific task*, [·||·||·] represents the feature matrix after the concatenation of vertices i, j and the corresponding public knowledge base information, a(·) represents mapping the feature to another real number, α i,j represents the attention coefficient from node i to node j after integrating the common common sense knowledge base. For the k introduced independent attention mechanisms, in order to further aggregate the node and the node's neighbor information, K-average is used to replace the connection, as shown in formula (2):
[0089]
[0090] Among them, h i is the feature information after node i completes the information fusion process for neighboring nodes, K represents the sequence number of the independent attention mechanism, and I * represents the public common sense knowledge base information matrix in a specific task*, σ(·) is the activation function, is the attention coefficient of node i to node j in the kth independent attention mechanism. The training process of integrating graph attention and common common sense knowledge base mechanism through L layers is shown in formula (3):
[0091]
[0092] in, is the feature representation corresponding to node i in layer l, Θ l-1 is the set of trainable parameters of GAT in the l-1th layer, A is the feature embedding, I * represents the d-dimensional public common sense knowledge base information matrix in a specific task*. Through the above process, user data, i.e. user information particles, can be converted into an embedding matrix e′ u ∈R |U|×d , d is the d-dimensional feature of user information, completing the abstract feature representation of user data and obtaining the user feature embedding matrix.
[0093] According to some specific embodiments of the present invention, step S130 includes but is not limited to the following steps:
[0094] Step S131: For text modal data, feature representation is performed using a Text-Transformer encoder and a public common sense knowledge base fusion mechanism to obtain a first modal feature embedding matrix of the text modal data, wherein the text modal data is represented as a word embedding sequence, absolute position information is incorporated into the word embedding, and the word embedding is integrated into a first common sense knowledge base feature embedding matrix corresponding to the absolute position information according to the absolute position information; relative position information is integrated into the word embedding according to the relative distance between the absolute position information and the word embedding, to obtain the first modal feature embedding matrix of the text modal data.
[0095] Specifically, refer to Figure 5 , for the text modality of resource data, the Text-Transformer model is used to integrate with the public common sense knowledge base. Considering that the dependency between tags will gradually weaken as the distance between text content increases, relative position and absolute position information are introduced in each encoding layer of the model to mark the order of words in the text. Among them, the encoding of absolute position and relative position information is trainable, and the relative position information will be injected into the attention calculation process.
[0096] First, in the process of embedding absolute position information, assume represents a sequence with N input tokens, w i is the i-th element, S N The corresponding word embedding can be expressed as x i ∈R is the wth i The self-attention mechanism will first integrate the position information into the word embedding and express it in the form of query, key and value. The representation forms are shown in formula (4):
[0097] q m =f q (x m , m);
[0098] k n =f k (x n , n);
[0099] v n =f v (x n ,n); (4)
[0100] Among them, q m , k n , v n Represent query, key, and value matrices respectively. To change the position embedding information, it is mainly necessary to select a suitable function form for formula (4).
[0101] The process design function of incorporating the common knowledge base into the absolute position information embedding process is shown in formula (5):
[0102] f l:l∈{q,k,v} (x i , i): =W l:l∈{q,k,v} (x i +p i +I i ); (5)
[0103] Among them, p i ∈Rd is a i The trainable d-dimensional matrix determined by the position information of p is calculated using the sine function i , I i is a i The d-dimensional public common sense knowledge base feature embedding determined by the position information of l:l∈{q,k,v} is the trainable weight matrix in the corresponding calculation formula.
[0104] Then, in the relative position information embedding process, a function setting different from formula (5) is used, and the specific calculations are shown in formula (6):
[0105] f q (x m ):=W q x m I m ;
[0106]
[0107]
[0108] in, is the trainable relative position information embedding, γ = clip(mn, γ min , γ max ) is the relative distance between positions m and n, and a threshold m,n∈[γ min , γ max ], beyond which the relative position information will lose its meaning. q , W k , W v is the trainable weight matrix in the corresponding function, I m is a m The d-dimensional public common sense knowledge base feature embedding is determined by the position information.
[0109] Based on the relative position acquisition formula (6), in order to obtain the absolute position of the text content, so as to realize the fusion information embedding of absolute position information and relative position information, the query matrix Decompose it into the following formula (7):
[0110]
[0111] Among them, To replace the absolute position embedding of position m and position n, W q , W kis the trainable weight matrix in the corresponding function. The self-attention mechanism used in the Text-Transformer encoder is described more generally, and the first modality feature embedding matrix for text modality data that integrates absolute position information, relative position information, and common common sense knowledge base information can be obtained. The first modality feature embedding matrix is shown in the following formula (8):
[0112]
[0113] in is an orthogonal matrix, represents a non-negative function, I * It is the d-dimensional public common sense knowledge base feature embedding in the specific task*. Using the above mechanism can highlight the values in the text modality that are more matched with the public common sense knowledge base information, thereby realizing the deep abstract semantic information learning of text information.
[0114] Step S132: For the image modal data, a CNN+Visual-Transformer encoder is used in combination with a common knowledge base fusion mechanism to perform feature representation, and a first modal feature embedding matrix of the image modal data is obtained, wherein the image modal data is subjected to a convolution pooling operation to obtain multiple first image feature embedding matrices, and position encoding and a second common knowledge base feature embedding matrix are integrated into each first image feature embedding matrix to obtain a second image feature embedding matrix, and the second image feature embedding matrix is input into a Visual-Transformer encoder based on a multi-head attention mechanism for encoding to obtain a first modal feature embedding matrix of the image modal data.
[0115] Specifically, refer to Figure 5 , for the image modality in the resource data, a CNN+Visual-Transformer encoder and a common knowledge base fusion mechanism are used. CNN is mainly used to learn the bias information of the image space. It can use rolling convolution kernels to perform convolution operations at different levels and realize feature extraction based on comprehensive consideration of local and global features of similar data. The feature extraction process is shown in the following formula (8):
[0116]
[0117] in, Represents the output result after convolution, w d represents the convolution kernel of size d, V i represents the eigenvalue corresponding to the input node i, b i is the bias term, f(·) represents the activation function, and the ReLU activation function is usually used.
[0118] After the convolution operation, the maximum pooling operation is added, that is, each feature map is divided into multiple regions. For a region Rd, the maximum activity value of all neurons in this region is selected as the new representation of this region. The pooling process is shown in the following formula (10):
[0119]
[0120] Among them, x i For Region Adding pooling operations can reduce the number of features and thus the number of parameters.
[0121] After completing the convolution and pooling operations to obtain the first image feature embedding matrix, adding the Visual-Transformer encoder and the public common sense knowledge base fusion mechanism as a supplementary component of CNN can improve the stability of the model and alleviate the long-range dependency problem of abstract semantic learning.
[0122] In the Visual-Transformer structure, the first image feature embedding matrix obtained after the pooling operation is recorded as Each first image feature embedding matrix x s Add a learnable positional encoding and the d-dimensional public common sense knowledge base information matrix I under the specific task * The second image feature embedding matrix is obtained, and the fusion process is shown in formula (11):
[0123]
[0124] The size of the position encoding is the same as the size of the image embedding.
[0125] Since a sufficient number of multi-head self-attention layers can directly simulate the interaction of short-distance and long-distance information and realize the learning of semantic associations between different parts of the image modality, each encoder in the Visual-Transformer encoder consists of a self-attention structure with H heads. For each encoder layer l∈{1, 2, 3, ..., H}, each attention head h∈{1, 2, 3, ..., H}, in order to prevent overfitting during training, regularization technology is introduced and the learnable parameter matrix As query, key, and value, the calculation process is shown in formula (12):
[0126]
[0127] Then scale the value D by the dimension of the attention head H=D / H, the output of the network layer is shown in formula (13):
[0128]
[0129] Among them, I * Represents the d-dimensional common common sense knowledge base information matrix in the specific task*. After obtaining the attention coefficient value, a two-layer fully connected neural network (MLP) is connected to obtain the final first modality feature embedding matrix for image modality data to realize the abstract semantic information extraction of image data.
[0130] Step S133, for the video modality data, use the Video-Transformer encoder and the public common sense knowledge base fusion mechanism to perform feature representation, and obtain the first modality feature embedding matrix of the video modality data, wherein the video modality data is input into the first layer encoder based on the standard multi-head self-attention mechanism for encoding to obtain the video feature embedding matrix, and the video feature embedding matrix is input into the second layer encoder based on the multi-head attention mechanism with a local sensing bias block for encoding to obtain the first modality feature embedding matrix of the video modality data, and the above-mentioned multi-head attention mechanism based on the local sensing bias block is determined by the local sensing bias block position information and the second common sense knowledge base feature embedding matrix.
[0131] Specifically, refer to Figure 5 , for the video modality of resource data, in order to make full use of the spatiotemporal position information in the resource video modality and determine the characteristics that pixels that are close in spatiotemporal distance are more likely to be associated with each other, a Video-Transformer encoder that combines spatial locality and translation invariance with a common sense knowledge base fusion mechanism is used to process the video modality. Unlike ordinary Transformer encoders for processing videos, Video-Transformer does not reduce sampling over the time dimension. In order to process the temporal information in the video modality, a local sensing bias is introduced in the multi-head self-attention module of each layer of the standard Transformer encoder. The input video size is denoted as S×H×W×3, that is, the frame S contains an image with a height of H, a width of W, and 3 RGB channels. The size of each local sensing bias block in the video is marked as S′×H′×W×3, S′≤S, H′≤H, W′≤W. To avoid the local sensing bias block being too large or too small, the size of the local sensing bias block is usually set to In this way, local information of the picture is captured. The position information of the local sensing bias block of each frame needs to be determined by the local sensing bias block to identify the main body of the picture in each frame. The local sensing bias block will determine the main body of the picture based on the degree of proximity and similarity with the pixel point, thereby capturing the key information of each frame and obtaining the position information.
[0132] In the Video-Transformer encoder, the standard multi-head self-attention mechanism and the multi-head attention mechanism with a local sensing bias block are combined, and then a regularization layer and a two-layer fully connected neural network (MLP) are introduced to enhance the stability of the model. The training process of the Video-Transformer encoder of the embodiment of the present invention is shown in formula combination (14):
[0133]
[0134] in, They represent the feature output after the multi-head self-attention mechanism and the feature output after two layers of fully connected neural network (MLP), respectively. MSA represents the standard multi-head self-attention mechanism, and CMSA represents the multi-head self-attention mechanism with the introduction of local induction bias.
[0135] Due to the introduction of bias and common common sense knowledge base, the parameter calculation formula of the attention mechanism also changes accordingly. The calculation process is shown in formula (15):
[0136]
[0137] Among them, Q, K, V ∈ R S×d represents the query, key, value matrix, b represents the deviation matrix for generating the position information of the local sensing bias block, I * Representing the d-dimensional public common sense knowledge base information matrix in the specific task*, the abstract semantic information of the video modality is extracted by inputting the video modality data into the Video-Transformer encoder of the embodiment of the present invention, and the first modality feature embedding matrix for the video modality is obtained.
[0138] According to some specific embodiments of the present invention, step S140 includes but is not limited to the following steps:
[0139] Assigning values at different positions in each first modal feature embedding matrix to different channels;
[0140] The second modal feature embedding matrix is obtained by performing channel exchange on the weight values of the modal fusion results according to the numerical values in the channels within each first modal feature embedding matrix, wherein, if there is a weight value less than the fusion threshold in the current first modal feature embedding matrix, the channel corresponding to the value is determined as the channel to be mapped, and the average value of the channels in the same position as the channel to be mapped in other first modal feature embedding matrices is substituted into the channel to be mapped.
[0141] Specifically, after various modal data are input into the corresponding Transformer encoder and integrated with the common sense knowledge base fusion mechanism to obtain various first modal feature embedding matrices, the length of the obtained first modal feature embedding matrices is not uniform due to the different data volumes and expressions of different modalities.
[0142] Note e′ f,m ∈R |F|×m Indicates the embedding of different modal data in the same resource data, based on the channel dynamic exchange mechanism, does not change the maximum value of the embedding dimension, and performs feature fusion on the embedding matrices of different modalities, while integrating the common information of different modalities, it also retains the specific form of each modality. That is, the values of different positions in each first modal feature embedding matrix of the same resource data are regarded as being in different channels, and then the importance of the values in the channel is calculated to determine whether to perform channel exchange. Specifically, in a certain first modal feature embedding matrix, if the numerical weight value in a channel is less than the fusion threshold, indicating that the numerical value has a small impact on the modal fusion result, then the numerical value of the channel is replaced with the average value of the channels at the same position of other modalities. It can be understood that the channels with missing information in the modality will also be filled with the average value of the channels at the same position of other modalities. Through the channel exchange mechanism of the embodiment of the present invention, the corresponding channels between different modal data layers will be embedded in the same mapping, thereby further compressing the multimodal structure and making it easier to construct modal common feature data. In addition, learnable parameters can also be set in the channel exchange mechanism so that the channel exchange process of the modality is dynamically adaptive. During the dynamic exchange of channels of the modalities, directional channel exchange is restricted to a specific channel range of each modality to ensure that data processing is performed within the modality only in the exact resource.
[0143] The information interaction process between modalities based on the channel exchange mechanism is shown in formula (16):
[0144]
[0145] Among them, x m,t,c represents the cth channel in the tth layer feature map of the mth modality data in the resource data, δ m,t represents the variance of the intra-channel information in the t-th layer feature map of the m-th modal data, ρ m,t represents the standard deviation of the intra-channel information in the t-th layer feature map of the m-th modal data, γ m,t,c , β m,t,c denote the trainable parameters and bias matrix respectively, ε is a custom constant to avoid the denominator being 0, θ is a threshold close to 0+ manually set according to task requirements, and the t+1 layer converts {x′ m,t,c} as new input.
[0146] In the process of information interaction between modalities, γ m,t,c To measure the input channel information x during training m,t,c And output channel information x′ m,t,c If the correlation between γ m,t,c tends to 0, then the loss function used in the model will also tend to 0, which means x m,t,c The value in the channel has little or no effect on the final prediction result. mt,c The values in the channel are taken as redundant values, and if γ m,t,c If the value in the current channel is smaller than the set fusion threshold θ, the value in the current channel is replaced by the average value of the corresponding channel of other modalities. Through the above process, a new second modality feature embedding matrix e for information interaction between modalities within a resource can be obtained. f,m ∈R |F|×m .
[0147] According to some specific embodiments of the present invention, the relationship between users in step S210 is determined in the following manner:
[0148] After getting the user feature embedding matrix e′ u ∈R |U|×d and the second modality feature embedding matrix e of the resource data f,m ∈R |F|×d After that, the implicit relationship between users is mined through the cosine similarity between multiple user feature embedding matrices. The similarity between users is determined by formula (17):
[0149]
[0150] Among them, A i , B i Represent two different user vertices, ||·|| u Indicates l u Regularization.
[0151] Then, the relationship between users is determined by a preset user similarity threshold, that is, the edge E1 between user vertices. For example, the user similarity threshold can be set to 0.85. If the similarity between users exceeds 0.85, the user vertex A is determined to be i and B i There is a similarity relationship, that is, there is an edge relationship E1.
[0152] Furthermore, if A i , B i There is an edge relationship E1 between the two vertices. When determining the user vertex A i Based on the feature embedding of user vertex A iIts adjacent user vertex B i The interaction relationship between user vertex A i The second modal feature embedding matrix of the resource data interacting with it, for vertex A i The missing information in the modality that appears in the matrix can be supplemented to obtain the updated user feature embedding matrix e u ∈R |U|×d .
[0153] According to some specific embodiments of the present invention, step S230 includes but is not limited to the following steps:
[0154] Embedding matrices of multiple second modal features of each resource data are horizontally spliced to obtain a multimodal representation matrix;
[0155] Extracting the maximum value of each row in the multimodal representation matrix to obtain a first vector, extracting the minimum value of each row in the multimodal representation matrix to obtain a second vector, and extracting the average of each row in the multimodal representation matrix to obtain a third vector;
[0156] Concatenate the first vector, the second vector, and the third vector to obtain a multimodal feature embedding matrix of each resource data;
[0157] The similarity between resources is determined according to the multimodal feature embedding matrix of different resource data, and the relationship between resources is obtained according to the similarity between resources.
[0158] Specifically, in order to mine edge information between resources, that is, between super-vertices, it is necessary to effectively utilize the embedded information of various modalities within the super-vertices. Some super-vertices with incomplete internal modal data can obtain relatively complete modal data after being processed by the channel dynamic exchange method.
[0159] Before calculating the similarity between resources, information integration can be performed on multiple second modality feature embedding matrices within each super vertex.
[0160] First, the second modal feature embedding matrix inside the super vertex is horizontally spliced to strengthen the internal connection between modal data. After splicing, the multimodal representation matrix is obtained as follows:
[0161]
[0162] Where n is the total number of dimensions of the second modal feature embedding matrices in the super-vertex after the concatenation operation is completed, m is determined by the second modal feature embedding matrix with the longest data in the super-vertex, and the multimodal representation matrix P∈R m×n If there is still missing information in the modal data after the channel dynamic exchange method, the vacant position is filled with 0.
[0163] After completing the concatenation operation, a search algorithm is used to determine the maximum, minimum, and average values of each row in the multimodal representation matrix P. The code-based search algorithm is shown in formula (18):
[0164] M i =max(P i );
[0165] m i =min(P i );
[0166]
[0167] Among them, i∈{1, 2, 3, …, m}, j∈{1, 2, 3, …, n}, P i represents the i-th row of the multimodal representation matrix P, max(·) represents the row vector maximum value search algorithm in the implementation code, min(·) represents the row vector minimum value search algorithm in the implementation code, Ave(·) represents the row vector averaging algorithm in the implementation code and the formula is given.
[0168] The maximum values M1, M2, ...M of each row are extracted respectively. m Construct column vector A Max , and A Max ∈R m×1 The minimum value of each row is m1, m2, ...m m Construct column vector A Min , and A Min ∈R m×1 . Take the average value of each row A1, A2, ...A m Construct column vector A Min And A Min ∈R m×1 .
[0169] The column vector A Max , A Min and A Min Spliced together, the new embedding of the i-th super vertex, that is, the multimodal feature embedding matrix, is expressed as and
[0170] Assign a learnable parameter matrix W∈R to each super vertex 3×n , and then the parameters are learned through a two-layer fully connected neural network (MLP). The multimodal feature embedding matrix update process of each super vertex is shown in formula (19):
[0171]
[0172] Then, the similarity between resources, that is, the similarity between super vertices, is calculated according to formula (20):
[0173]
[0174] Among them, K i , K j Represent the super vertex V i 、V j The number of modes included, w k represents the weight parameter, M k , N k They correspond to the super vertex V i and V i The updated multimodal feature embedding matrix.
[0175] According to the preset resource similarity threshold Y∈{0,1}, the relationship between resources is determined, that is, whether there is a similar edge E3 between super vertices. Through the information transmission of edge E3, the incomplete information of different resource modalities can be supplemented again, so as to obtain the complementary information feature embedding matrix e between resource f1 and resource f2 f1,f2 ∈R |F|×|F| .
[0176] According to some specific embodiments of the present invention, Figure 6 , obtain user data, resource data, and the interaction relationship between users and resources, calculate the user feature embedding matrix of each user according to the user data, calculate various second modal feature embedding matrices of each resource according to the resource data based on the channel dynamic exchange mechanism, and then determine the relationship between users through the similarity between users, update the user feature embedding matrix according to the relationship between users, the interaction relationship between users and resources, and the second modal feature embedding matrix of the corresponding user's resource data, determine the relationship between resources through the similarity between resources and determine the complementary information feature embedding matrix according to the relationship between resources, and construct a heterogeneous composite graph according to the user feature embedding matrix, the complementary information feature embedding matrix, the second modal feature embedding matrix, the relationship between resources, the relationship between users, and the interaction relationship between users and resources. Among them, each user is recorded as a common node of the heterogeneous composite graph, each resource is recorded as a supernode of the heterogeneous composite graph, the edges between supervertices are constructed based on the relationship between resources, the edges of common nodes are constructed based on the relationship between users, the edges of common vertices and supervertices are constructed based on the interactive relationship between users and resources, various second modality feature embedding matrices of resource data are correspondingly embedded in supervertices, the updated user feature embedding matrix is correspondingly embedded in common nodes, and the complementary information feature embedding matrix is embedded in corresponding supervertices. Through the heterogeneous composite graph of the embodiment of the present invention, embedded feature information in the graph can be extracted as needed, such as user information e u , resource information f,m , the user's resource information under the corresponding resource u,f and complementary information between different resourcesf1,f2 .
[0177] On the other hand, an embodiment of the present invention further provides a composite graph construction system for multimodal data, referring to Figure 2 ,include:
[0178] The first module is used to obtain a user data set, a resource data set, and a relationship between a user and a resource, wherein the resource data set includes a plurality of resource data, and the resource data includes a plurality of modal data;
[0179] The second module is used to determine multiple user feature embedding matrices based on the user data set;
[0180] The third module is used to process the multiple modal data respectively to obtain multiple first modal feature embedding matrices;
[0181] The fourth module is used to obtain multiple second modal feature embedding matrices by performing inter-modal information interaction based on multiple first modal feature embedding matrices of each resource data based on a channel dynamic exchange mechanism;
[0182] The fifth module is used to construct a composite graph based on the user feature embedding matrix, multiple second modality feature embedding matrices of resource data and the relationship between users and resources.
[0183] It can be understood that the contents of the above-mentioned embodiments of the method for constructing a composite graph of multimodal data are all applicable to the embodiments of the present system, the functions specifically implemented by the embodiments of the present system are the same as those in the above-mentioned embodiments of the method for constructing a composite graph of multimodal data, and the beneficial effects achieved are also the same as the beneficial effects achieved by the above-mentioned embodiments of the method for constructing a composite graph of multimodal data.
[0184] Reference Figure 3 , Figure 3 Schematic diagram of a composite graph construction device for multimodal data provided by an embodiment of the present invention. The composite graph construction device for multimodal data of the embodiment of the present invention includes one or more control processors and a memory. Figure 3 A control processor and a memory are taken as an example.
[0185] The control processor and the memory can be connected via a bus or other means. Figure 3 The example of connecting through bus is taken in the following.
[0186] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the control processor, and these remote memories may be connected to the composite graph construction device of the multimodal data via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0187] Those skilled in the art will understand that Figure 3 The device structure shown in the figure does not constitute a limitation on the composite graph construction device for multimodal data, and may include more or less components than those shown in the figure, or a combination of certain components, or a different arrangement of components.
[0188] The non-transient software program and instructions required to implement the composite graph construction method for multimodal data of the composite graph construction device applied to multimodal data in the above-mentioned embodiment are stored in the memory, and when executed by the control processor, the composite graph construction method for multimodal data of the composite graph construction device applied to multimodal data in the above-mentioned embodiment is executed.
[0189] In addition, an embodiment of the present invention also provides a computer-readable storage medium, which stores computer-executable instructions. The computer-executable instructions are executed by one or more control processors, so that the one or more control processors can execute the composite graph construction method of multimodal data in the above method embodiment.
[0190] It will be appreciated by those skilled in the art that all or some of the steps and systems in the disclosed method above may be implemented as software, firmware, hardware and appropriate combinations thereof. Some physical components or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor or a microprocessor, or may be implemented as hardware, or may be implemented as an integrated circuit, such as an application specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or a non-transitory medium) and a communication medium (or a temporary medium). As known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, disk storage or other magnetic storage devices, or any other medium that may be used to store desired information and may be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically embodies computer readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0191] The embodiments of the present invention are described in detail above with reference to the accompanying drawings, but the present invention is not limited to the above embodiments, and various changes can be made within the knowledge scope of ordinary technicians in the relevant technical field without departing from the purpose of the present invention.
Claims
1. A method for constructing a composite graph of multimodal data, characterized in that: The following steps are involved: Acquire a user data set, a resource data set, and a relationship between a user and a resource, wherein the resource data set includes a plurality of resource data, and the resource data includes a plurality of modal data; Determine a plurality of user feature embedding matrices according to the user data set; Processing the plurality of modal data respectively to obtain a plurality of first modal feature embedding matrices; Based on the channel dynamic exchange mechanism, multiple second modal feature embedding matrices are obtained by performing inter-modal information interaction according to multiple first modal feature embedding matrices of each resource data; Constructing a composite graph according to the user feature embedding matrix, multiple second modality feature embedding matrices of resource data, and the user-resource relationship; The method of performing inter-modal information interaction based on the multiple first modal feature embedding matrices of each resource data to obtain multiple second modal feature embedding matrices based on the channel dynamic exchange mechanism includes the following steps: Allocating values at different positions in each of the first modal features embedding matrix into different channels; The second modal feature embedding matrix is obtained by performing channel exchange on the weight values of the modal fusion results according to the numerical values in the channels in each of the first modal feature embedding matrices, wherein if there is a numerical value in the current first modal feature embedding matrix where the weight value is less than the fusion threshold, the channel corresponding to the numerical value is determined as the channel to be mapped, and the average value of the channels in other first modal feature embedding matrices with the same position as the channel to be mapped is substituted into the channel to be mapped; The step of constructing a composite graph according to the user feature embedding matrix, multiple second modality feature embedding matrices of resource data, and the relationship between the user and resources comprises the following steps: Determine the similarity between users according to the plurality of user feature embedding matrices, and obtain the relationship between users according to the similarity between users; updating the user feature embedding matrix according to the relationship between users, the relationship between user resources and a plurality of the second modality feature embedding matrices; Determine the similarity between resources according to the plurality of second modality feature embedding matrices, and obtain the relationship between resources according to the similarity between resources; Determine a complementary information feature embedding matrix between resources based on the relationship between the resources and multiple second modality feature embedding matrices of resource data; The composite graph is constructed according to the updated user feature embedding matrix and the complementary information feature embedding matrix.
2. The method for constructing a composite graph of multimodal data according to claim 1, characterized in that: Determining a plurality of user feature embedding matrices according to the user data set comprises the following steps: Granulating the user data set according to the information metric function to obtain an information granule set, wherein the information granule set includes a plurality of user information granules; The information particle set is integrated into a multi-head attention mechanism and a common common sense knowledge base mechanism for abstract feature representation to obtain the user feature embedding matrix.
3. The method for constructing a composite graph of multimodal data according to claim 2, characterized in that: The step of respectively processing the plurality of modal data to obtain a plurality of first modal feature embedding matrices comprises the following steps: For text modal data, a Text-Transformer encoder and a common knowledge base fusion mechanism are used to perform feature representation to obtain a first modal feature embedding matrix of the text modal data, wherein the text modal data is represented as a word embedding sequence, absolute position information is integrated into the word embedding sequence, and the word embedding is integrated into a first common knowledge base feature embedding matrix corresponding to the absolute position information according to the absolute position information, and the relative position information is integrated into the word embedding according to the relative distance between the absolute position information and the word embedding, to obtain the first modal feature embedding matrix of the text modal data; For image modality data, a CNN+Visual-Transformer encoder and a common knowledge base fusion mechanism are used for feature representation to obtain a first modality feature embedding matrix of the image modality data, wherein the image modality data is subjected to a convolution pooling operation to obtain multiple first image feature embedding matrices, and a position code and a second common knowledge knowledge base feature embedding matrix are integrated into each first image feature embedding matrix to obtain a second image feature embedding matrix, and the second image feature embedding matrix is input into a Visual-Transformer encoder based on a multi-head attention mechanism for encoding to obtain a first modality feature embedding matrix of the image modality data; For video modality data, a Video-Transformer encoder is used in combination with a common sense knowledge base fusion mechanism for feature representation to obtain a first modality feature embedding matrix of the video modality data, wherein the video modality data is input into a first layer encoder based on a standard multi-head self-attention mechanism for encoding to obtain a video feature embedding matrix, and the video feature embedding matrix is input into a second layer encoder based on a multi-head attention mechanism with a local sensing bias block for encoding to obtain a first modality feature embedding matrix of the video modality data, wherein the multi-head attention mechanism with a local sensing bias block is determined by the local sensing bias block position information and the second common sense knowledge base feature embedding matrix.
4. The method for constructing a composite graph of multimodal data according to claim 3, characterized in that: Determining the similarity between resources according to the plurality of second modal feature embedding matrices, and obtaining the relationship between resources according to the similarity between resources comprises the following steps: Transversely splicing the multiple second modal feature embedding matrices of each resource data to obtain a multimodal representation matrix; Extracting the maximum value of each row in the multimodal representation matrix to obtain a first vector, extracting the minimum value of each row in the multimodal representation matrix to obtain a second vector, and extracting the average of each row in the multimodal representation matrix to obtain a third vector; Concatenating the first vector, the second vector, and the third vector to obtain a multimodal feature embedding matrix of each resource data; The similarity between resources is determined according to the multimodal feature embedding matrix of different resource data, and the relationship between resources is obtained according to the similarity between resources.
5. The method for constructing a composite graph of multimodal data according to claim 4, characterized in that: The feature embedding of the composite graph is expressed as: Message=(and u ,And u,f ,And f,m ,And f1,f2 |u∈U,f,f1,f2∈F,m∈N + ); Among them, e u ∈R |U|×d Used to characterize user data, d is the d-dimensional feature of user data, e u,f ∈R |F|×d The feature embedding matrix used to characterize user u under resource f, e f,m ∈R |F|×m It is used to characterize each modal data in resource f, m is the m-dimensional feature of the modal data in resource f, and e f1,f2 ∈R |F|×|F| Used to represent the complementary information between resource f1 and resource f2.
6. A composite graph construction system for multimodal data, characterized in that: include: The first module is used to obtain a user data set, a resource data set, and a relationship between a user and a resource, wherein the resource data set includes a plurality of resource data, and the resource data includes a plurality of modal data; A second module is used to determine a plurality of user feature embedding matrices according to the user data set; A third module is used to process the multiple modal data respectively to obtain multiple first modal feature embedding matrices; A fourth module is used to obtain multiple second modal feature embedding matrices by performing inter-modal information interaction based on multiple first modal feature embedding matrices of each resource data based on a channel dynamic exchange mechanism; A fifth module is used to construct a composite graph according to the user feature embedding matrix, multiple second modality feature embedding matrices of resource data and the relationship between the user and resources; The fourth module is specifically used to perform the following steps: Allocating values at different positions in each of the first modal features embedding matrix into different channels; The second modal feature embedding matrix is obtained by performing channel exchange on the weight values of the modal fusion results according to the numerical values in the channels in each of the first modal feature embedding matrices, wherein if there is a numerical value in the current first modal feature embedding matrix where the weight value is less than the fusion threshold, the channel corresponding to the numerical value is determined as the channel to be mapped, and the average value of the channels in other first modal feature embedding matrices with the same position as the channel to be mapped is substituted into the channel to be mapped; The fifth module is specifically used to perform the following steps: Determine the similarity between users according to the plurality of user feature embedding matrices, and obtain the relationship between users according to the similarity between users; updating the user feature embedding matrix according to the relationship between users, the relationship between user resources and a plurality of the second modality feature embedding matrices; Determine the similarity between resources according to the plurality of second modality feature embedding matrices, and obtain the relationship between resources according to the similarity between resources; Determine a complementary information feature embedding matrix between resources based on the relationship between the resources and multiple second modality feature embedding matrices of resource data; The composite graph is constructed according to the updated user feature embedding matrix and the complementary information feature embedding matrix.
7. A composite graph construction device for multimodal data, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the composite graph construction method for multimodal data as described in any one of claims 1 to 5.
8. A computer-readable storage medium storing a program executable by a processor, characterized in that: The program executable by the processor is used to implement the method for constructing a composite graph of multimodal data as described in any one of claims 1 to 5 when executed by the processor.
Citation Information
Patent Citations
Multi-modal knowledge graph construction method
CN112200317A
Multi-modal emotion recognition method and device, electronic equipment and storage medium
CN112418034A