Lightweight multi-modal knowledge graph representation learning method based on fourier transform
By employing Fourier transform and filtering gate techniques, the problems of insufficient efficiency and quality in multimodal knowledge graph representation learning are solved, achieving efficient social multimodal knowledge graph representation and reducing computational resource consumption.
Patent Information
- Application Number
- CN202310786378.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-29
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2043-06-29
AI Technical Summary
Existing multimodal knowledge graph representation learning methods are insufficient in terms of efficiency and quality in social scenarios, especially in terms of excessive computational resource consumption, making it difficult to balance high-quality representation and model efficiency.
A lightweight method based on Fourier transform is adopted, which eliminates visual modal noise through filtering gates, introduces a deconvolution strategy to process visual modal input, and uses the Fourier operator AFNO for modal fusion to reduce computational resource consumption.
While ensuring the quality of multimodal knowledge graph representation, it significantly reduces computational resource consumption, improves efficiency, effectively filters visual noise, and reduces the waste of time and space resources.
Smart Images

Figure CN117171351B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of knowledge graph technology, and in particular relates to a lightweight multimodal knowledge graph representation learning method based on Fourier transform. Background Technology
[0002] A knowledge graph is a large-scale semantic network based on semantic relationships, where nodes represent entities and concepts, and edges represent semantic relationships between concepts. Knowledge graphs are typically represented in triplet form (e...). h ,r,e t ), where e h and e t Let 'r' represent the head and tail entities, and 'r' represent the semantic relation. Traditional knowledge graphs often use pure symbolic text to represent concepts and entities, but this representation method has obvious limitations and cannot fully describe the characteristics of entities and the complexities of the real world. Therefore, multimodal knowledge graphs that include image and text information have emerged. With the widespread use of multimodal information in social media platforms like Weibo and Twitter, multimodal knowledge graphs have been widely applied in various fields such as information retrieval and recommendation systems, and representation learning methods for multimodal knowledge graphs are constantly emerging. Multimodal knowledge graphs consist of text information and image information corresponding to entities, where the text information also includes structural information and descriptions of the entities. The data for multimodal knowledge graphs often comes from user posts on social media platforms that contain rich information. The text and image information contained therein has a strong timeliness characteristic, which further enhances the representation capabilities of knowledge graphs. Information from the visual modality can sometimes enhance entity representations, but noise it contains can also negatively impact representation learning, thus affecting the performance of multimodal knowledge graphs in downstream applications (such as multimodal knowledge graph completion and information extraction). Therefore, obtaining high-quality multimodal knowledge graph representations is crucial.
[0003] Previous work, in an effort to enhance information extraction and fusion analysis capabilities and better handle information from various modalities, often introduced complex model structures, such as the transformer-based visual processing model ViT and the language processing model BERT. These complex models frequently employ a segmentation and stacking approach to process input. While this can enhance representation learning to some extent, it also significantly increases the overall number of model parameters, leading to a substantial increase in computational complexity and space consumption. This significant expenditure of computational resources not only reduces the efficiency of representation learning tasks but also hinders the practical application of multimodal knowledge graph representation learning models. Considering the real-time updating and complexity of multimodal information in social scenarios, time and space efficiency are more crucial aspects for the practical application of multimodal knowledge graph representation learning models. Therefore, how to ensure high-quality representation capabilities while maintaining model efficiency is a pressing issue that needs to be addressed. Summary of the Invention
[0004] To address the aforementioned shortcomings in existing technologies, this invention provides a lightweight multimodal knowledge graph representation learning method based on Fourier transform. This method solves the problem that previous multimodal knowledge graph representation learning methods still fall short in terms of efficiency and quality, considering the real-time updates and complexity of multimodal information in social scenarios.
[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0006] This solution provides a lightweight multimodal knowledge graph representation learning method based on Fourier transform, including the following steps:
[0007] S1. Use a filtering gate to process the original image from the multimodal knowledge graph input and form an embedded representation;
[0008] S2. Based on the formed embedding representation, the Fourier operator AFNO is used to fuse the output embedding representation, thus completing the lightweight multimodal knowledge graph representation learning.
[0009] The beneficial effects of this invention are: it can significantly reduce the consumption of computational resources while ensuring the representation learning effect of multimodal knowledge graphs. This invention first processes and embeds the original input from the multimodal knowledge graph, and then aligns and fuses the information from each modality to retain the information from each modality to the greatest extent, obtaining a high-quality social knowledge graph representation. In the modality processing part, this invention introduces a filter gate to eliminate the noise influence from the visual modality and reduces the number of images processed by the model, thereby improving efficiency. Furthermore, this invention introduces a deconvolution-free strategy to process the images input from the visual modality, reducing the computational resources consumed by visual modality processing. In the modality fusion part, this invention uses the Fourier operator AFNO to fuse the information from each modality, ensuring the information extraction effect while reducing the waste of computational resources by the model, thus achieving efficient social multimodal knowledge graph representation learning.
[0010] Further, step S1 includes the following steps:
[0011] S101. Calculate the similarity of the original images from the multimodal knowledge graph using a filtering gate, and select the image with the highest similarity as the representative of the image set.
[0012] S102. Using a deconvolutional visual processing method, linear flattening is applied to the images with the highest similarity to obtain a visual modality embedding representation.
[0013] S103, Transfer entity structure information and entity description information Concatenated into a text sequence representation And splice together visual modal embedding representation and text sequence representation Forming an embedded representation
[0014] The beneficial effects of the above-mentioned further solutions are as follows: This invention can remove noise from the visual modality while retaining effective information from each modality, thus reducing the negative impact of visual noise on representation learning. This invention eliminates the influence of noise from the visual modality by introducing a filter gate, and reduces the number of images processed by the model, thereby improving efficiency. It effectively filters noise information from the dataset level and effectively reduces the number of images that the model needs to process to 1 / 10 of the original. Furthermore, this invention introduces a deconvolution-free strategy to process the images input from the visual modality, reducing the computational resources consumed by visual modality processing. It can eliminate the previous convolutional methods while ensuring the embedding and representation effect of modality information, saving time and space resources.
[0015] Furthermore, step S101 includes the following steps:
[0016] S1011. Scale the original image from the multimodal knowledge graph to a fixed size;
[0017] S1012. Convert the scaled image to grayscale.
[0018] S1013. Use Fourier transform to convert the grayscale processed image from the pixel domain to the frequency domain, generate a discrete cosine transform (DCT) matrix, and retain the 8×8 low-frequency matrix in the upper left corner.
[0019] S1014. Calculate the DCT mean of the current Discrete Cosine Transform (DCT) matrix;
[0020] S1015. Compare the grayscale of each image pixel with the DCT mean to obtain the hash matrix of the current image, and combine them into a 64-bit binary integer to form the current image fingerprint.
[0021] S1016. Compare the fingerprints of different images, compare each bit, calculate the Hamming distance between the images, and obtain the similarity between the images.
[0022] S1017. Select the image with the highest similarity as the representative of the image set.
[0023] The beneficial effects of the above-mentioned further solutions are: the modal fusion structure based on Fourier operators in this invention can effectively extract information from various modalities, while reducing the computational load of the model.
[0024] Furthermore, the expression for the image with the highest similarity is as follows:
[0025]
[0026] in, Let represent the image with the highest similarity, max denotes the maximization operation, m represents the total number of images, j represents the index of the j-th image in the image set, and pHash(·) denotes the perceptual hash similarity operation. Represents the head entity e i The corresponding k-th image, Represents the head entity e i The corresponding j-th image.
[0027] The beneficial effect of the above-mentioned further solution is that the present invention can quickly and easily identify the image in the image set that is most relevant to the entity by using the image with the highest similarity.
[0028] Furthermore, the embedded representation The expression is as follows:
[0029]
[0030]
[0031] in, This indicates an embedded representation. Represents the head entity e i The corresponding position representation of the visual modality embedding representation, Represents the head entity e i The corresponding visual modality embedding representation, Represents the head entity e i The corresponding position representation of the text modality embedding. Represents the head entity e i The corresponding text modality embedding representation is as follows: [CLS] represents the beginning of the text sequence, and [SEP] represents the segmentation symbol. Represents the head entity e i Corresponding relational entities The text contained herein, [MASK] represents the tail entity to be predicted, i represents the index of the i-th entity in the entity set, and n represents the number of entities in the entity set.
[0032] The beneficial effect of the above-mentioned further scheme is that the obtained embedding representation can retain information from each modality to the greatest extent and transform it into a commonly used representation in multimodal pre-training tasks.
[0033] Furthermore, step S2 includes the following steps:
[0034] S201. Divide the formed embedding representation into h×w blocks to obtain the input tensor, where each block is represented as a d-dimensional word segmentation sequence, h represents the height, and w represents the width;
[0035] S202. Spatial mixing of the input tensor is performed using the Fourier operator AFNO to obtain the output embedding representation, thus completing the lightweight multimodal knowledge graph representation learning.
[0036] The beneficial effects of the above-mentioned further solutions are: the present invention utilizes the Fourier operator AFNO to fuse information from various modalities, thereby reducing the waste of computational resources by the model while ensuring the information extraction effect, thus achieving efficient social multimodal knowledge graph representation learning.
[0037] Furthermore, step S202 includes the following steps:
[0038] S2021. Perform a Fourier transform on the input tensor to obtain the intermediate representation z:
[0039] z = FFT(X)
[0040] Where FFT(·) represents Fourier transform, and X represents the input tensor;
[0041] S022. Based on the intermediate representation z, use a two-layer MLP structure to share weights across all inputs.
[0042]
[0043] Where MLP(·) represents the MLP structure;
[0044] S2023, Based on input shared weights The intermediate output X' is obtained using the inverse Fourier transform:
[0045]
[0046] Wherein, IFFT(·) represents the inverse Fourier transform;
[0047] S2024. Normalize the intermediate output X' using the double normalization algorithm to obtain the output embedding representation. Complete lightweight multimodal knowledge graph representation learning.
[0048] The beneficial effect of the above-mentioned further solution is that, while achieving the same effect as the attention module in the transformer, it reduces the amount of computation required by the original attention module and improves the computation speed.
[0049] Furthermore, the output embedding representation The expression is as follows:
[0050]
[0051] LayerNorm(·) represents the normalization operation. Attached Figure Description
[0052] Figure 1 This is a flowchart of the method of the present invention.
[0053] Figure 2 This is a flowchart illustrating the method of applying the present invention to a social scenario.
[0054] Figure 3 This diagram illustrates the performance and computational resource consumption of the present invention in the multimodal knowledge graph completion task in this embodiment.
[0055] Figure 4 This diagram illustrates the performance and time complexity comparison of each model on the WN18-IMG dataset in this embodiment.
[0056] Figure 5 This is a comparison chart of the performance and space complexity of each model on the WN18-IMG dataset in this embodiment. Detailed Implementation
[0057] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.
[0058] Example 1
[0059] like Figure 1 As shown, this invention provides a lightweight multimodal knowledge graph representation learning method based on Fourier transform: MLFormer, the implementation method of which is as follows:
[0060] S1. The original image from the multimodal knowledge graph input is processed and embedded using a filtering gate. The implementation method is as follows:
[0061] S101. A similarity calculation is performed on the original images from the multimodal knowledge graph using a filtering gate, and the image with the highest similarity is selected as the representative of the image set. The implementation method is as follows:
[0062] S1011. Scale the original image from the multimodal knowledge graph to a fixed size;
[0063] S1012. Convert the scaled image to grayscale.
[0064] S1013. Use Fourier transform to convert the grayscale processed image from the pixel domain to the frequency domain, generate a discrete cosine transform (DCT) matrix, and retain the 8×8 low-frequency matrix in the upper left corner.
[0065] S1014. Calculate the DCT mean of the current Discrete Cosine Transform (DCT) matrix;
[0066] S1015. Compare the grayscale of each image pixel with the DCT mean to obtain the hash matrix of the current image, and combine them into a 64-bit binary integer to form the current image fingerprint.
[0067] S1016. Compare the fingerprints of different images, compare each bit, calculate the Hamming distance between the images, and obtain the similarity between the images.
[0068] S1017. Select the image with the highest similarity as the representative of the image set:
[0069]
[0070] in, Let represent the image with the highest similarity, max denotes the maximization operation, m represents the total number of images, j represents the index of the j-th image in the image set, and pHash(·) denotes the perceptual hash similarity operation. Represents the head entity e i The corresponding k-th image, Represents the head entity e i The corresponding j-th image;
[0071] S102. Using a deconvolutional visual processing method, linear flattening is applied to the images with the highest similarity to obtain a visual modality embedding representation.
[0072] S103, Transfer entity structure information and entity description information Concatenated into a text sequence representation And splice together visual modal embedding representation and text sequence representation Forming an embedded representation
[0073]
[0074]
[0075] in, This indicates an embedded representation. Represents the head entity e i The corresponding position representation of the visual modality embedding representation, Represents the head entity e i The corresponding visual modality embedding representation, Represents the head entity e i The corresponding position representation of the text modality embedding. Represents the head entity e i The corresponding text modality embedding representation is as follows: [CLS] represents the beginning of the text sequence, and [SEP] represents the segmentation symbol. Represents the head entity e i Corresponding relational entities The text contained herein, [MASK] represents the tail entity to be predicted, i represents the index of the i-th entity in the entity set, and n represents the number of entities in the entity set;
[0076] S2. Based on the formed embedding representation, the Fourier operator AFNO is used to fuse them to obtain the output embedding representation, thus completing the lightweight multimodal knowledge graph representation learning. The implementation method is as follows:
[0077] S201. Divide the formed embedding representation into h×w blocks to obtain the input tensor, where each block is represented as a d-dimensional word segmentation sequence, h represents the height, and w represents the width;
[0078] S202. Spatial mixing of the input tensor is performed using the Fourier operator AFNO to obtain the output embedding representation, thus completing the lightweight multimodal knowledge graph representation learning. The implementation method is as follows:
[0079] S2021. Perform a Fourier transform on the input tensor to obtain the intermediate representation z:
[0080] z = FFT(X)
[0081] Where FFT(·) represents Fourier transform, and X represents the input tensor;
[0082] S022. Based on the intermediate representation z, use a two-layer MLP structure to share weights across all inputs.
[0083]
[0084] Where MLP(·) represents the MLP structure;
[0085] S2023, Based on input shared weights The intermediate output X' is obtained using the inverse Fourier transform:
[0086]
[0087] Wherein, IFFT(·) represents the inverse Fourier transform;
[0088] S2024. Normalize the intermediate output X' using the double normalization algorithm to obtain the output embedding representation. The lightweight multimodal knowledge graph representation learning is completed, and the output embedding representation is... The expression is as follows:
[0089]
[0090] LayerNorm(·) represents the normalization operation.
[0091] This invention first processes and embeds the raw input from a multimodal knowledge graph, then aligns and fuses the information from each modality to retain the information from each modality to the greatest extent possible, resulting in a high-quality social knowledge graph representation. In the modality processing part, this invention introduces a filter gate to eliminate the noise influence from the visual modality and reduces the number of images processed by the model, thereby improving efficiency. Furthermore, this invention introduces a deconvolution-free strategy to process the images input from the visual modality, reducing the computational resources consumed by visual modality processing. In the modality fusion part, this invention utilizes the Fourier operator AFNO to fuse the information from each modality, ensuring effective information extraction while reducing the waste of computational resources by the model, thus achieving efficient learning of social multimodal knowledge graph representations.
[0092] Example 2
[0093] The present invention will be further described below.
[0094] As explained in the background section, considering the real-time updating and complexity of multimodal information in social scenarios, time and space efficiency are more important aspects for multimodal knowledge graph representation learning models in practical applications. Therefore, how to ensure high-quality representation capabilities while maintaining model efficiency is an urgent problem to be solved. Figure 2 As shown, this invention provides a lightweight multimodal knowledge graph representation learning method based on Fourier transform (MLFormer, i.e., the method proposed in this invention), the implementation method of which is as follows:
[0095] S1. The original image input from the social multimodal knowledge graph is processed and embedded using a filtering gate. The implementation method is as follows:
[0096] S101. A similarity calculation is performed on the original images from the social multimodal knowledge graph using a filtering gate, and the image with the highest similarity is selected as the representative of the image set. The implementation method is as follows:
[0097] S1011, Scale the original image from the social multimodal knowledge graph to a fixed size;
[0098] S1012. Convert the scaled image to grayscale.
[0099] S1013. Use Fourier transform to convert the grayscale processed image from the pixel domain to the frequency domain, generate a discrete cosine transform (DCT) matrix, and retain the 8×8 low-frequency matrix in the upper left corner.
[0100] S1014. Calculate the DCT mean of the current Discrete Cosine Transform (DCT) matrix;
[0101] S1015. Compare the grayscale of each image pixel with the DCT mean to obtain the hash matrix of the current image, and combine them into a 64-bit binary integer to form the current image fingerprint.
[0102] S1016. Compare the fingerprints of different images, compare each bit, calculate the Hamming distance between the images, and obtain the similarity between the images.
[0103] S1017. Select the image with the highest similarity as the representative of the image set;
[0104] S102. Using a deconvolutional visual processing method, linear flattening is applied to the images with the highest similarity to obtain a visual modality embedding representation.
[0105] S103, Social entity structure information and social entity description information Concatenated into a text sequence representation And splice together visual modal embedding representation and text sequence representation Forming an embedded representation
[0106] S2. Based on the formed embedding representation, the Fourier operator AFNO is used to fuse them to obtain the output embedding representation, thus completing the lightweight multimodal knowledge graph representation learning. The implementation method is as follows:
[0107] S201. Divide the formed embedding representation into h×w blocks to obtain the input tensor, where each block is represented as a d-dimensional word segmentation sequence, h represents the height, and w represents the width;
[0108] S202. Spatial mixing of the input tensor is performed using the Fourier operator AFNO to obtain the output embedding representation, thus completing the lightweight multimodal knowledge graph representation learning. The implementation method is as follows:
[0109] S2021. Perform a Fourier transform on the input tensor to obtain the intermediate representation z:
[0110] z = FFT(X)
[0111] Where FFT(·) represents Fourier transform, and X represents the input tensor;
[0112] S022. Based on the intermediate representation z, use a two-layer MLP structure to share weights across all inputs.
[0113]
[0114] Where MLP(·) represents the MLP structure;
[0115] S2023, Based on input shared weights The intermediate output X' is obtained using the inverse Fourier transform:
[0116]
[0117] Wherein, IFFT(·) represents the inverse Fourier transform;
[0118] S2024. Normalize the intermediate output X' using the double normalization algorithm to obtain the output embedding representation. Complete lightweight multimodal knowledge graph representation learning.
[0119] In this embodiment, the lightweight social multimodal knowledge graph representation learning method based on Fourier transform is divided into two parts: modality processing and modality fusion. The modality processing part is further divided into two steps: raw dataset processing and modality feature processing.
[0120] Raw Dataset Processing: In processing the raw dataset based on social knowledge graphs, considering that information from image modalities may introduce noise that interferes with the representation of the knowledge graph, a filtering gate is used to process the raw input images. Given the i-th social entity e in the social entity set... i Its corresponding social image set is Where m represents the number of social media photos, Let m be the m-th social image. The filtering gate calculates the similarity of images in the social image set and then selects the image with the highest similarity. As a representative of this social photo collection, the calculation formula is as follows:
[0121]
[0122] in, Let represent the image with the highest similarity, max denotes the maximization operation, m represents the total number of images, j represents the j-th image in the social image set, and pHash(·) denotes the perceptual hash similarity operation. Represents the social head entity e i The corresponding k-th image, Represents the social head entity e i The corresponding j-th image.
[0123] The basic process of the similarity algorithm is as follows:
[0124] Resize the input image to a suitable size (usually 32×32);
[0125] The image is converted to grayscale to simplify the image matrix and improve processing speed;
[0126] The image is transformed from the pixel domain to the frequency domain using Fourier transform, generating a Discrete Cosine Transform (DCT) matrix while retaining the 8×8 low-frequency matrix in the upper left corner. The formula for the two-dimensional DCT transform is as follows:
[0127] F = AfA T
[0128]
[0129] Where F represents the generated transformation matrix, A represents the transformation coefficient matrix, f represents the original information input transformed to the frequency domain, and A T Let A(i',j') represent the transpose of A, A(i',j') represent the coefficient of the input at the (i',j')th two-dimensional pixel, N represent the number of points in the original signal, i' and j' represent the two-dimensional pixel input, with values ranging from [0,N-1], and c(i') represent the compensation coefficient, expressed as:
[0130]
[0131] Calculate the DCT mean of the current DCT coefficient matrix;
[0132] The grayscale value of each pixel is compared with the DCT mean to obtain the hash matrix of the current image. These hash matrices are then combined into a 64-bit binary integer to form the fingerprint of the current image.
[0133]
[0134] Where hash(i”,j”) represents the fingerprint of the current image, and F(i”,j”) represents the hash matrix. denoted as DCT mean, i”,j” represent the location of the pixel, and N' represents the total number of pixels.
[0135] By comparing fingerprints from different images, comparing each digit, and calculating the Hamming distance between the images, the similarity between social images can be obtained.
[0136] In this embodiment, the filtering gate operation not only eliminates dataset-level noise but also reduces the number of social images the model needs to process, saving computational resources. The training, testing, and validation sets of the multimodal knowledge graph consist of a large number of social triples, which are derived from one-to-one, one-to-many, and many-to-many relationship pairs in real-world social knowledge graph datasets. The social relationship and social entity dictionaries in the social knowledge graph dataset are a one-to-one correspondence between text descriptions and structural information. Considering the noise in the text modality, for text modality input, this invention removes social triples that cannot be found in the dictionary during the original dataset processing stage.
[0137] In this embodiment, modal feature processing includes visual modal feature processing and text modal feature processing.
[0138] In this embodiment, for visual modality feature processing, after filtering, the number of social images corresponding to each social entity is reduced to one. For missing images of certain social entities, this invention uses a method of randomly generating a blank image to fill in the gaps. In the visual feature processing stage, this invention abandons the previous convolutional processing method and uses a deconvolutional visual processing method to perform linear flattening mapping on the original image to obtain the visual modality embedding representation.
[0139] In this embodiment, for text modality feature processing, it is considered that the text modality information of the knowledge graph comes from two aspects, namely social entity structure information. and social entity description information Both parts of the information are preserved and concatenated into a text sequence representation.
[0140]
[0141] in, Represents the social head entity e i Corresponding social relationship entities The included social text, where i represents the index of the i-th social entity in the entity set, and n represents the number of entities in the social entity set. Indicates when e i Relationship information corresponding to the head social entity.
[0142] In this embodiment, during training, the text feature processing module mimics the text masking task in the pre-training task to further process the text sequence representation as follows:
[0143]
[0144] The embedded representation of visual and textual modalities will be combined with modality type information. and location information The concatenation forms the final embedded representation:
[0145]
[0146] in, This indicates an embedded representation. Represents the social head entity e i The corresponding position representation of the visual modality embedding representation, Represents the social head entity e i The corresponding visual modality embedding representation, Represents the social head entity e i The corresponding position representation of the text modality embedding. Represents the social head entity e i The corresponding text modality embedding representation is as follows: [CLS] represents the beginning of the text sequence, and [SEP] represents the segmentation symbol. Represents the head entity e i Corresponding social relationship entities The text contained herein, [MASK] represents the tail entity to be predicted, i represents the index of the i-th social entity in the entity set, and n represents the number of social entities in the entity set.
[0147] In this embodiment, considering the superiority of the transformer model in multimodal information processing, the modality fusion part uses a transformer encoder to fuse the embedded representations from each modality. Since the computational complexity of the self-attention mechanism in the transformer is quadratically related to the length of the input sequence, this method replaces the self-attention mechanism with the Fourier operator AFNO, whose computational complexity is logarithmically related to the input, to improve the model's inference speed.
[0148] It is known that the embedding representation is divided into h×w patches before being input to the AFNO layer, where h represents the height and w represents the width, and each patch can be represented as a d-dimensional token sequence. Then the input to the AFNO layer of the Fourier operator is a block tensor X∈R of size h×w×d. h×w×d , where R h×w×d Representing a block tensor, the Fourier operator AFNO layer performs spatial mixing on the input tensor according to the channel mixing method. The spatial mixing steps are as follows:
[0149] Performing a Fourier transform on the input tensor yields the intermediate representation z:
[0150] z = FFT(X)
[0151] Where FFT(·) represents Fourier transform, and X represents the input tensor;
[0152] A two-level MLP structure is used to share weights across all inputs.
[0153]
[0154] Where MLP(·) represents the MLP structure;
[0155] The intermediate output X' is obtained using the inverse Fourier transform:
[0156]
[0157] Wherein, IFFT(·) represents the inverse Fourier transform;
[0158] The intermediate outputs are normalized using a double normalization algorithm to obtain the final output embedding.
[0159]
[0160] LayerNorm(·) represents the normalization operation.
[0161] In this embodiment, in the social multimodal knowledge graph completion task, given a query q = (eh, r, unknown), the model assigns different probability scores to each social tail entity from the given candidate social tail entities and sorts them from high to low scores. The social tail entity with the highest ranking is selected as the final predicted social entity and compared with the real social entity. This idea is consistent with the text masking task MLM in pre-training tasks, both belonging to multi-label classification tasks. Therefore, during training, we chose the cross-entropy loss function to train the model, as shown in the following formula:
[0162] loss = -log(p) i )
[0163] Where, p i This represents the probability score that the i-th candidate social entity is the correctly predicted entity.
[0164] In the scenario of social knowledge graph completion, given a social knowledge graph query, the model will generate a candidate set of social entities based on a scoring function. Each social entity in the query is scored, and the most likely predicted answer is selected based on the ranking.
[0165] In this embodiment, the advantage of using the above algorithm is that the representation learning method provided by this invention can not only effectively extract various types of information from multimodal knowledge graphs and reduce the impact of visual modal noise from social knowledge graphs, but also improve the inference speed of the representation learning model in social scenarios and save the model's space usage. The representation performance of the model was verified on two real-world datasets, FB15k-237 and WN18-IMG.
[0166] FB15k-237 is a subset of the commonly used dataset FB15k for knowledge graph completion tasks. It contains 237 classes of relations and 14,451 entities. Among them, there are 272,115 training triples, 17,535 validation triples, and 20,466 test triples.
[0167] WN18-IMG: WN18 is a knowledge graph dataset from WordNet, and WN18-IMG is the multimodal dataset corresponding to the WN18 knowledge graph data. Each entity corresponds to ten images, containing 18 classes of relations and 40,943 entities. It includes 141,442 training triples, 5,000 validation triples, and 5,000 test triples.
[0168] This invention verifies the performance and computational resource consumption of the proposed MLFormer method on multimodal knowledge graph completion tasks, see [link to documentation]. Figure 3 Here, Hits@n (n = 1, 3, 10) refers to the average proportion of triples with a ranking less than n in link prediction, and MR is the average ranking of the correct entity scoring function. FLOPs and Parameters are commonly used metrics for measuring the time and space complexity of a model. This invention compares the performance and computational resource consumption of MLFormer with current multimodal knowledge graph representation learning methods and pre-trained multimodal models on multimodal knowledge graph completion tasks. Figure 3 The bolded part shows the effect of the present invention. It can be seen that the performance of the present invention in the completion task is basically on par with the best multimodal knowledge graph representation model, and is far superior to the existing methods in terms of time and space complexity.
[0169] In this embodiment, regarding the comparison of time complexity: Figure 4 This section compares the performance and time complexity of various models on the WN18-IMG dataset. Figure 4As can be seen, MLFormer reduces time complexity by 99.4% compared to MKGformer, the current best multimodal knowledge graph representation learning method, 95.5% compared to RSME, and 98.5% compared to pre-trained multimodal models. It is known that the size of FLOPs is proportional to the size of the model input. The input size of MKGformer is (224, 64), where 224 represents the image size and 64 represents the sequence length; while the input size of RSME is (224, 100), and the input size of MLFormer is (384, 32). It can be seen that MLFormer still has the lowest time complexity even when the image input size is larger than other models, which proves that this invention effectively improves the inference speed of the model in social media scenarios.
[0170] In this embodiment, regarding the comparison of space complexity: Figure 5 This is a comparison chart of the performance and space complexity of each model on the WN18-IMG dataset. Figure 5 As can be seen, MLFormer reduces the number of parameters by 48.4% compared to MKGformer, 45.9% compared to TransAE, and 16.2% compared to RSME. MLFormer achieves the same performance as the best model while significantly reducing the number of parameters, which proves that this invention is very effective in reducing the complexity of the reasoning space in social knowledge graphs.
Claims
1. A lightweight multimodal knowledge graph representation learning method based on Fourier transform, characterized in that, Includes the following steps: S1. Use a filtering gate to process the original image from the multimodal knowledge graph input and form an embedded representation; Step S1 includes the following steps: S101. Calculate the similarity of the original images from the multimodal knowledge graph using a filtering gate, and select the image with the highest similarity as the representative of the image set. S102. Using a deconvolutional visual processing method, linear flattening is applied to the images with the highest similarity to obtain a visual modality embedding representation. ; S103, Transfer entity structure information and entity description information Concatenated into a text sequence representation And splice together visual modal embedding representation and text sequence representation To form an embedded representation ; S2. Based on the formed embedding representation, the Fourier operator AFNO is used to fuse them to obtain the output embedding representation, thus completing the lightweight multimodal knowledge graph representation learning. Step S2 includes the following steps: S201, Segment the formed embedded representation into We obtain the input tensor by dividing the input into blocks, where each block is represented as... d Dimensional word segmentation sequence, h Indicates altitude, w Indicates width; S202. Spatial mixing of the input tensor is performed using the Fourier operator AFNO to obtain the output embedding representation, thus completing the lightweight multimodal knowledge graph representation learning. Step S202 includes the following steps: S2021. Perform a Fourier transform on the input tensor to obtain an intermediate representation. : in, Indicates Fourier transform, Indicates the input tensor; S2022, According to the intermediate representation A two-layer MLP structure is used to share weights across all inputs. : in, Indicates the MLP structure; S2023, Based on input shared weights The intermediate output is obtained using the inverse Fourier transform. : in, Indicates the inverse Fourier transform; S2024. Use the dual normalization algorithm to process intermediate outputs. After normalization, the output embedding representation is obtained. Complete lightweight multimodal knowledge graph representation learning.
2. The lightweight multimodal knowledge graph representation learning method based on Fourier transform according to claim 1, characterized in that, Step S101 includes the following steps: S1011. Scale the original image from the multimodal knowledge graph to a fixed size; S1012. Convert the scaled image to grayscale. S1013. Use Fourier transform to convert the grayscale processed image from the pixel domain to the frequency domain, generate a Discrete Cosine Transform (DCT) matrix, and retain the upper left corner. Low-frequency matrix; S1014. Calculate the DCT mean of the current Discrete Cosine Transform (DCT) matrix; S1015. Compare the grayscale of each image pixel with the DCT mean to obtain the hash matrix of the current image, and combine them into a 64-bit binary integer to form the current image fingerprint. S1016. Compare the fingerprints of different images, compare each bit, calculate the Hamming distance between the images, and obtain the similarity between the images. S1017. Select the image with the highest similarity as the representative of the image set.
3. The lightweight multimodal knowledge graph representation learning method based on Fourier transform according to claim 2, characterized in that, The expression for the image with the highest similarity is as follows: in, This indicates the image with the highest similarity. This indicates a maximization operation. Indicates the total number of images. j Indicates the first image in the image set. j The index number of the image. This represents the perceptual hash similarity operation. Represents head entity The corresponding number k One picture, Represents head entity The corresponding number j [Number of pictures] 4. The lightweight multimodal knowledge graph representation learning method based on Fourier transform according to claim 1, characterized in that, The embedding representation The expression is as follows: in, This indicates an embedded representation. Represents head entity The corresponding position representation of the visual modality embedding representation, Represents head entity The corresponding visual modality embedding representation, Represents head entity The corresponding position representation of the text modality embedding. Represents head entity The corresponding text modality embedding representation, Indicates the beginning of a text sequence. Indicates a segmentation symbol. Represents head entity Corresponding relational entities The text contained therein This indicates the tail entity that needs to be predicted. i Represents the first entity in the entity set i The index of an entity, n This indicates the number of entities in the entity set.
5. The lightweight multimodal knowledge graph representation learning method based on Fourier transform according to claim 1, characterized in that, The output embedding representation The expression is as follows: in, This indicates a normalization operation.