Multimodal Data Processing Method, Device, Storage Medium and Electronic Device

By pre-training fusion vocabulary and VQGAN model, image feature representation is optimized, and image embedding vectors are combined with text embedding vectors, which solves the problems of multimodal large language model's computational resource occupation and low information fusion efficiency, and realizes more efficient multimodal information processing.

CN119226992BActive Publication Date: 2025-07-18PEKING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411058820.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-02
Publication Date
2025-07-18
Estimated Expiration
2044-08-02

AI Technical Summary

Technical Problem

When processing multimodal information such as images and text, the multimodal large language model occupies a lot of computing resources, making it difficult to effectively understand and fuse multimodal information, and its overall performance is poor.

Method used

The image embedding vector is converted into a pre-fusion coded vector using a pre-trained fusion vocabulary, and combined with the text embedding vector to form a target multimodal vector, and the image feature representation and text feature fusion are optimized through the VQGAN model and the Transformer model.

Benefits of technology

Reduce the number of encodings of image representations, save computing resources, and improve the understanding and fusion performance of multimodal large language models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119226992B_ABST
    Figure CN119226992B_ABST
Patent Text Reader

Abstract

The present invention discloses a multimodal data processing method, device, storage medium and electronic device. Among them, the method includes: obtaining multimodal data to be recognized, where the multimodal data includes image data and text data; obtaining an image embedding vector corresponding to the image data, and converting the image embedding vector into a pre-fused coding vector based on a pre-trained fusion vocabulary; the pre-trained fusion vocabulary is a coding book obtained according to image training samples for reducing the coding amount of image features; combining the pre-fused coding vector and a text embedding vector corresponding to the text data to obtain a target multimodal vector. The present invention solves the technical problem in the related art that multimodal large language models occupy a large amount of computing resources, are difficult to effectively understand and fuse multimodal information, and have poor overall performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a multi-modal data processing method, device, storage medium, and electronic device. Background Art

[0002] When processing multi-modal information such as images and texts, multi-modal large language models in related technologies usually process images into a large number of independent encodings. This method not only occupies a large amount of computing resources but also makes it difficult to capture the structured information and semantic associations in images, making it difficult for multi-modal large language models to effectively understand and fuse multi-modal information. In addition, large language models usually have a fixed input length limit, which will cause information loss when processing high-dimensional information such as images and also limit the number of multi-modal information that can be processed simultaneously. Therefore, the overall performance of multi-modal large language models is not good. Summary of the Invention

[0003] Embodiments of the present application provide a multi-modal data processing method, device, storage medium, and electronic device to at least solve the technical problems in related technologies that multi-modal large language models occupy more computing resources, are difficult to effectively understand and fuse multi-modal information, and have poor overall performance.

[0004] According to one aspect of the embodiments of the present application, a multi-modal data processing method is provided, including: obtaining multi-modal data to be recognized, where the multi-modal data includes image data and text data; obtaining an image embedding vector corresponding to the image data, and converting the image embedding vector into a pre-fused coding vector based on a pre-trained fusion vocabulary; the pre-trained fusion vocabulary is a coding book obtained according to image training samples for reducing the coding amount of image features; combining the pre-fused coding vector and a text embedding vector corresponding to the text data to obtain a target multi-modal vector.

[0005] According to another aspect of the embodiments of the present application, a multi-modal data processing device is further provided, including: a first obtaining unit for obtaining multi-modal data to be recognized, where the multi-modal data includes image data and text data; a second obtaining unit for obtaining an image embedding vector corresponding to the image data, and converting the image embedding vector into a pre-fused coding vector based on a pre-trained fusion vocabulary; the pre-trained fusion vocabulary is a coding book obtained according to image training samples for reducing the coding amount of image features; a combining unit for combining the pre-fused coding vector and a text embedding vector corresponding to the text data to obtain a target multi-modal vector.

[0006] According to another aspect of the embodiments of the present application, there is also provided an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to execute the above multi-modal data processing method through the computer program.

[0007] According to another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the above multi-modal data processing method when running.

[0008] In the embodiments of the present application, a method is adopted, which includes obtaining multi-modal data to be recognized, where the multi-modal data includes image data and text data; obtaining an image embedding vector corresponding to the image data, and converting the image embedding vector into a pre-fused coding vector based on a pre-trained fusion vocabulary; the pre-trained fusion vocabulary is a coding book obtained according to image training samples for reducing the coding amount of image features; combining the pre-fused coding vector and a text embedding vector corresponding to the text data to obtain a target multi-modal vector. In the above method, based on a preset pre-trained fusion vocabulary, the model can more easily learn the connection between visual features and training corpora, and can also greatly reduce the number of codes required to represent each image, enabling the model to process more image information or larger images within a limited context length. It can not only greatly save the computing resources occupied by the large language model, but also improve the performance of the multi-modal large language model in understanding and fusing multi-modal information; solving the technical problems in the related art that the multi-modal large language model occupies more computing resources, is difficult to effectively understand and fuse multi-modal information, and has poor overall performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of this application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention, and do not constitute an improper limitation to the present invention. In the drawings:

[0010] Figure 1 is a schematic diagram of an application environment of an optional multi-modal data processing method according to the embodiments of the present application;

[0011] Figure 2 is a schematic diagram of an application environment of another optional multi-modal data processing method according to the embodiments of the present application;

[0012] Figure 3 is a schematic flowchart of an optional multi-modal data processing method according to the embodiments of the present application;

[0013] Figure 4 is a schematic diagram of a multi-modal large language model framework of an optional pre-fused vocabulary according to the embodiments of the present application;

[0014] Figure 5 It is a schematic diagram of an optional VQGAN model architecture according to an embodiment of the present application;

[0015] Figure 6 It is an expanded training flow chart of an optional adjacency matrix vocabulary according to an embodiment of the present application;

[0016] Figure 7 It is a schematic diagram of an optional Transformer network structure based on an attention mechanism according to an embodiment of the present application;

[0017] Figure 8 It is a schematic diagram of the structure of an optional multi-modal data processing device according to an embodiment of the present application;

[0018] Figure 9 It is a schematic diagram of the structure of an optional electronic device according to an embodiment of the present application. Detailed implementation manners

[0019] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0020] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0021] According to one aspect of the embodiments of the present application, a multi-modal data processing method is provided. Optionally, as an optional implementation manner, the above multi-modal data processing method can be but is not limited to being applied to, for example Figure 1In the application environment shown. This application environment includes: a terminal device 102 for human-computer interaction with the user, a network 104, and a server 106. There can be human-computer interaction between the user 108 and the terminal device 102, and a multimodal data processing application client runs in the terminal device 102. The above terminal device 102 includes a human-computer interaction screen 1022, a processor 1024, and a memory 1026. The human-computer interaction screen 1022 is used to display image data and text data; the processor 1024 is used to obtain the multimodal data to be recognized. The memory 1026 is used to store the multimodal data to be recognized and the target multimodal vector corresponding to the multimodal data.

[0022] In addition, the server 106 includes a database 1062 and a processing engine 1064. The database 1062 is used to store the multimodal data to be recognized and the target multimodal vector corresponding to the multimodal data. The processing engine 1064 is used to obtain the multimodal data to be recognized, and the multimodal data includes image data and text data; obtain the image embedding vector corresponding to the image data, and convert the image embedding vector into a pre-fused coding vector based on a pre-trained fusion vocabulary; the pre-trained fusion vocabulary is a coding book obtained according to image training samples for reducing the coding amount of image features; combine the pre-fused coding vector and the text embedding vector corresponding to the text data to obtain a target multimodal vector.

[0023] In one or more embodiments, the above multimodal data processing method of the present application can be applied to Figure 2 the application environment shown. As Figure 2 shown, there can be human-computer interaction between the user 202 and the user device 204. The user device 204 includes a memory 206 and a processor 208. In this embodiment, the user device 204 can, but is not limited to, perform the operations performed by the above terminal device 102 with reference to obtain the target multimodal vector corresponding to the multimodal data.

[0024] Optionally, the above terminal device 102 and user device 204 include, but are not limited to, terminals such as mobile phones, tablets, laptops, PCs, in-vehicle electronic devices, wearable devices, etc. The above network 104 can include, but is not limited to, a wireless network or a wired network. Among them, the wireless network includes: WIFI and other networks for implementing wireless communication. The above wired network can include, but is not limited to: wide area network, metropolitan area network, local area network. The above server 106 can include, but is not limited to, any hardware device capable of performing calculations. The above server can be a single server, or a server cluster composed of multiple servers, or a cloud server. The above is only an example, and no limitation is made in this embodiment.

[0025] As an optional implementation manner, as Figure 3As shown in the figure, an embodiment of the present application provides a multi-modal data processing method, which includes the following steps:

[0026] S302, obtain the multi-modal data to be recognized, where the multi-modal data includes image data and text data.

[0027] S304, obtain the image embedding vector corresponding to the image data, and convert the image embedding vector into a pre-fusion coding vector based on a pre-trained fusion vocabulary; the pre-trained fusion vocabulary is a coding book obtained according to image training samples for reducing the coding amount of image features.

[0028] S306, combine the pre-fusion coding vector and the text embedding vector corresponding to the text data to obtain a target multi-modal vector.

[0029] Specifically, in the embodiment of the present application, for the text data in the multi-modal data to be recognized, a common pre-trained text encoder (such as a BPE encoder) can be used to convert it into a text vector t; for the image data in the multi-modal data to be recognized, first use the VQGAN model to preprocess it to obtain an initial coding vector z, and then use the pre-trained fusion vocabulary to replace the existing combination patterns in the vocabulary with the corresponding new image codes (token IDs) by looking up the table, so as to convert the initial coding vector z into a pre-fusion coding vector z m .

[0030] Concatenate the text vector t with the pre-fusion coding vector z obtained after processing the image m to obtain a multi-modal data vector (target multi-modal vector) y = [t; z m .

[0031] In the embodiments of the present application, a method is adopted to obtain multi-modal data to be recognized, where the multi-modal data includes image data and text data; obtain an image embedding vector corresponding to the image data, and convert the image embedding vector into a pre-fusion coding vector based on a pre-trained fusion vocabulary; the pre-trained fusion vocabulary is a coding book obtained according to image training samples for reducing the coding amount of image features; combine the pre-fusion coding vector and a text embedding vector corresponding to the text data to obtain a target multi-modal vector. In the above method, based on a preset pre-trained fusion vocabulary, the model can more easily learn the connection between visual features and training corpora, and can also greatly reduce the number of codes required to represent each image, enabling the model to process more image information or larger images within a limited context length. This can not only greatly save the computing resources occupied by the large language model, but also improve the performance of the multi-modal large language model in understanding and fusing multi-modal information; it solves the technical problem in the related art that the multi-modal large language model occupies more computing resources, is difficult to effectively understand and fuse multi-modal information, and has poor overall performance.

[0032] In one or more embodiments, the multi-modal data processing method further includes:

[0033] Obtain an initial coding book corresponding to a preset image processing model, and image training samples;

[0034] Based on the image processing model, obtain an image embedding vector corresponding to the image training samples;

[0035] Based on the image embedding vector corresponding to the image training samples, update the initial coding book to obtain the pre-trained fusion vocabulary.

[0036] Specifically, in the embodiments of the present application, the above image processing model may be a VQGAN model; taking the VQGAN model as an example, as shown in combination with Figure 4 and Figure 5 , first use the VQGAN model to perform quantization preprocessing on the input image. First, adjust the input image to pixels of size h×h, and then use the pre-trained VQGAN model to encode the image. Through the initial coding book (Codebook) of length N in the VQGAN model, a quantized two-dimensional discrete codebook index array (image embedding vector) of size N×N is obtained; where h and N are positive integers;

[0037] In an example, the array can be flattened into a one-dimensional vector of length N×N. Then, based on the one-dimensional vector of length N×N corresponding to the image training samples, update the initial coding book to obtain the pre-trained fusion vocabulary. For example, adjust the length of the initial coding book to N+K, where K is a positive integer. Here, each character in the initial coding book represents a coding mode.

[0038] In one or more embodiments, the preset image processing model includes a VQGAN model, and updating the initial codebook based on the image embedding vectors corresponding to the image training samples to obtain the pre-trained fusion vocabulary includes the following steps:

[0039] Obtain the image embedding vectors corresponding to the image training samples obtained based on the VQGAN model;

[0040] For each image embedding vector corresponding to an image in the image training samples, perform the following codebook update operations in sequence until the initial codebook reaches the preset code quantity:

[0041] Determine the image coding connection status of the image embedding vector corresponding to the current image, where the image coding connection status includes the horizontal image coding connection status and the vertical image coding connection status;

[0042] According to the image coding connection status, determine whether there is a first coding pair in the current image whose occurrence times are greater than the preset times; the first coding pair indicates two adjacent image codings in the horizontal or vertical direction;

[0043] If there is a first coding pair in the current image whose occurrence times are greater than the preset times, replace the first coding pair with a first updated coding value to obtain a first image embedding vector, and add the first updated coding value to the initial codebook;

[0044] Judge whether there is a second coding pair in the updated image embedding vector whose occurrence times are greater than the preset times; the second coding pair indicates two adjacent image codings in the horizontal or vertical direction; if so, replace the second coding pair with a second updated coding value to obtain a second image embedding vector, and add the second updated coding value to the initial codebook.

[0045] Specifically, in the embodiments of the present application, first obtain the image embedding vectors corresponding to the image training samples obtained based on the VQGAN model;

[0046] Create a first adjacency matrix H and a second adjacency matrix V based on the image embedding vectors. The first adjacency matrix H is used to record the horizontal image coding connection status corresponding to the image training samples, and the second adjacency matrix V is used to record the vertical image coding connection status corresponding to the image training samples;

[0047] Determine all coding pairs (i, j) where H[i][j] is greater than a preset number of times or V[i][j] is greater than the preset number of times; H[i][j] represents the number of times coding i is adjacent to coding j in the horizontal direction; V[i][j] represents the number of times coding i is adjacent to coding j in the vertical direction;

[0048] For each found combination of coding pairs (i, j), if the size of the initial codebook is N, then taking the coding pair (i, j) as a whole, for example, here a new code q can be used to replace the (i, j) coding pair, update the original image embedding vector to obtain a first image embedding vector, and then add code q to the initial codebook. At this time, the size of the initial codebook becomes N + 1;

[0049] Then continue the extended search in the vertical coding direction and the vertical encoding direction to determine all coding pairs (q, m) where H[q][m] is greater than the preset number of times or V[q][m] is greater than the preset number of times; H[q][m] represents the number of times code q is adjacent to code m in the horizontal direction; V[q][m] represents the number of times code q is adjacent to code k in the vertical direction; if any exist, form a second coding pair (q, m), where m is the code adjacent to code q; for example, here a new code n can be used to replace the (q, m) coding pair, update the original image embedding vector to obtain a second image embedding vector, and then add code n to the initial codebook. At this time, the size of the initial codebook becomes N + 2; if there are no coding pairs (q, m) where H[q][m] is greater than the preset number of times or V[q][m] is greater than the preset number of times, then take the next image of the image training sample as the current image and continue to perform the codebook update operation.

[0050] Then, loop through the above steps for multiple image training samples. When the initial codebook reaches the preset number of codes, determine the pre-trained fusion vocabulary.

[0051] Specifically, as Figure 6 shown, in the embodiments of the present application, the sizes of the above-mentioned first adjacency matrix H and second adjacency matrix V can be N×N, where N is the size of the initial codebook. The present application can further loop-train and expand the initial codebook using a large number of image data sets. For example, set the above steps as the number of loops K, and finally train to obtain an expanded vocabulary of size N + K. This expanded vocabulary is the above-mentioned pre-trained fusion vocabulary. Based on this pre-fusion vocabulary, for any obtained picture tensor of length N×N, the K expanded combination patterns that appear in the vocabulary can be replaced with corresponding new image codes, thereby realizing the pre-fusion of the spatial prior information of the one-dimensional picture tensor.

[0052] In one or more embodiments, the multimodal data processing method further includes:

[0053] If there is no first coding pair in the current image whose occurrence times are greater than the preset times, then use the next image of the image training sample as the current image and continue to perform the codebook update operation.

[0054] When the number of codings in the initial codebook reaches the preset number of codings, the pre-trained fusion vocabulary is determined.

[0055] For example, if it is determined that all H[i][j] corresponding to the current image are less than the preset times, or the coding pair (i, j) where V[i][j] is less than the preset times, then use the next image of the image training sample as the current image and continue to perform the codebook update operation. After performing the above codebook update operation multiple times, when the number of codings in the initial codebook reaches the preset number of codings, the pre-trained fusion vocabulary is determined.

[0056] In one or more embodiments, the multimodal data processing method further includes:

[0057] Map the image training sample to a latent space through the encoder of the VQGAN model to generate the latent vector corresponding to the image training sample.

[0058] Discretize the latent vector into a set of discrete vectors through the quantizer of the VQGAN model, and determine the code vector closest to the discrete vector through nearest neighbor search.

[0059] Map the code vector back to the image space through the decoder of the VQGAN model to generate a reconstructed image.

[0060] Adjust the reconstructed image based on the generative adversarial network in the VQGAN model to obtain the image embedding vector corresponding to the image training sample.

[0061] As Figure 5 shown, the VQGAN (Vector Quantized Generative Adversarial Network) model is a model that combines vector quantization technology with a generative adversarial network (GAN), mainly used for the generation and compression of high-quality images. The VQGAN model decomposes an image into a set of discrete vectors (or codes), and then performs generation and reconstruction through GAN, thereby retaining the details and textures of the image. The VQGAN model consists of the following main parts:

[0062] 1. Encoder: Map the input image to a latent space to generate a latent vector.

[0063] 2. Quantizer: Discretize the continuous latent vector into a set of finite codes (or vectors).

[0064] 3. Decoder: Maps the discretized code back to the image space to generate a reconstructed image.

[0065] 4. Generative Adversarial Network (GAN): Consists of a generator and a discriminator, used to improve the quality of image generation.

[0066] In the encoder of VQGAN, the input image x is mapped to a latent space through a deep neural network to generate a latent vector z e (x), and the mapping formula includes formula (1): z e (x) = Encoder(x) (1)

[0067] After that, VQGAN uses Vector Quantization (VQ) technology to map the continuous latent vector z e (x) to a set of discrete vectors, and finds the nearest code vector z q (x) through nearest neighbor search. Specifically, configure e i as the i-th vector in the codebook (Codebook), then the quantization formula includes formula (2):

[0068] z q (x) = e i , where i = arg min j ||z e (x) - e j || (2);

[0069] Among them, z q (x) is the discretized code vector.

[0070] The decoder maps the discretized code vector z q (x) back to the image space to generate a reconstructed image Among them, the mapping formula of the reconstructed image includes formula (3):

[0071] In one or more embodiments, the multimodal data processing method further includes:

[0072] Jointly training the VQGAN model based on a preset reconstruction loss function, quantization loss function, and adversarial loss function, and when the VQGAN model reaches the convergence condition, determining a preset image processing model.

[0073] To ensure that the encoding after the quantization and discretization processing of the image can fully contain the original image information, the VQGAN model is jointly trained based on a preset reconstruction loss function, quantization loss function, and adversarial loss function;

[0074] Regarding the reconstruction loss, it is necessary to measure the difference between the reconstructed image and the original image. Specifically, the corresponding reconstruction loss function includes Equation (4):

[0075]

[0076] Regarding the quantization loss, it is necessary to minimize the difference between the latent vector and its quantized vector. Specifically, the corresponding loss function is shown in Equation (5):

[0077]

[0078] where β is an adjustable weight parameter, and sg represents the Stop-Gradient operation commonly used in deep learning, which is used to retain the gradient of its calculation to ensure the smooth backpropagation of the gradient.

[0079] Regarding the adversarial loss, it is necessary to improve the reconstruction quality of the image. The loss function of the configured generator and the loss function of the discriminator are shown in Equation (6) and Equation (7):

[0080]

[0081] where D(x) represents the discrimination result of the discriminator on the input image.

[0082] In one or more embodiments, the multi-modal data processing method further includes:

[0083] Using the target multi-modal vector as sample data to train a Transformer model;

[0084] When the accuracy of the Transformer model in identifying the data sequence in the sample data reaches a preset threshold, a pre-trained multi-modal data recognition model is determined.

[0085] As Figure 7 shown, the core of the Transformer model is the attention mechanism, which allows the model to dynamically focus on different parts of the input sequence when processing the current element. Its calculation process is as follows: Given a query matrix a key matrix and a value matrix where n is the number of queries, m is the number of key-value pairs, d k and dv They are the dimensions of the key and the value respectively. The attention weights are calculated by Equation (8):

[0086]

[0087] where is a scaling factor used to mitigate the problem that the inner product value becomes larger as the dimension increases.

[0088] To further enhance the model's expressive power, the Transformer introduces the multi-head attention mechanism, including Equation (9):

[0089] MultiHead(Q, K, V) = Concat(head1,..., head h )W O (9)

[0090] The calculation formula for each attention head includes Equation (10):

[0091]

[0092] where W O , and are all trainable parameter matrices.

[0093] After calculating the dependencies between input sequences through the multi-head attention mechanism, a feed-forward neural network (Feed-Forward Neural Network Layer, FFN) is applied to each position. The feed-forward neural network used consists of two linear transformations and a ReLU activation function, Equation (11):

[0094] FFN(x) = max(0, xW1 + b1)W2 + b2 (11)

[0095] where and are trainable weight matrices, and b1 and b2 are bias terms.

[0096] Layer normalization (Layer Normalization) and residual connection (ResidualConnection) are performed after each sublayer, including Equation (12):

[0097] LayerNorm(x + Sublayer(x)) (12)

[0098] In addition, since the Transformer does not contain a convolutional or recurrent structure and cannot capture the position information in the sequence, positional encoding is introduced. The positional encoding formula includes Equation (13):

[0099]

[0100] Among them, pos represents the position, and i represents the dimension index. The position encoding is added to the input embeddings so that the model can utilize the position information.

[0101] The training process of the Transformer model includes two steps: forward propagation and backward propagation. Using the autoregressive model architecture, the attention is calculated and the target sequence is finally generated by iterating the above-described structure through multiple layers, and the model parameters are optimized through backpropagation.

[0102] When the accuracy of the Transformer model in identifying the data sequence in the sample data reaches a preset threshold, a pre-trained multi-modal data recognition model is determined.

[0103] In one or more embodiments, the multi-modal data processing method further includes:

[0104] Obtaining a target multi-modal vector corresponding to an image and text input by a user;

[0105] Identifying image features and text features corresponding to the target multi-modal vector based on a pre-trained multi-modal data recognition model;

[0106] Generating semantic information corresponding to the image and text based on the image features and text features.

[0107] Specifically, as Figure 4 shown, the image input by the user includes the grassland and animals in the image, and the text can be the user's question about the image, such as text information like what animals are there in the picture, etc. First, obtain the image embedding vector corresponding to the image data (for example, an encoding matrix with an initial encoding of 3*3), and convert the image embedding vector into a pre-fused encoding vector [111, 3, 22, 333] based on a pre-trained fusion vocabulary (pre-trained VQGAN codebook); combine the pre-fused encoding vector and the text embedding vector corresponding to the text data to obtain a target multi-modal vector. Based on the pre-trained Transformer model and the pre-trained fusion vocabulary, identify the 3*3-sized original encoding matrix represented by the above pre-fused encoding vector. Generate semantic information corresponding to the image and text based on the image features and text features, for example, including: the language features of this image and this image (depicting a dog on the grassland in the picture, etc.).

[0108] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0109] According to another aspect of the embodiments of the present application, there is also provided a multimodal data processing device for implementing the above multimodal data processing method. As Figure 8 shown, the device includes:

[0110] A first acquisition unit 802, configured to acquire multimodal data to be recognized, where the multimodal data includes image data and text data;

[0111] A second acquisition unit 804, configured to acquire an image embedding vector corresponding to the image data, and convert the image embedding vector into a pre-fusion coding vector based on a pre-trained fusion vocabulary; the pre-trained fusion vocabulary is a coding book obtained according to image training samples for reducing the coding amount of image features;

[0112] A combination unit 806, configured to combine the pre-fusion coding vector and a text embedding vector corresponding to the text data to obtain a target multimodal vector.

[0113] In the embodiments of the present application, a method is adopted, which includes: acquiring multimodal data to be recognized, where the multimodal data includes image data and text data; acquiring an image embedding vector corresponding to the image data, and converting the image embedding vector into a pre-fusion coding vector based on a pre-trained fusion vocabulary; the pre-trained fusion vocabulary is a coding book obtained according to image training samples for reducing the coding amount of image features; combining the pre-fusion coding vector and a text embedding vector corresponding to the text data to obtain a target multimodal vector. In the above method, based on the preset pre-trained fusion vocabulary, the model can more easily learn the connection between visual features and training corpora, and can also greatly reduce the number of codes required to represent each image, enabling the model to process more image information or larger images within a limited context length. It can not only greatly save the computing resources occupied by the large language model, but also improve the performance of the multimodal large language model in understanding and fusing multimodal information; it solves the technical problems in the related art that the multimodal large language model occupies more computing resources, is difficult to effectively understand and fuse multimodal information, and has poor overall performance.

[0114] In one or more embodiments, the multimodal data processing device further includes:

[0115] A third acquisition unit, configured to acquire an initial codebook corresponding to a preset image processing model and an image training sample;

[0116] A fourth acquisition unit, configured to acquire an image embedding vector corresponding to the image training sample based on the image processing model;

[0117] An update unit, configured to update the initial codebook based on the image embedding vector corresponding to the image training sample to obtain the pre-trained fusion vocabulary.

[0118] In one or more embodiments, the update unit includes:

[0119] A first acquisition module, configured to acquire an image embedding vector corresponding to the image training sample obtained based on the VQGAN model;

[0120] A first determination module, configured to sequentially perform the following codebook update operations on the image embedding vector corresponding to each image in the image training sample until the initial codebook reaches a preset code quantity:

[0121] Determine the image coding connection state of the image embedding vector corresponding to the current image, where the image coding connection state includes the image coding connection state in the horizontal direction and the image coding connection state in the vertical direction;

[0122] A second determination module, configured to determine whether there is a first coding pair with an occurrence frequency greater than a preset frequency in the current image according to the image coding connection state; the first coding pair indicates two adjacent image codings in the horizontal direction or the vertical direction;

[0123] A replacement module, configured to, if there is a first coding pair with an occurrence frequency greater than a preset frequency in the current image, replace the first coding pair with a first updated coding value to obtain a first image embedding vector, and add the first updated coding value to the initial codebook;

[0124] A judgment update module, configured to judge whether there is a second coding pair with an occurrence frequency greater than a preset frequency in the updated image embedding vector; the second coding pair indicates two adjacent image codings in the horizontal direction or the vertical direction; if so, replace the second coding pair with a second updated coding value to obtain a second image embedding vector, and add the second updated coding value to the initial codebook.

[0125] In one or more embodiments, the multi-modal data processing device further includes:

[0126] A judgment unit, configured to, if there is no first coding pair with an occurrence frequency greater than a preset frequency in the current image, use the next image of the image training sample as the current image to continue performing the codebook update operation;

[0127] A first determination unit, configured to determine the pre-trained fusion vocabulary when the encoding quantity of the initial codebook reaches a preset encoding quantity.

[0128] In one or more embodiments, the multimodal data processing device further includes:

[0129] A first generation unit, configured to map an image training sample to a latent space through an encoder of the VQGAN model, and generate a latent vector corresponding to the image training sample;

[0130] A discretization unit, configured to discretize the latent vector into a set of discrete vectors through a quantizer of the VQGAN model, and determine the code vector closest to the discrete vector through nearest neighbor search;

[0131] A mapping unit, configured to map the code vector back to the image space based on a decoder of the VQGAN model, and generate a reconstructed image;

[0132] An adjustment unit, configured to adjust the reconstructed image based on a generative adversarial network in the VQGAN model to obtain an image embedding vector corresponding to the image training sample.

[0133] In one or more embodiments, the multimodal data processing device further includes:

[0134] A first training unit, configured to jointly train the VQGAN model based on a preset reconstruction loss function, quantization loss function, and adversarial loss function, and determine a preset image processing model when the VQGAN model reaches a convergence condition.

[0135] In one or more embodiments, the multimodal data processing device further includes:

[0136] A second training unit, configured to use the target multimodal vector as sample data to train a Transformer model;

[0137] A second determination unit, configured to determine a pre-trained multimodal data recognition model when the accuracy rate of the Transformer model in identifying a data sequence in the sample data reaches a preset threshold.

[0138] In one or more embodiments, the multimodal data processing device further includes:

[0139] A fifth acquisition unit, configured to acquire a target multimodal vector corresponding to an image and text input by a user;

[0140] An identification unit, configured to identify image features and text features corresponding to the target multimodal vector based on a pre-trained multimodal data recognition model;

[0141] A generation unit, configured to generate semantic information corresponding to the image and the text based on the image feature and the text feature.

[0142] According to another aspect of the embodiments of the present application, there is also provided an electronic device for implementing the above multi-modal data processing method. The electronic device may be Figure 9 the terminal device or the server shown in the figure. In this embodiment, the electronic device is taken as an example of the server for illustration. As Figure 9 shown in the figure, the electronic device includes a memory 902 and a processor 904. A computer program is stored in the memory 902, and the processor 904 is configured to execute the steps in any of the above method embodiments through the computer program.

[0143] Optionally, in this embodiment, the above electronic device may be at least one network device among multiple network devices in a computer network.

[0144] Optionally, in this embodiment, the above processor may be configured to execute the following steps through a computer program:

[0145] S1. Obtain multi-modal data to be recognized, where the multi-modal data includes image data and text data;

[0146] S2. Obtain an image embedding vector corresponding to the image data, and convert the image embedding vector into a pre-fused coding vector based on a pre-trained fusion vocabulary; the pre-trained fusion vocabulary is a coding book obtained according to image training samples for reducing the coding amount of image features;

[0147] S3. Combine the pre-fused coding vector and a text embedding vector corresponding to the text data to obtain a target multi-modal vector.

[0148] Optionally, those of ordinary skill in the art can understand that Figure 9 the structure shown in the figure is only schematic. The electronic device may also be a terminal device such as a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, and a Mobile Internet Device (MID), a PAD, etc. Figure 9 It does not limit the structure of the above electronic device. For example, the electronic device may further include more or fewer components (such as a network interface, etc.) than those shown in Figure 9 , or have a different configuration from that shown in Figure 9 .

[0149] Among them, the memory 902 can be used to store software programs and modules, such as the program instructions / modules corresponding to the multi-modal data processing method and device in the embodiments of the present application. The processor 904 executes various functional applications and data processing by running the software programs and modules stored in the memory 902, that is, implements the above-mentioned multi-modal data processing method. The memory 902 may include a high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 902 may further include a memory remotely disposed relative to the processor 904, and these remote memories can be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof. Among them, the memory 902 can specifically but not limitedly be used to store image data and text data. As an example, as Figure 9 shown, the above memory 902 may include, but not limited to, the first acquisition unit 802, the second acquisition unit 804, and the combination unit 806 in the above multi-modal data processing device. In addition, it may also include, but not limited to, other module units in the above multi-modal data processing device, which will not be elaborated in this example.

[0150] Optionally, the above transmission device 906 is used to receive or send data via a network. Specific examples of the above network may include wired networks and wireless networks. In one instance, the transmission device 909 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers through a network cable, thereby enabling communication with the Internet or a local area network. In one instance, the transmission device 906 is a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0151] In addition, the above electronic device further includes: a display 908, which is used to display the image data and text data input by the user; and a connection bus 910, which is used to connect the various module components in the above electronic device.

[0152] In other embodiments, the above terminal device or server may be a node in a distributed system. Among them, the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting the multiple nodes through network communication. Among them, the nodes can form a peer-to-peer (P2P, Peer To Peer) network, and any form of computing device, such as servers, terminals and other electronic devices, can become a node in the blockchain system by joining the peer-to-peer network.

[0153] According to one aspect of the present application, there is provided a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above multi-modal data processing method, wherein the computer program is configured to execute the steps in any one of the above method embodiments when running.

[0154] Optionally, in this embodiment, the above computer-readable storage medium may be configured to store a computer program for executing the following steps:

[0155] S1. Obtain multi-modal data to be recognized, where the multi-modal data includes image data and text data;

[0156] S2. Obtain an image embedding vector corresponding to the image data, and convert the image embedding vector into a pre-fusion coding vector based on a pre-trained fusion vocabulary; the pre-trained fusion vocabulary is a coding book obtained according to image training samples for reducing the coding amount of image features;

[0157] S3. Combine the pre-fusion coding vector and a text embedding vector corresponding to the text data to obtain a target multi-modal vector.

[0158] Optionally, in this embodiment, those of ordinary skill in the art can understand that all or part of the steps in the above various methods can be completed by a program instructing the relevant hardware of the terminal device. The program can be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0159] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.

[0160] If the integrated unit in the above embodiment is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in the above computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing one or more computer devices (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present invention.

[0161] In the above embodiments of the present invention, the descriptions of the various embodiments each have their own emphasis. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0162] In several embodiments provided by the present application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.

[0163] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0164] In addition, the functional units in the various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0165] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A multimodal data processing method, characterized in that Including: Obtain multi-modal data to be recognized, where the multi-modal data includes image data and text data; Obtain the image embedding vector corresponding to the image data, the initial codebook corresponding to the preset image processing model, and image training samples; Update the initial codebook based on the image embedding vectors corresponding to the image training samples to obtain a pre-trained fusion vocabulary, including: Obtain the image embedding vector corresponding to the image training sample obtained based on the image processing model; For each image embedding vector corresponding to an image in the image training sample, sequentially perform the following codebook update operations until the initial codebook reaches the preset code quantity: Determine the image coding connection state of the image embedding vector corresponding to the current image, where the image coding connection state includes the horizontal image coding connection state and the vertical image coding connection state; According to the image coding connection state, determine whether there is a first coding pair in the current image whose occurrence times are greater than the preset times; the first coding pair indicates two adjacent image codings in the horizontal or vertical direction; If there is a first coding pair in the current image whose occurrence times are greater than the preset times, replace the first coding pair with a first updated coding value to obtain a first image embedding vector, and add the first updated coding value to the initial codebook; Judge whether there is a second coding pair in the updated image embedding vector whose occurrence times are greater than the preset times; the second coding pair indicates two adjacent image codings in the horizontal or vertical direction; if so, replace the second coding pair with a second updated coding value to obtain a second image embedding vector, and add the second updated coding value to the initial codebook; Convert the image embedding vector into a pre-fused coding vector based on the pre-trained fusion vocabulary; the pre-trained fusion vocabulary is a codebook obtained according to image training samples for reducing the coding amount of image features; Combine the pre-fused coding vector and the text embedding vector corresponding to the text data to obtain a target multi-modal vector.

2. The method according to claim 1, wherein The method further includes: Obtain the image embedding vector corresponding to the image training sample based on the image processing model.

3. The method according to claim 2, wherein The preset image processing model includes a VQGAN model.

4. The method according to claim 3, wherein The method further includes: If there is no first coding pair in the current image whose occurrence times are greater than the preset times, use the next image of the image training sample as the current image and continue to perform the codebook update operation; When the coding quantity of the initial codebook reaches the preset coding quantity, determine that the pre-trained fusion vocabulary is obtained.

5. The method according to claim 3 or 4, characterized in that, The method further includes: Map the image training sample to a latent space through the encoder of the VQGAN model to generate a latent vector corresponding to the image training sample; Discretize the latent vector into a set of discrete vectors through the quantizer of the VQGAN model, and determine the code vector closest to the discrete vector through nearest neighbor search; Map the code vector back to the image space through the decoder of the VQGAN model to generate a reconstructed image; Adjust the reconstructed image based on the generative adversarial network in the VQGAN model to obtain the image embedding vector corresponding to the image training sample.

6. The method according to claim 5, characterized in that, The method further includes: Jointly train the VQGAN model based on a preset reconstruction loss function, quantization loss function, and adversarial loss function. When the VQGAN model reaches the convergence condition, determine the preset image processing model.

7. The method according to any one of claims 1 to 4, characterized in that The method further includes: Use the target multi-modal vector as sample data to train the Transformer model; When the accuracy of the Transformer model in recognizing the data sequence in the sample data reaches a preset threshold, determine the pre-trained multi-modal data recognition model.

8. The method according to claim 7, wherein The method further includes: Obtain the target multi-modal vector corresponding to the image and text input by the user; Based on the pre-trained multi-modal data recognition model, recognize the image features and text features corresponding to the target multi-modal vector; Based on the image features and text features, generate the semantic information corresponding to the image and text.

9. A multimodal data processing device, characterized in that, Includes: A first acquisition unit for acquiring the multi-modal data to be recognized, where the multi-modal data includes image data and text data; A second acquisition unit for acquiring the image embedding vector corresponding to the image data, the initial codebook corresponding to the preset image processing model, and the image training sample; Update the initial codebook based on the image embedding vector corresponding to the image training sample to obtain a pre-trained fusion vocabulary, including: Obtain the image embedding vector corresponding to the image training sample obtained based on the image processing model; For each image embedding vector corresponding to an image in the image training sample, sequentially perform the following codebook update operations until the initial codebook reaches the preset code quantity: Determine the image coding connection state of the image embedding vector corresponding to the current image, where the image coding connection state includes the horizontal image coding connection state and the vertical image coding connection state; According to the image coding connection state, determine whether there is a first coding pair in the current image whose occurrence times are greater than the preset times; the first coding pair indicates two adjacent image codings in the horizontal or vertical direction; If there is a first coding pair in the current image whose occurrence times are greater than the preset times, replace the first coding pair with a first updated coding value to obtain a first image embedding vector, and add the first updated coding value to the initial codebook; Judge whether there is a second coding pair in the updated image embedding vector whose occurrence times are greater than the preset times; the second coding pair indicates two adjacent image codings in the horizontal or vertical direction; if so, replace the second coding pair with a second updated coding value to obtain a second image embedding vector, and add the second updated coding value to the initial codebook; Convert the image embedding vector into a pre-fused coding vector based on the pre-trained fusion vocabulary; the pre-trained fusion vocabulary is a codebook obtained according to the image training sample for reducing the coding amount of image features. A combination unit for combining the pre-fused encoded vector and the text embedding vector corresponding to the text data to obtain a target multi-modal vector.

10. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to execute the method described in any one of claims 1 to 8 through the computer program.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program executes the method described in any one of claims 1 to 8 when running.

Citation Information

Patent Citations

  • Knowledge injection-based text and graph pre-training model processing method and text and graph retrieval system

    CN115759062A

  • Pre-training method, device and equipment of image-text understanding model and storage medium

    CN116796287A