A multi-modal information source joint encoding method

By extracting and decoupling the characteristics of multimodal sources and introducing knowledge bases, the problem of multimodal signals redundancy in the existing technology is solved, efficient source coding is achieved, and bandwidth and storage space requirements are reduced.

CN115604475BActive Publication Date: 2025-06-10XIDIAN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210969884.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-12
Publication Date
2025-06-10
Estimated Expiration
2042-08-12

AI Technical Summary

Technical Problem

The existing source encoding methods cannot effectively handle the association between multiple modal signals, resulting in repeated transmission of redundant information and increasing the bandwidth and storage space requirements.

Method used

By extracting the feature maps of multiple modal sources, the second encoder is used to decouple common features and personality features, and a knowledge base is introduced during the encoding process to remove the correlation between different modal signals and reduce redundancy.

Benefits of technology

The joint encoding of multimodal sources is realized, reducing the need for transmission bandwidth and storage space, and at the same time it has modal scalability, and can restore different modal sources as needed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115604475B_ABST
    Figure CN115604475B_ABST
Patent Text Reader

Abstract

A multi-modal source joint coding method. First, multiple modal sources are processed through corresponding first encoders to extract features and remove redundancy within each modal signal, obtaining corresponding feature maps. Then, multiple groups of feature maps are concatenated and input into a second encoder, which decouples them into a common feature map and an individual feature map. The common feature map represents the common part among different modal sources, and the individual feature map represents the unique features of each modal source. Finally, the individual feature maps and the common feature map of multiple modal sources are decoded through corresponding decoders and the corresponding modal sources are reconstructed, that is, through entropy coding, they are converted into binary bitstreams for storage or transmission. At the decoding end, after entropy decoding of the binary bitstream, the corresponding modal sources are restored through corresponding decoders respectively. The present invention utilizes the correlation between different sources, reduces the repeated transmission of relevant information, reduces the transmission bandwidth, and reduces the storage space. At the decoding end, different modal sources are restored, having modal scalability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of source coding, and particularly to a multi-modal source joint coding method. Background Art

[0002] As a basic technology, source coding is widely used in various fields. Source coding is the product of the combination of multimedia technology and Internet technology in the information age, aiming to represent the source with the fewest bits on the premise of allowing a certain degree of distortion or not allowing distortion. High-efficiency source coding technology can greatly improve the quality of the decoded source and reduce the storage space under limited bandwidth. For example, currently, there are text compression, image compression (such as compression standards like PNG, BMP, JPEG, BPG, WEBP, etc.), video compression (such as H.264 / AVC, H.265 / HEVC, H.266 / VVC, VP9, AV1, AVS1, AVS2, AVS3, etc.), audio coding (such as AAC, etc.). These standards have a common feature that they only target a single type of input. For example, text compression only targets text input, image compression only targets images, video compression targets images or videos, and audio coding only targets audio input. They cannot process other forms, and even if they do, they need preprocessing and are inefficient. For example, video compression coding standards cannot directly compress text. Although text can be organized into video form through preprocessing, its content is very different from normal videos and has no practical physical meaning. The technologies in video codec standards are not designed for such abnormal signals. Therefore, even forced coding will be inefficient.

[0003] In practice, several modalities of data are often combined for a certain expression. For example, the most common modalities in TV dramas, movies, etc. include three modalities: video, audio, and subtitles. According to the above standards, almost all current solutions encode the three modalities separately. However, in fact, there is a correlation between these three modality signals, that is, there is a certain degree of redundancy, and existing independent coding methods cannot eliminate such redundancy. Therefore, it is a waste of bandwidth or storage space. Therefore, a method capable of jointly coding multiple modality signals is needed to remove the correlation between different modality signals, reduce redundancy, and thus achieve the purpose of reducing bandwidth and saving storage space. Summary of the Invention

[0004] In order to overcome the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a multi-modal source joint coding method, which reduces the repeated transmission of relevant information during the coding and compression process by utilizing the correlation between different sources, thereby reducing the transmission bandwidth and storage space; the decoding end can restore different modality sources according to needs, that is, it has modality scalability.

[0005] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0006] A multi-modal source joint coding method, comprising the following steps:

[0007] 1) Pass multiple modal sources through corresponding first encoders to extract features and remove redundancy within each modal signal, obtaining corresponding feature maps;

[0008] 2) To remove the correlation between different modal signals, connect multiple groups of feature maps and input them into a second encoder, which decouples them into a common feature map and an individual feature map; the common feature map represents the common part between different modal sources, and the individual feature map represents the unique features of each modal source;

[0009] 3) Pass the individual feature maps and common feature maps of multiple modal sources through corresponding decoders to decode and reconstruct the corresponding modal sources, that is, respectively perform entropy coding, convert them into binary bitstreams for storage or transmission; at the decoding end, after entropy decoding the binary bitstream, the corresponding modal sources are respectively restored through corresponding decoders.

[0010] A knowledge base is introduced to perform joint coding on multi-modal sources; the knowledge base is multi-modal or single-modal. A multi-modal knowledge base means that the knowledge base stores information containing various different forms from different modal sources; single or multiple modal sources obtain indexes for retrieving the knowledge base through "modal parsing", and "modal parsing" is to obtain knowledge base node entities for querying and reasoning.

[0011] In one manifestation form of the multi-modal knowledge base, there is text and image, which are represented by nodes and edges. Each node represents an entity, or represents text, or represents an image, and each edge represents the relationship between different nodes.

[0012] The beneficial effects of the present invention are as follows: The present invention proposes a multi-modal source joint coding method, which represents each modal source as common features and individual features, and the common features between different modal sources are the same, thereby realizing the joint coding of multiple modal sources. Compared with the independent coding of multiple modal sources, the present invention reduces the repeated transmission of relevant information by utilizing the correlation between different sources during the coding compression process, thereby reducing the transmission bandwidth and storage space. At the same time, the decoding end can restore different modal sources according to needs, that is, it has the advantage of modal scalability.

[0013] Based on the above multi-modal joint coding method, the present invention introduces a knowledge base (where there is strong relevant known information for the source to be coded), increases prior knowledge, explicitly associates sources of different modalities, and uses the prior knowledge in the knowledge base to guide the multi-modal coding process during the coding process. Therefore, compared with the multi-modal joint coding without a knowledge base, it can further save storage space and reduce bandwidth. Brief Description of the Drawings

[0014] Figure 1 This is a flowchart of a multi-modal source joint encoding method according to Embodiment 1 of the present invention.

[0015] Figure 2 This is a flowchart of a knowledge base-assisted multi-modal source joint encoding method according to Embodiment 2 of the present invention.

[0016] Figure 3 This is an image and text multi-modal knowledge base in Embodiment 2 of the present invention.

[0017] Figure 4 This is a flowchart of a knowledge base-assisted multi-modal source joint encoding method according to Embodiment 3 of the present invention. Detailed Embodiments

[0018] The present invention will be described in detail below with reference to the drawings and embodiments.

[0019] Embodiment 1. Embodiment 1 gives an example of taking two sources as inputs. A multi-modal source joint encoding method includes the following steps:

[0020] 1) Given two modal sources "Modal 1" and "Modal 2", denoted as src 1 and src 2 respectively. The two modal signals pass through the first encoder A and the first encoder B respectively to extract features and remove the redundancy within each modal signal, obtaining the feature map feat 1 and the feature map feat 2 . The first encoder A and the second encoder B are not particularly limited and can be a convolutional neural network CNN in a neural network, or a recurrent neural network RNN; the feature map feat 1 and the feature map feat 2 can be a one-dimensional vector, a two-dimensional matrix or even a tensor of a higher dimension;

[0021] 2) In order to remove the correlation between different modal signals, the two groups of feature maps are concatenated and input into the second encoder C, which is decoupled into a common feature map and a personalized feature map; the common feature map represents the common part between different modal sources, usually at the semantic level; the personalized feature map represents the unique features of each modal source; taking two modal sources of video and audio as an example, the common feature may be the words spoken by the people in the video, and this information is usually also included in the audio; the personalized feature of the video can be the appearance of the people in the video or other background information such as flowers and plants outside the people, and the personalized feature of the audio may include other non-related audio, or it can also be the tone that is usually difficult to express in the video;

[0022] This embodiment decouples common and individual features and outputs the individual features feati of modality 1 1 , the common features featc of the two modalities and the individual features feati of modality 2 2 , the second encoder C may include a quantization process to achieve lossy coding, and its structure has no special requirements. It can be a CNN, an RNN or can also include a hyper prior model; in addition, it should be noted that feati 1 , featc and feati 2 The internal characteristics of the three types of features are not necessarily the same. For example, feati 1 may internally include side information featis 1 and feature featii 1 , where the side information featis 1 is used to assist in the generation of featii 1 , and the same applies to featc and feati 2 ;

[0023] 3) featc, feati 1 and feati 2 The three types of features are respectively entropy-coded and converted into binary bitstreams for storage or transmission; at the decoding end, the binary bitstream is entropy-decoded to recover feati 1 , featc and feati 2 ; then feati 1 and featc are jointly input into decoder A to recover modality 1, labeled as featci 1 and featc are jointly input into decoder B to recover modality 2, denoted as

[0024] The above is the process during testing. During the training process, only paired multi-modal data is required for training. During the training process, the encoders and decoders of multiple modalities are trained end-to-end together, and the loss function is designed in the following form:

[0025]

[0026] where quality 1 (·,·) and quality 2 (·,·) are respectively used to measure the quality loss of modality 1 and modality 2 caused by coding. For example, for videos or images, PSNR (Peak Signal-to-Noise Ratio), MS-SSIM (Multi-Scale Structural Similarity) or perceptual loss can be used for measurement; and are used to measure and The number of bits consumed to convert to a binary bitstream can usually be obtained by estimation. For example, in the above description, it can be assumed that featc, feati 1 and feati 2 The three types of features follow a Gaussian distribution. Use part of the features in featis 1 to represent the mean of the Gaussian distribution, and the other part of the features to represent the variance. That is, if the encoder adopts the variational autoencoder (VAE) structure, then the bitrate and can be estimated using Shannon entropy; λ in the formula 1 , λ, λ 3 belongs to hyperparameters. λ 1 controls the trade-off between the reconstruction qualities of modality 1 and modality 2. That is, when it is more desirable that the source distortion of modality 1 is smaller, λ 1 can be set smaller, and vice versa; λ 3 allocates the bitrate between modality 1 and modality 2. That is, when the total bandwidth or storage space requirement for the two modalities is fixed, when λ 3 is larger, it tends to have a larger bitrate for modality 1 and a smaller bitrate for modality 2, and vice versa; λ is used to control the trade-off between quality and bitrate. Usually, the higher the quality, the larger the bitrate consumed, and the lower the quality, the smaller the bitrate consumed. That is, λ is used to select the final bitrate point. The larger λ is, the lower the selected bitrate point, which is applicable to scenarios with lower bandwidth, and the corresponding reconstruction quality will be lower, and vice versa.

[0027] Example 2. Refer to Figure 2 , Example 2 introduces a knowledge base on the basis of Example 1, which can more efficiently perform joint encoding of multi-modal sources.

[0028] Figure 2 The knowledge base in Figure 3 can be either multi-modal or single-modal. A multi-modal knowledge base refers to a knowledge base that stores information in different forms (usually from different modal sources); Figure 3An image of Claude Shannon is given in the lower right corner. "Claude Shannon" and its image are connected by a directed edge "imageOf". "Deep Thought" participates in the "World Computer Chess Championship". Two nodes respectively represent "Deep Thought" and the "World Computer Chess Championship", and the relationship between them is represented by "attend".

[0029] In Example 2, a knowledge base is introduced on the basis of Example 1. On the basis of Example 1, the source of Modality 1 can obtain the index for retrieving the knowledge base through "Modality 1 Parsing", and the source of Modality 2 can also obtain the index for retrieving the knowledge base through "Modality 2 Parsing". Either one of them is also acceptable. With two types of parsing, more relevant information can be retrieved from the knowledge base or the robustness can be enhanced, which has a greater effect on improving the coding efficiency of the multi-modal source. Among them, "Modality 1 Parsing" and "Modality 2 Parsing" are mainly for obtaining the knowledge base node entities for query and reasoning. After the reasoning and query of the knowledge base, the relevant information can be embedded and encoded by the third encoder D to obtain the knowledge base features, which are jointly encoded with the source features by the second encoder C to remove the redundancy between the source coding and the knowledge base, thereby improving the coding efficiency. Correspondingly, in the decoding process, the decoder A and the decoder B also need to input the knowledge base features to decode the sources of Modality 1 and Modality 2.

[0030] The purpose of the knowledge base introduced in Example 2 is to increase the prior knowledge and explicitly associate the sources of different modalities.

[0031] The specific process of Example 2 is as follows: A multi-modal source joint coding method includes the following steps:

[0032] 1) Given two modal sources "Modality 1" and "Modality 2", denoted as src 1 and src 2 respectively. The two modal signals respectively pass through the first encoder A and the first encoder B to extract features and remove the redundancy within each modal signal, obtaining the feature map feat 1 and the feature map feat 2 ;

[0033] The source of Modality 1 obtains the index for retrieving the knowledge base through "Modality 1 Parsing", and the source of Modality 2 obtains the index for retrieving the knowledge base through "Modality 2 Parsing". Among them, "Modality 1 Parsing" and "Modality 2 Parsing" are mainly for obtaining the knowledge base node entities for query and reasoning. After the reasoning and query of the knowledge base, the relevant information is embedded and encoded by the encoder D to obtain the knowledge base features;

[0034] 2) To remove the correlation between different modality signals, the two sets of feature maps are concatenated and input into the second encoder C, which decouples them into a common feature map and a personalized feature map. The common feature map represents the common part between different modality sources, usually at the semantic level. The personalized feature map represents the unique features of each modality source. Taking video and audio as two modality sources as an example, the common feature may be the words spoken by the people in the video, which is usually also included in the audio. The personalized feature of the video can be the appearance of the people in the video or other background information such as flowers and plants outside the people. The personalized feature of the audio may include other non-related audio or the tone that is usually difficult to express in the video.

[0035] In this embodiment, the common and personalized features are decoupled, and the personalized feature feati of modality 1 is output. 1 , the common feature featc of the two modalities and the personalized feature feati of modality 2 2 , and the second encoder C may include a quantization process to achieve lossy coding.

[0036] The knowledge base feature and the source feature are jointly encoded through the second encoder C to remove the redundancy between the source coding and the knowledge base, thereby improving the coding efficiency.

[0037] 3) featc, feati 1 and feati 2 These three types of features are respectively entropy encoded and converted into binary bitstreams for storage or transmission. At the decoding end, after entropy decoding of the binary bitstream, feati 1 , featc and feati 2 are restored; then feati 1 and featc are jointly input into decoder A to restore modality 1, marked as feati 1 and featc are jointly input into decoder B to restore modality 2, denoted as

[0038] During the decoding process, decoder A and decoder B also need to input the knowledge base feature to decode the modality 1 and modality 2 sources.

[0039] Example 3, referring to Figure 4 Example 3 gives an example of introducing a knowledge base. The role of the knowledge base is to query the image of the person himself in the knowledge base according to the keyword "Claude Shannon" in the "text" source, so that there is no need to encode the image part corresponding to Claude Shannon in the "image" source. Therefore, the image and text can be encoded more efficiently.

[0040] Referring to Figure 4, the inputs of this embodiment are two modal sources, "text" and "image", corresponding to "Modal 1" and "Modal 2" in Embodiment 2 respectively. For the text source, "Named Entity Recognition: BERT" corresponds to "Modal 1 Parsing", that is, the BERT technology in the field of natural language processing can be used to parse the named entities in the text to obtain entity names, such as "Claude Shannon" and "Deep Thought", which are input into the knowledge base for query and reasoning. After encoding, knowledge base features are generated. These features are usually embedded feature vectors; Figure 2 does not parse Modal 2, that is, it does not utilize Figure 4 "Modal 2 Parsing" in Figure 2 . For the main branch, the "text" modality is encoded into text features by a text encoder, such as GRU. The "image" modality detects the objects in the image and establishes the relationships between the objects through the scene graph generation technology. The scene graph generates an image feature map through a convolutional network, which is marked as image features. Then, after the text features and image features are concatenated, they are jointly used as inputs with the knowledge base features and fed into the second encoder C for encoding to generate text personality features, image personality features, and common features of text and image. Figure 4 does not show the process of losslessly encoding the features into a binary bitstream and the part of decoding the binary bitstream to generate the corresponding features. In addition, the parsed "entity names" also need to be encoded and transmitted to the decoding end.

[0041] At the decoding end, the text personality features, common features, and knowledge base features are jointly used as inputs and the text is output through a text decoder; the image personality features, common features, and knowledge base features are jointly used as inputs and the image is output through an image decoder. From Figure 4 it can be seen that by introducing the knowledge base, the encoding end does not need to transmit the part corresponding to Claude Shannon in the image, but only needs to transmit the parsed Claude Shannon entity. The decoding end can obtain the image corresponding to Claude Shannon in the knowledge base through the knowledge base; in addition, it is not necessary to transmit the encoding of "Edmonton" and "1989", which can be obtained through transmission and reasoning by the knowledge base. The personality features of the image at the encoding end mainly include the clothing, posture, and position features of "Feng-hsiung Hsu". The personality features in the text mainly include "Feng-hsiung Hsu" and "first prize"; the common features include information such as "Claude Shannon" and "Deep Thought". Therefore, adding the knowledge base makes the encoding more efficient. The training process of this embodiment is similar to that of Embodiment 1, and the design of the loss function is also similar.

Claims

1. A multi-modal source joint encoding method, characterized in that, it includes the following steps: 1) Pass multiple modal sources through corresponding first encoders to extract features and remove the redundancy within each modal signal, obtaining corresponding feature maps; 2) In order to remove the correlation between different modal signals, connect multiple groups of feature maps and input them into a second encoder, which decouples them into a common feature map and a personalized feature map; The common feature map represents the common part between different modal sources, and the personalized feature map represents the unique features of each modal source; 3) Pass the personalized feature maps and the common feature map of multiple modal sources through corresponding decoders to decode and reconstruct the corresponding modal sources, that is, perform entropy encoding respectively, convert them into binary bitstreams for storage or transmission; after entropy decoding the binary bitstreams at the decoding end, obtain the first personalized feature, the common feature, and the second personalized feature; After that, the first personalized feature and the common feature are jointly input into the first decoder to recover the first modal source; the second personalized feature and the common feature are jointly input into the second decoder to recover the second modal source.

2. The method according to claim 1, characterized in that: Introduce a knowledge base to perform joint encoding on multi-modal sources; the knowledge base is multi-modal or single-modal. A multi-modal knowledge base means that the knowledge base stores information containing multiple different forms from different modal sources; single or multiple modal sources obtain the index for retrieving the knowledge base through "modal parsing", and "modal parsing" is to obtain the node entities of the knowledge base for querying and reasoning; the obtained results are subjected to embedding encoding by a third encoder to obtain knowledge base features, which are jointly encoded with the source features through a second encoder to remove the redundancy between the source encoding and the knowledge base; correspondingly, in the decoding process, the first decoder and the second decoder also need to input the knowledge base features to decode the first modal source and the second modal source.

3. The method according to claim 2, characterized in that: In one manifestation form of the multi-modal knowledge base, there are text and images, which are represented by nodes and edges. Each node represents an entity, either representing text or representing an image, and each edge represents the relationship between different nodes.

Citation Information

Patent Citations

  • Encoding integration system and method and decoding integration system and method

    CN101141644A

  • Multi-modal entity linking method and device and computer readable storage medium

    CN110928961A

  • Complex scene voice recognition method and device based on multiple modes

    CN112151030A