Machine-oriented multi-modal collaborative coding device and method of using the same

CN117440163BActive Publication Date: 2026-09-25XIDIAN UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311154946.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-07
Publication Date
2026-09-25
Estimated Expiration
2043-09-07

AI Technical Summary

Technical Problem

当前大多数VCM方法的实现方式主要是对特征进行编解码,因此可以简单拓展到多模态任务中,即将每个模态对应的特征独立使用VCM方法进行编码,但是这么做就无法利用多模态数据之间的相关性来降低带宽

Benefits of technology

[0026]与现有技术相比,本发明实施例公开的一种面向机器的多模态协同编码装置及其运用方法,包括特征抓取模块,用于从各模态的原生特征中抓取对应的专属嵌入特征;特征编解码模块,用于对所述专属嵌入特征进行编码压缩,通过信道传输在解码端进行特征解码获得重构特征;模态想象模块分为前向模态想象模块和后向模态想象模块,用于通过已有模态信息想象出丢弃模态信息;所述前向模态想象模块,以所述重构特征为输入,获得前向丢弃模态特征以及联合多模态特征;所述后向模态想象模块,以所述前向丢弃模态特征为输入,获得后向丢弃模态特征;分类器模块,用于以所述联合多模态特征为输入,获得概率分布结果,以进行多模态任务情感识别。因此,本发明实施例能够舍弃数据级别的保真而采用了语义级别的保真,对中间特征进行压缩、传输来降低带宽负载;在编码端丢弃部分模态信息来减轻数据传输时的带宽负载,在解码端时利用模态之间的相关性,通过其他模态来恢复丢弃的模态信息,以保证智能任务的效能。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117440163B_ABST
    Figure CN117440163B_ABST
Patent Text Reader

Abstract

The application discloses a kind of machine-oriented multimodal collaborative coding device and its use method, the device includes feature capture module, for from the native feature of each modality Capture corresponding exclusive embedded feature;Feature encoding and decoding module, for the exclusive embedded feature is encoded and compressed, is transmitted in decoding end by channel and is obtained reconstruction feature by feature decoding;Modality imagination module, for discarded modality information is imagined by existing modality information, and obtains forward discarded modality feature and joint multimodal feature;Classifier module, for the joint multimodal feature is input, obtains probability distribution result, to carry out multimodal task sentiment recognition.Therefore, the embodiment of the application can discard data level fidelity and adopt semantic level fidelity, compress, transmission to reduce bandwidth load.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of collaborative coding technology, and in particular to a machine-oriented multimodal collaborative coding device and its application method. Background Technology

[0002] Artificial intelligence (AI) is one of the hottest topics in the technology field today, and it has been widely applied in areas such as natural language processing, computer vision, and machine learning. As AI applications continue to expand, many AI systems require extensive machine-to-machine communication to process massive amounts of data and information. This results in significant bandwidth consumption, making reducing bandwidth load and network computation a major concern. Encoding and decoding methods are crucial for addressing this issue.

[0003] Multimodal tasks are a typical type of task in machine intelligence analysis with wide applications and research value. They aim to enable machines to understand and analyze richer and more accurate information through different multimodal data inputs (such as images, videos, text, and audio). Real life often contains complex and diverse information, and using multimodal technology can better solve these practical problems.

[0004] Currently, for multimodal tasks, the methods for encoding multimodal data mainly fall into two categories: traditional encoding / decoding methods and machine-oriented encoding / decoding methods. Traditional encoding / decoding methods are a common information compression and transmission approach. Their main goal is to represent source information using as few bits as possible while ensuring data reconstruction quality, thereby reducing bandwidth load. This method emphasizes data-level fidelity and can encode and compress data such as images, audio, and video. When encoding multimodal data, traditional encoding / decoding methods encode and compress each modality using a corresponding encoder, then decode and reconstruct the data at the decoding end. The reconstructed data is then used as input to the model for intelligent analysis.

[0005] Video Coding for Machines (VCM) focuses on completing intelligent tasks, aiming to reduce bandwidth load while ensuring the performance of these tasks, and emphasizing semantic-level fidelity. Intermediate feature compression has always been a topic of great interest in the VCM field. Figure 4The demonstration showcases three implementations of VCM: The first involves splitting the network model into two parts: Module 1 is placed on an edge device (encoder end) to extract intermediate features from the data; while Module 2 is placed in the cloud (decoder end) to perform intelligent analysis using the features as input and obtain the analysis results. The feature encoding and decoding method uses common data compression tools (GZIP, ZLIB, BZIP2, and LZMA, etc.) to encode and compress the features, and then decodes them at the decoder end to obtain the features. The second method divides the network model into two parts, Module 1 and Module 2, and places them on the encoder and decoder ends respectively. At the encoder end, the features are packaged into a format similar to F... C×W×H The three-dimensional form is treated as a video and compressed using a video encoder. At the decoding end, the video containing feature information is unpacked to obtain the features as input to module 2. Thirdly, the correlation between source information and features is used to jointly encode source information and features. The upper part uses the aforementioned method to encode and decode features, while the lower part performs collaborative analysis of features and source information to reduce the redundancy of source information before encoding the source information. At the decoding end, source information and features are reconstructed simultaneously. This achieves both the purpose of VCM and the purpose of traditional encoding and decoding. It ensures the efficiency of intelligent tasks, the ability to reconstruct data-level information, and reduces bandwidth load.

[0006] Traditional encoding and decoding methods offer good data-level fidelity. However, due to the different properties between different modalities, specific modalities can only be encoded and decoded using specific methods. Directly using traditional encoding and decoding methods to compress multimodal data cannot utilize the correlation between modalities. Furthermore, due to the data fidelity requirement, the compressed information still contains considerable redundancy, which also leads to wasted bandwidth.

[0007] VCM reduces data redundancy because it does not have the constraint of data-level fidelity. Most current VCM methods are implemented by encoding and decoding features, so they can be easily extended to multimodal tasks, that is, encoding the features corresponding to each modality independently using the VCM method. However, this will not be able to take advantage of the correlation between multimodal data to reduce bandwidth. Summary of the Invention

[0008] This invention provides a machine-oriented multimodal cooperative coding device and its application method, which abandons data-level fidelity and adopts semantic-level fidelity, compresses and transmits intermediate features to reduce bandwidth load, and further reduces bandwidth load by utilizing the correlation of multimodalities.

[0009] To achieve the above objectives, in a first aspect, embodiments of the present invention provide a machine-oriented multimodal cooperative coding apparatus, comprising:

[0010] The feature extraction module is used to extract corresponding exclusive embedded features from the native features of each modality;

[0011] The feature encoding and decoding module is used to encode and compress the dedicated embedded features, and then perform feature decoding at the decoding end through channel transmission to obtain the reconstructed features.

[0012] The modal imagination module is divided into a forward modal imagination module and a backward modal imagination module, which are used to imagine discarded modal information based on existing modal information. The forward modal imagination module takes the reconstructed features as input to obtain forward discarded modal features and joint multimodal features. The backward modal imagination module takes the forward discarded modal features as input to obtain backward discarded modal features.

[0013] A classifier module is used to obtain probability distribution results by taking the joint multimodal features as input, so as to perform multimodal task emotion recognition;

[0014] Each modality includes at least one or more of the following: audio modality, video modality, and text modality.

[0015] Secondly, embodiments of the present invention provide a method for using a machine-oriented multimodal cooperative coding device, which applies the aforementioned machine-oriented multimodal cooperative coding device to the testing phase, including:

[0016] The native features of each modality are used as input to the machine-oriented multimodal co-coding device, and the corresponding exclusive embedded features are extracted from the native features of each modality through the feature extraction module;

[0017] The specific embedded features are input into the feature encoding and decoding module for encoding and compression, and then transmitted through the channel to the decoding end for feature decoding to obtain the reconstructed features.

[0018] Using the reconstructed features as input to the modal imagination module, discarded modal information is imagined based on existing modal information, thereby obtaining forward discarded modal features and joint multimodal features;

[0019] The joint multimodal features are input into the classifier module to obtain the probability distribution results for multimodal task emotion recognition.

[0020] Thirdly, embodiments of the present invention also provide a method for using a machine-oriented multimodal cooperative coding device, which applies the aforementioned machine-oriented multimodal cooperative coding device to the training phase, including:

[0021] The native features and all-zero native features of each modality are used as inputs to the machine-oriented multimodal co-coding device. The feature extraction module extracts the corresponding exclusive embedding features from the native features, and the pre-trained feature extraction module obtains the corresponding pre-trained exclusive embedding features from the all-zero native features.

[0022] The reconstructed features are obtained by uniformly processing the specific embedded features;

[0023] Using the reconstructed features as input to the modal imagination module, the forward discard modal features and joint multimodal features are obtained through the forward modal imagination module; using the forward discard modal features as input to the backward modal imagination module, the backward discard modal features are obtained, thereby obtaining the backward loss;

[0024] The joint multimodal features are input into the classifier module to obtain the classification loss;

[0025] The forward loss is calculated based on the pre-trained dedicated embedding features and the forward discard modal features. The joint loss function of the machine-oriented multimodal co-coding device is obtained according to the calculation formula of the joint loss function, so as to optimize the training of the machine-oriented multimodal co-coding device.

[0026] Compared with existing technologies, the present invention discloses a machine-oriented multimodal collaborative coding device and its application method, including a feature extraction module for extracting corresponding dedicated embedded features from the native features of each modality; a feature encoding and decoding module for encoding and compressing the dedicated embedded features, and performing feature decoding at the decoding end through channel transmission to obtain reconstructed features; a modality imagination module divided into a forward modality imagination module and a backward modality imagination module for imagining discarded modality information based on existing modality information; the forward modality imagination module takes the reconstructed features as input to obtain forward discarded modality features and joint multimodal features; the backward modality imagination module takes the forward discarded modality features as input to obtain backward discarded modality features; and a classifier module for taking the joint multimodal features as input to obtain probability distribution results for performing multimodal task emotion recognition. Therefore, the embodiments of the present invention can abandon data-level fidelity and adopt semantic-level fidelity, compress and transmit intermediate features to reduce bandwidth load; discard some modal information at the encoding end to reduce bandwidth load during data transmission; and utilize the correlation between modalities at the decoding end to recover the discarded modal information through other modalities to ensure the performance of intelligent tasks. Attached Figure Description

[0027] Figure 1 This is a schematic diagram of the structure of a machine-oriented multimodal cooperative coding device provided in an embodiment of the present invention;

[0028] Figure 2 This is an architecture diagram of a machine-oriented multimodal cooperative coding device applied in the testing phase, provided by an embodiment of the present invention.

[0029] Figure 3 This is an architecture diagram of a machine-oriented multimodal cooperative coding device applied during the training phase, provided by an embodiment of the present invention.

[0030] Figure 4 This is a schematic diagram illustrating three implementation methods of VCM provided in an embodiment of the present invention;

[0031] Figure 5 This is a feature encoding / decoding module architecture diagram provided in an embodiment of the present invention;

[0032] Figure 6 This is a schematic diagram of a modal imagination module architecture provided in an embodiment of the present invention;

[0033] Figure 7 This is an example diagram of an RD curve provided in an embodiment of the present invention;

[0034] Figure 8 This is a comparison diagram of a traditional encoding / decoding method and a feature encoding / decoding method provided in an embodiment of the present invention under the AVT case;

[0035] Figure 9 This is a radar chart comparing the results of a traditional encoding / decoding method and a feature encoding / decoding method under AVT conditions, as provided in an embodiment of the present invention.

[0036] Figure 10 This is a comparison diagram of a traditional encoding / decoding method and a feature encoding / decoding method provided in the embodiments of the present invention under the ZZT case;

[0037] Figure 11 This is a comparison diagram of discarded modal results under a feature encoding / decoding method provided in an embodiment of the present invention. Detailed Implementation

[0038] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0039] It should be noted that the terms "comprising" and "specific" in this invention, and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such process, method, product, or device.

[0040] Please see Figure 1 , Figure 1 This is a schematic diagram of a machine-oriented multimodal cooperative coding device provided in an embodiment of the present invention. The machine-oriented multimodal cooperative coding device includes:

[0041] The feature extraction module 11 is used to extract the corresponding exclusive embedded features from the native features of each modality;

[0042] Feature encoding and decoding module 12 is used to encode and compress the exclusive embedded features, and to obtain reconstructed features by decoding the features at the decoding end through channel transmission;

[0043] Modal imagination module 13 is divided into a forward modal imagination module and a backward modal imagination module, which are used to imagine discarded modal information based on existing modal information. The forward modal imagination module takes the reconstructed features as input to obtain forward discarded modal features and joint multimodal features. The backward modal imagination module takes the forward discarded modal features as input to obtain backward discarded modal features.

[0044] The classifier module 14 is used to obtain the probability distribution result by taking the joint multimodal features as input, so as to perform multimodal task emotion recognition;

[0045] The modalities include audio modal, video modal, and text modal.

[0046] Specifically, the feature extraction module 11 includes:

[0047] The audio capture unit is used to obtain corresponding audio-specific embedding features by taking the original audio features as input. Specifically, it takes the original audio features as input and uses the max pooling method on the hidden state features of the LSTM network to obtain audio-specific embedding features. Each audio segment corresponds to a vector of length 128.

[0048] The video capture unit is used to obtain corresponding video-specific embedding features by taking the original video features as input. Specifically, it uses an LSTM network to obtain video-specific embedding features by taking the original video features as input, with one vector of length 128 for each video.

[0049] The text extraction unit is used to obtain corresponding text-specific embedding features by taking the original text features as input. Specifically, it uses the TextCNN model to obtain text-specific embedding features by taking the original text features as input, with one sentence corresponding to one vector of length 128.

[0050] The feature splicing unit is used to splice the audio, video and text-specific embedding features to obtain specific embedding features containing full modality information.

[0051] For example, with native feature x a x v and x t As input, obtain the corresponding specific embedding feature h. a h v and h t For the specific embedding feature h a h v and h t By concatenating the features, we can obtain the exclusive embedding feature H containing full modal information; the exclusive embedding feature h a h v and h t These are video-level features with consistent dimensions. When a modality is discarded, the corresponding x is all zeros. Taking the discarded text modality as an example, in this case... Set to all zeros but Not all zero.

[0052] H is composed of the specific embedding feature h. a h v and h t It is composed of multiple parts, so the amount of data transmitted can be reduced by discarding some modal information. Taking discarding text modality as an example, only the specific embedded feature h is processed. a and h v To obtain H, concatenation is performed, which means that compared to concatenating the specific embedding feature h... a h v and h t Theoretically, this reduces the amount of data by one-third. After H is encoded and decoded, it displays all-zero native features at the decoding end. The input is obtained through the text extraction module. With specific embedding feature h a and h v The decoded features are obtained by splicing the data together and then fed into subsequent modules.

[0053] It should be noted that the audio extraction process uses an LSTM (Long Short-Term Memory) network to obtain audio-specific embedding features. Max pooling is used on the hidden state features of the LSTM network to obtain audio-specific embedding features, with each audio segment corresponding to a 128-bit vector. The video extraction process is similar to the audio extraction module, using an LSTM network to obtain video-specific embedding features, with each video corresponding to a 128-bit vector. The text extraction process uses the TextCNN model to obtain text-specific embedding features. The three convolutional blocks in TextCNN have sizes of {3, 4, 5}, and the final output layer has 128 channels, with each sentence corresponding to a 128-bit vector.

[0054] Specifically, the feature decoding module 12 includes:

[0055] The feature encoding unit is used to quantize and tile the dedicated embedded features into a feature image, and to encode and compress the feature image using an image encoding and decoding method to obtain an encoded feature image.

[0056] The feature decoding unit is used to decode the encoded feature image at the decoding end and obtain the reconstructed features after segmentation and dequantization;

[0057] The quantification formula is as follows:

[0058]

[0059] In the formula, max(H) and min(H) represent the maximum and minimum values ​​in the specific embedded feature H, respectively. i This represents the i-th feature value in the specific embedding feature H. Representing the H i The quantized feature values, depth represents the bit depth of the resulting feature image, and floor() is a function that represents rounding down.

[0060] In a specific embodiment, the feature encoding / decoding module architecture is as follows: Figure 5 As shown, after quantizing (converting floating-point to integer) and tiling (converting multi-dimensional features to two-dimensional features) the feature H into a feature image, it is encoded and compressed using image encoding / decoding methods PNG and JPEG. After decoding at the decoding end, it is segmented and dequantized to obtain the reconstructed feature H. c .

[0061] In image encoding and decoding methods, pixel information is represented using integers, while features are represented using floating-point numbers. Therefore, it is necessary to convert the features from floating-point to integer, a process called quantization. This invention uses the following formula to quantize the features:

[0062]

[0063] In the formula, max(H) and min(H) represent the maximum and minimum values ​​in the specific embedded feature H, respectively. i This represents the i-th feature value in the specific embedding feature H. Representing the H i The quantized feature values, where depth represents the bit depth of the resulting feature image, and floor() is a function that rounds down. After quantization, the value of feature H will be converted to an integer.

[0064] Furthermore, features are generally multidimensional, with H... C×H×W For example, it is a feature with C channels, H height, and W width, and the feature image is in two-dimensional form, so H needs to be... C×H×W The feature image is transformed into a two-dimensional form. The method used in this embodiment of the invention is to tile the features, processing each feature block H... i×H×W (i = 1, 2, ..., C), it will be tiled in the feature image from left to right and from top to bottom in the form of pixel blocks.

[0065] It's important to note that the feature encoding / decoding module involves two key steps: 1) quantization uses the `floor()` function for truncation. 2) The feature image is encoded / decoded using PNG and JPEG methods. These two key steps raise two issues: 1) H and the reconstructed feature H. C The values ​​of H and H are not the same, and such errors are called quantization errors. 2) The process is not differentiable, so end-to-end training is not possible. Therefore, this embodiment of the invention simulates quantization error by adding uniform noise to H during the training phase, replacing the feature encoding and decoding module; while in the testing phase, the feature encoding and decoding module is used to encode and decode H.

[0066] Regarding the implementation of the feature encoding / decoding module, this embodiment of the invention employs the lossless compression method PNG and the lossy method JPEG with quality parameters of 95, 90, 85, 80, 75, 70, 65, and 60 to compress the feature images, with a uniform image bit depth of 8. A dedicated embedded feature h is then defined. a h v and h t Since they are vectors of length 128, their shape was changed to compress them into feature blocks of size (16,8).

[0067] Specifically, the modal imagination module 13 is as follows:

[0068] The modal imagination module is located on the decoding end and is implemented through a CRA structure;

[0069] The CRA structure consists of multiple interconnected RA blocks; the forward dropout modal feature or backward dropout modal feature is output by the last RA block; the joint multimodal feature is obtained by extracting the intermediate features of each RA block and concatenating them.

[0070] For example, the modal imagination module architecture is as follows: Figure 6 As shown, it is located at the decoding end, outputting both the dropped modality features containing dropped modality information and the joint multimodal features y. The modality imagination module is implemented using a CRA (Cascade Residual Autoencoder) structure, which utilizes a residual mechanism and has stronger learning and convergence capabilities compared to other autoencoder structures. The CRA structure consists of multiple interconnected RA blocks, with the output of the last RA block being the dropped modality feature. The modality imagination module also extracts intermediate features from each RA block and concatenates them to obtain the joint multimodal features y.

[0071] Taking the discarded text modality as an example, in order to ensure that the modality imagination module can accurately imagine the features of the discarded modality, the feature extraction module uses x a x v and The input is the forward discard modality feature H' obtained through the forward modality imagination module, and the pre-trained feature grabbing module uses... and x t To obtain pre-trained specific embedding features H for input y By letting H' and H y The device is designed to accurately predict and discard modal information. Furthermore, the Cycle Consistency Learning (CCL) method is used to narrow down the information obtained from reconstructing features H. c To forward discard modal feature H ' The mapping space is used to input the forward discard modal feature H' into the backward modal imagination module to obtain the backward discard modal feature H'", and let H' be... c Try to make it as close as possible to "H".

[0072] In this embodiment of the invention, the modal imagination module is constructed using five RA blocks. The feature sizes output by each RA layer in the RA block are 384-256-128-64-128-256-384, and the feature with a size of 64 is used as the intermediate feature of the RA block.

[0073] Specifically, the classifier module 14 includes:

[0074] The classifier module consists of several fully connected layers. The joint multimodal features are passed through several fully connected layers to obtain output features. The output features are then passed through the softmax() function to obtain the probability distribution results for multimodal task emotion recognition.

[0075] The classifier module in this embodiment of the invention consists of three fully connected layers connected sequentially, with output sizes of 128, 128, and 4 respectively. Output features are obtained after passing through the three fully connected layers, and these features are then processed by a function to obtain the probability distribution.

[0076] Furthermore, the machine-oriented multimodal cooperative coding device further includes:

[0077] The pre-trained feature extraction module is used to extract the corresponding pre-trained specific embedding features from the all-zero native features of each modality.

[0078] Specifically, the joint loss function of the machine-oriented multimodal cooperative coding device includes:

[0079] For classification loss, cross-entropy is used as the loss function.

[0080] The forward loss aims to make the forward discarded modal features and the pre-trained proprietary embedding features as similar as possible, and uses mean squared error as the loss function.

[0081] The backward loss, similar to the forward loss term, aims to make the reconstructed features and the backward discarded mode features as similar as possible, and uses the mean squared error as the loss function.

[0082] The formula for calculating the forward loss is:

[0083]

[0084] The formula for calculating the joint loss function is as follows:

[0085] L=λ cls L cls +λ forward L forward +λ backward L backward ,

[0086] In the formula, L forward For the forward loss term, N represents the number of feature values ​​of the forward discard modal features or the pre-trained proprietary embedding features, and H' i This represents the i-th feature value in the forward discard modal feature H'. The pre-trained specific embedding feature H represents y The i-th eigenvalue, λ cls , λforward and λ backward The hyperparameter is used to control the classification loss term L. cls The forward loss term L forward and the backward loss term L backward Weight allocation.

[0087] Specifically, the native features of each modality are frame-level features, and the methods for capturing the native features of each modality include:

[0088] The audio modality is extracted using the tool OpenSMILE to obtain the original audio features. For example, the "ComParE_2016" configuration file in OpenSMILE is used to extract the original audio features, with each 0.1-second audio segment corresponding to a vector of length 130.

[0089] The video is sampled, faces are detected and cropped, and the original video features are obtained based on the DenseNet model. For example, taking facial features as an example, after sampling the video, faces are detected and cropped, and the cropped facial images are 64×64 pixels in size. The facial images are then fed into a DenseNet model pre-trained on the FER+ dataset (Facial Expression Recognition Plus), and the output features are the original video features, with one vector of length 342 for each image.

[0090] The BERT-large model is used to extract the native features of the text modality, thus obtaining the native text features. For example, in this embodiment of the invention, the BertModel class in Python's transformers package is used to obtain them, and the term "bert-large-uncased" is used to load the pre-trained model, with each word corresponding to a vector of length 1024.

[0091] It should be noted that this embodiment of the invention uses the IEMOCAP dataset to test the proposed machine-oriented multimodal cooperative coding device. The IEMOCAP dataset is an emotion recognition dataset containing five different dialogue scenarios, each featuring multiple dialogues between a male actor and a female actor. For each dialogue, IEMOCAP provides four types of data: video, audio, text, and emotion recognition labels. In this embodiment, the four labels {happy, angry, neutral, sad} are used as emotion recognition labels, and the obtained data statistics are shown in Table 1.

[0092] Table 1 IEMOCAP Data Statistics

[0093] IEMOCAP 1636 1103 1708 1084 5531

[0094] This embodiment uses K-fold cross-validation to validate the proposed machine-oriented multimodal cooperative coding device. The IEMOCAP dataset contains five different dialogue scenarios, each featuring a male actor and a female actor, for a total of ten different actors. Therefore, the dataset is divided by actor. Taking the male actor's data from dialogue scenario 1 as the test set, the female actor's data from dialogue scenario 1 serves as the validation set, and the data from the other actors in dialogue scenarios 2, 3, 4, and 5 serve as the training set. Ten tests are required, and the average result of these ten tests is used as the emotion recognition evaluation metric.

[0095] To accommodate all dropout modalities, a dataset was created based on the IEMOCAP dataset to fit all dropout modalities. The dropout modalities are shown in Table 2, and seven different dropout modalities were simulated for each dataset. During training, the dropout modality was randomly determined, and the corresponding x was set to zero. During testing, all seven dropout modalities were simulated for each dataset, and evaluation metrics were obtained for each of the seven scenarios. Thus, the dataset used for both training and testing is actually seven times larger than the IEMOCAP dataset.

[0096] Table 2 Different cases of discarding modalities

[0097]

[0098] The training parameters are set as shown in Table 3. All programs are written in Python and use the PyTorch framework. The model is trained on an NVIDIA GeForce RTX 3080Ti.

[0099] Table 3 Hyperparameter Settings

[0100]

[0101] This embodiment uses the rate-distortion curve (RD) as the evaluation metric, and its sample graph is shown below. Figure 7 As shown, R on the horizontal axis represents bitrate, indicating bandwidth load, while D on the vertical axis represents the intelligent task evaluation metric, indicating model performance. The RD curve reflects how the evaluation metric changes with bandwidth load during intelligent task execution, and allows for comparison of the advantages and disadvantages of different methods by fixing the values ​​of the horizontal and vertical axes (e.g., ...). Figure 7 The evaluation metrics D1 and D2 of Experiment A and Experiment B are compared using a fixed bandwidth load value of R1.

[0102] In this experiment, Accuracy, Recall, and F1 were selected as evaluation metrics for the emotion recognition task, while the bit rate (Kbps) represents the bandwidth load. Encoding using PNG and JPEG with quality parameters of 95, 90, 85, 80, 75, 70, 65, and 60 yields nine points. Connecting these points sequentially generates the RD curve.

[0103] It's important to note that in the feature encoding / decoding module, features are compressed and transmitted as feature images. Since images lack a temporal dimension, their Kbps value cannot be calculated. Therefore, the following method is used to calculate the image's Kbps value:

[0104]

[0105] Where pngsize represents the image size in bytes, and videotime represents the video duration corresponding to the feature.

[0106] To verify the effectiveness of the present invention and the proposed technical solution, a comparison was made between traditional encoding / decoding methods and feature encoding / decoding methods.

[0107] The traditional encoding and decoding methods used for each modal data are implemented as follows:

[0108] Audio modalities: The Opus method is used to encode and decode audio modalities, which can achieve good quality at low bitrates and excellent performance at high bitrates.

[0109] Video modality: H.266 / VVC is used for video encoding and decoding. On the encoding side, this invention preprocesses the video, first cropping out facial image sequences from the video before synthesizing the video for encoding.

[0110] Text modality: Huffman coding is used to encode and decode text modality. It can represent frequently occurring characters with shorter codes, thereby reducing storage space.

[0111] The traditional method for calculating Kbps in encoding and decoding is shown in the following formula, kbps a kbps v and kbps t These represent Kbps for audio, video, and text, respectively, and binsize. a binsize v and binsize t This represents the size of the binary stream file after encoding audio, video, and text modalities, in bytes (kbps). trad This is the bitrate of the traditional encoding and decoding method.

[0112]

[0113] For video modality, the bitrate was controlled by setting the parameter QP, with experiments conducted at QP values ​​of 56, 39, and 34. For audio modality, the bitrate was controlled by setting the parameter bitrate, with experiments conducted at bitrate values ​​of 1, 3, 5, 7, and 9. For text modality, Huffman coding is lossless, therefore only a single fixed bitrate can be obtained. The RD curves of traditional encoding / decoding methods are obtained from points obtained under different permutations and combinations of parameters QP and bitrate, totaling 15 points.

[0114] Since the experimental results of the feature encoding and decoding methods were obtained using the same parameters, to ensure fairness, the same parameters are also used for traditional encoding and decoding methods (this model is called the baseline). trad This is to obtain experimental results rather than training different devices as the parameters QP and bitrate change. (Baseline) trad The training method involves using unencoded data as input for training, while the testing method uses encoded data as input to obtain experimental results.

[0115] Since there are seven possible scenarios for discarding modes, resulting in numerous and repetitive results, only a subset of meaningful results are presented. The results obtained when modes are not discarded are as follows: Figure 8 As shown. Figure 8 In the diagram, (a), (b), and (c) show the RD curves when the vertical axis is Accuracy, Recall, and F1, and the horizontal axis is Kbps, respectively. AVT_Feat represents the RD curve obtained using the feature-based encoding / decoding method in the AVT case, while AVT_Trad represents the RD curve obtained using the traditional encoding / decoding method in the AVT case. AVT indicates that Audio, Video, and Text are not discarded.

[0116] For the feature-based encoding / decoding method, with a Kbps range of 0-1, it can transmit data at extremely low bitrates, achieving the same or even higher performance as traditional encoding / decoding methods with only a small amount of data transmitted. To more intuitively compare the results, a radar chart was created comparing the results of the feature-based encoding / decoding method compressed using the PNG encoding / decoding method with the results of the traditional encoding / decoding method using QP of 34 and a bitrate of 9. The results are shown below. Figure 9As shown, the Kbps of the feature-based encoding / decoding method is 0.49, while the Kbps of the traditional encoding / decoding method is 11.26. However, the task evaluation metrics obtained using the feature-based encoding / decoding method are superior to those of the traditional method. Quantitative analysis shows that compared to the traditional encoding / decoding method, the feature-based encoding / decoding method reduces the bitrate by 95% while improving accuracy by 3%.

[0117] The results obtained when transmitting only the text modality are as follows Figure 10 As shown. Figure 10 In the diagram, (a), (b), and (c) show the RD curves when the vertical axis is Accuracy, Recall, and F1, and the horizontal axis is Kbps, respectively. ZZT_Feat represents the RD curve obtained using the feature-based encoding / decoding method in ZZT mode, while ZZT_Trad represents the RD curve obtained using the traditional encoding / decoding method in ZZT mode. The Z in ZZT stands for Zero, representing the case where video and audio modalities are discarded and only the text modal is transmitted.

[0118] Feature encoding / decoding methods achieve a bitrate between 0.05 and 0.15 Kbps, while traditional encoding / decoding methods, due to Huffman coding being a lossless method, can only achieve a fixed bitrate. Using the PNG method for feature encoding / decoding yields 0.142 Kbps, while traditional encoding / decoding achieves 0.053 Kbps. Traditional encoding / decoding methods outperform feature encoding / decoding methods in both bandwidth load and model performance. This is because the amount of data used to represent textual modal information is relatively small, while feature encoding / decoding methods require transmitting feature images, which necessitates a larger number of bits to represent information. Furthermore, the presence of quantization errors in feature encoding / decoding methods can negatively impact model performance.

[0119] To investigate the changes in network load and model performance when dropping modalities, a modal dropout experiment was conducted using the feature encoding / decoding method. The results are as follows: Figure 11 As shown in the figure, the meaning of each label is shown in Table 4.

[0120] Compared to the AVT case, when discarding a modality, both the upper and lower limits of Kbps decrease, and the value range also shrinks. The shrinking upper and lower limits indicate that the amount of data transmitted is reduced when discarding a modality, while the shrinking value range is because the size of the transmitted feature image becomes smaller, resulting in less redundancy. Consequently, the amount of redundancy that can be eliminated when using image encoding / decoding methods for compression is also reduced.

[0121] Analyzing the AVT and AZT cases, when Kbps is greater than 0.15, the AVT case slightly outperforms the AZT case, while when Kbps is less than 0.15, the AZT case slightly outperforms the AVT case. Since the dominant text and audio modalities more easily evoke information from the video modal, the model performance in the AZT case is not significantly different from that in the AVT case. Furthermore, when Kbps is less than 0.15, at the same Kbps, the AVT JPEG encoding / decoding method has smaller quality parameters and is more susceptible to quantization errors, thus the task evaluation index drops faster in the AVT case. Therefore, this invention suggests that when a smaller bandwidth load is desired for data transmission, discarding the video modal and transmitting only the video and audio modalities can achieve better results.

[0122] Table 4 Figure 11 Meaning of each label

[0123] AVT The case of no modal discard AZZ Cases of discarding video and text modalities ZVZ Cases of discarding audio and text modalities ZZT Cases of discarding video and audio modalities AVZ Discarding the text modality AZT Case of discarding video modalities ZVT Case of discarding audio modalities

[0124] Figure 2 This is an embodiment of the invention providing an architecture diagram of a machine-oriented multimodal cooperative coding device applied in the testing phase. The method for using this machine-oriented multimodal cooperative coding device in the testing phase includes:

[0125] The native features of each modality are used as input to the machine-oriented multimodal co-coding device, and the corresponding exclusive embedded features are extracted from the native features of each modality through the feature extraction module;

[0126] The specific embedded features are input into the feature encoding and decoding module for encoding and compression, and then transmitted through the channel to the decoding end for feature decoding to obtain the reconstructed features.

[0127] Using the reconstructed features as input to the modal imagination module, discarded modal information is imagined based on existing modal information, thereby obtaining forward discarded modal features and joint multimodal features;

[0128] The joint multimodal features are input into the classifier module to obtain the probability distribution results for multimodal task emotion recognition.

[0129] For example, with native feature x a x v and x t As input, the feature extraction module extracts the corresponding exclusive embedded features from the original features of each modality; the feature encoding and decoding module is used to encode and compress the features, and the features are decoded at the decoding end through channel transmission; the modality imagination module can imagine the discarded modality information based on the existing modality information and obtain joint multimodal features; the classifier module takes the joint multimodal features as input to obtain the probability distribution of emotion recognition.

[0130] Figure 3 This is an embodiment of the invention providing an architecture diagram of a machine-oriented multimodal co-coding device applied during the training phase. The method of using this machine-oriented multimodal co-coding device, applying it during the training phase, includes:

[0131] The native features and all-zero native features of each modality are used as inputs to the machine-oriented multimodal co-coding device. The feature extraction module extracts the corresponding exclusive embedding features from the native features, and the pre-trained feature extraction module obtains the corresponding pre-trained exclusive embedding features from the all-zero native features.

[0132] The reconstructed features are obtained by uniformly processing the specific embedded features;

[0133] Using the reconstructed features as input to the modal imagination module, the forward discard modal features and joint multimodal features are obtained through the forward modal imagination module; using the forward discard modal features as input to the backward modal imagination module, the backward discard modal features are obtained, thereby obtaining the backward loss;

[0134] The joint multimodal features are input into the classifier module to obtain the classification loss;

[0135] The forward loss is calculated based on the pre-trained dedicated embedding features and the forward discard modal features. The joint loss function of the machine-oriented multimodal co-coding device is obtained according to the calculation formula of the joint loss function, so as to optimize the training of the machine-oriented multimodal co-coding device.

[0136] For example, uniform noise was used instead of the feature encoding / decoding module, and a pre-trained feature grabbing module (with fixed parameters during training) and a backward modality imagination module were added to ensure that the model could accurately imagine the discarded modal features.

[0137] In summary, the machine-oriented multimodal collaborative coding device and its application method disclosed in this embodiment of the invention include: a feature extraction module for extracting corresponding dedicated embedded features from the native features of each modality; a feature encoding and decoding module for encoding and compressing the dedicated embedded features, transmitting them through a channel, and performing feature decoding at the decoding end to obtain reconstructed features; a modality imagination module divided into a forward modality imagination module and a backward modality imagination module for imagining discarded modality information based on existing modality information; the forward modality imagination module, taking the reconstructed features as input, obtains forward discarded modality features and joint multimodal features; the backward modality imagination module, taking the forward discarded modality features as input, obtains backward discarded modality features; and a classifier module, taking the joint multimodal features as input, obtains a probability distribution result for performing multimodal task sentiment recognition. Therefore, this embodiment of the invention can employ both compressed transmission of intermediate features and discarded modality, sacrificing data-level fidelity for semantic-level fidelity. This invention reduces bandwidth load by compressing and transmitting intermediate features; it also reduces bandwidth load during data transmission by discarding some modal information at the encoding end; and at the decoding end, it utilizes the correlation between modalities to recover the discarded modal information from other modalities, thus ensuring the performance of intelligent tasks. This invention employs two methods: compressing and transmitting intermediate features and discarding modalities. It leverages the correlation between multimodalities to reduce bandwidth load while ensuring the performance of intelligent tasks. Furthermore, discarded modal features are intermediate features and can still be used in other intelligent tasks that use those features as input.

[0138] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A machine-oriented multimodal cooperative coding device, characterized in that, include: The feature extraction module is used to extract corresponding exclusive embedded features from the native features of each modality; The feature encoding and decoding module is used to encode and compress the dedicated embedded features, and then perform feature decoding at the decoding end through channel transmission to obtain the reconstructed features. The modal imagination module is divided into a forward modal imagination module and a backward modal imagination module, which are used to imagine discarded modal information based on existing modal information. The forward modal imagination module takes the reconstructed features as input to obtain forward discarded modal features and joint multimodal features. The backward modal imagination module takes the forward discarded modal features as input to obtain backward discarded modal features. A classifier module is used to obtain probability distribution results by taking the joint multimodal features as input, so as to perform multimodal task emotion recognition; Each modality includes at least one or more of the following: audio modality, video modality, and text modality; Specifically, the modal imagination module is: The modal imagination module is located on the decoding end and is implemented through a CRA (Cascade Residual Autoencoder) structure; The CRA structure consists of multiple interconnected RA (Residual Autoencoder) blocks; the forward discarded modal features or backward discarded modal features are output by the last RA block; the joint multimodal features are obtained by extracting the intermediate features of each RA block and concatenating them. The joint loss function of the machine-oriented multimodal cooperative coding device specifically includes: The classification loss term uses cross-entropy as the loss function. The forward loss term aims to make the forward discarded modal features and the pre-trained proprietary embedding features as similar as possible, and uses mean squared error as the loss function. The backward loss term, similar to the forward loss term, aims to make the reconstructed features and the backward discarded mode features as similar as possible, and uses the mean squared error as the loss function. The formula for calculating the forward loss term is: , The formula for calculating the joint loss function is as follows: , In the formula, The forward loss term is defined as N, where N represents the number of feature values ​​in the forward discard modal features or the pre-trained proprietary embedding features. Represents the forward discard mode feature The i-th eigenvalue, Represents the pre-trained proprietary embedding features The i-th eigenvalue, , and These are hyperparameters used to control the classification loss term. The forward loss term and the backward loss term Weight allocation.

2. The machine-oriented multimodal cooperative coding device as described in claim 1, characterized in that, The feature extraction module specifically includes: The audio capture unit is used to obtain corresponding audio-specific embedded features through feature extraction processing, taking the original audio features as input. The video capture unit is used to take the original features of the video as input and obtain the corresponding video-specific embedded features through feature extraction processing. The text capture unit is used to take the original text features as input and obtain the corresponding text-specific embedding features through feature extraction processing. The feature splicing unit is used to splice the audio, video and text-specific embedding features to obtain specific embedding features containing full modality information.

3. The machine-oriented multimodal cooperative coding device as described in claim 1, characterized in that, The feature encoding / decoding module specifically includes: The feature encoding unit is used to quantize and tile the dedicated embedded features into a feature image, and to encode and compress the feature image using an image encoding and decoding method to obtain an encoded feature image. The feature decoding unit is used to decode the encoded feature image at the decoding end and obtain the reconstructed features after segmentation and dequantization.

4. The machine-oriented multimodal cooperative coding device as described in claim 1, characterized in that, The classifier module specifically includes: The classifier module consists of several fully connected layers. The joint multimodal features are passed through several fully connected layers to obtain output features, which are then processed... The function obtains the probability distribution results for multimodal emotion recognition tasks.

5. The machine-oriented multimodal cooperative coding device as described in claim 1, characterized in that, Also includes: The pre-trained feature extraction module is used to extract the corresponding pre-trained specific embedding features from the all-zero native features of each modality.

6. The machine-oriented multimodal cooperative coding device as described in claim 1, characterized in that, The native features of each modality are frame-level features, and the methods for capturing the native features of each modality specifically include: The original features of the audio modality are extracted to obtain the original audio features; The video is sampled, faces are detected and cropped, and the original features of the video are obtained based on the DenseNet model; The BERT-large model is used to extract the native features of the text modality to obtain the native text features.

7. A method for using a machine-oriented multimodal cooperative coding device, characterized in that, Applying the machine-oriented multimodal cooperative coding apparatus as described in any one of claims 1-6 to the testing phase includes: The native features of each modality are used as input to the machine-oriented multimodal co-coding device, and the corresponding exclusive embedded features are extracted from the native features of each modality through the feature extraction module; The specific embedded features are input into the feature encoding and decoding module for encoding and compression, and then transmitted through the channel to the decoding end for feature decoding to obtain the reconstructed features. Using the reconstructed features as input to the modal imagination module, discarded modal information is imagined based on existing modal information, thereby obtaining forward discarded modal features and joint multimodal features; The joint multimodal features are input into the classifier module to obtain the probability distribution results for multimodal task emotion recognition.

8. A method for using a machine-oriented multimodal cooperative coding device, characterized in that, Applying the machine-oriented multimodal cooperative coding apparatus as described in any one of claims 1-6 to the training phase includes: The native features and all-zero native features of each modality are used as inputs to the machine-oriented multimodal co-coding device. The feature extraction module extracts the corresponding exclusive embedding features from the native features, and the pre-trained feature extraction module obtains the corresponding pre-trained exclusive embedding features from the all-zero native features. The reconstructed features are obtained by uniformly processing the specific embedded features; Using the reconstructed features as input to the forward modality imagination module, forward discard modality features and joint multimodal features are obtained through the forward modality imagination module; using the forward discard modality features as input to the backward modality imagination module, backward discard modality features are obtained, thereby obtaining the backward loss; The joint multimodal features are input into the classifier module to obtain the classification loss; The forward loss is calculated based on the pre-trained dedicated embedding features and the forward discard modal features. The joint loss function of the machine-oriented multimodal co-coding device is obtained according to the calculation formula of the joint loss function, so as to optimize the training of the machine-oriented multimodal co-coding device.