Transform algorithm-based single-mode label generation and multi-mode emotion discrimination method
A single-modal and multi-modal technology, applied in character and pattern recognition, computing, computer parts, etc., can solve the problem of time-consuming and laborious manual labeling of single-modal labels, and achieve the effect of improving understanding and generalization ability
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Publication Date
- 2022-04-22
Smart Images

Figure 1 
Figure 2 
Figure 3
Abstract
Description
technical field
[0001] The present invention relates to time series one-dimensional convolutional neural network Conv1D, BiLSTM, Transformer self-attention mechanism and multimodal interactive attention mechanism, involves different fusion strategies of modalities, and realizes multimodal (voice, text, video) emotion Evaluation, and using a self-supervised mechanism with weighted voting, the prediction of single-modal labels and the final multi-modal emotional discrimination are realized, which belongs to the field of multi-modal multi-task emotional computing. Background technique
[0002] With the advent of the big data era, the data content is complicated and the data forms are also extremely rich. Human cognition of a certain event is a response to combining multiple modal information perceptions. It is difficult to fully interpret information using only a single modality, especially the judgment of human emotion. For example, the frowning person said to the escort rob...
Examples
Embodiment Construction
[0069] In this embodiment, a single-modal label generation and multi-modal emotion discrimination method based on the Transformer algorithm, the overall algorithm flow is as follows figure 1 As shown, the steps include: first obtain multi-modal non-aligned data sets, and perform preprocessing to obtain embedded expression features of corresponding modalities; then establish ITE network modules to extract intra-modal features; combine single-modal label prediction with multi-modal Fusion generation of modal emotion decision-making discriminant labels, establishment of inter-modal BTE network module and modal enhanced MTE network module, and acquisition of inter-modal features and modal enhancement features through the global self-attention STE network module to obtain multi-modal emotions The label of the deep prediction; finally, iterative training is carried out in combination with the design of the loss function. Specifically, it is characterized in that it proceeds in the f...