Method and system for a multimodal fusion model

The multimodal fusion system addresses synchronization issues in video description by generating content vectors from image, motion, and audio signals, improving description quality and efficiency.

DE112017006685B4Active Publication Date: 2025-06-18MITSUBISHI ELECTRIC CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE112017006685
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2017-03-29
Filing Date
2017-12-25
Publication Date
2025-06-18
Estimated Expiration
2037-12-25

AI Technical Summary

Technical Problem

Existing video description systems face challenges in synchronizing the sequence of video features with the sequence of words in the description, leading to irrelevant features and missing events, which affects the quality of the generated descriptions.

Method used

A multimodal fusion system that generates content vectors from input data comprising multiple modalities, including image, motion, and audio signals, using feature extractors, weight estimation, and attention mechanisms to synchronize and select relevant features for generating descriptive words.

Benefits of technology

The system improves the quality of video descriptions by selectively utilizing different features across modalities, reducing CPU usage and energy consumption while enhancing description accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A system for generating a word sequence from multimodal input vectors, comprising: one or more processors (120) in communication with a memory (140) and one or more storage devices (130) storing instructions executable when executed by the one or more processors to cause the one or more processors to perform operations comprising: Receiving (110, 118) first and second input vectors according to first and second consecutive intervals; Extracting (211-231) first and second feature vectors from the first and second input vectors using first and second feature extractors, respectively; estimating (212-232) a first set of weights and a second set of weights respectively from the first and second feature vectors and a pre-step context vector of a sequence generator; Calculating (213-233) a first content vector from the first set of weights and the first feature vectors, and calculating a second content vector from the second set of weights and the second feature vectors; Transforming (214-234) the first content vector into a first modal content vector having a predetermined dimension, and transforming the second content vector into a second modal content vector having the predetermined dimension; estimating (255) a set of modal attention weights from the pre-step context vector and the first and second content vectors or the first and second modal content vectors; Generating (245) a weighted content vector having the predetermined dimension from the set of modal attention weights and the first and second modal content vectors; and Generating (250) a predicted word using the sequence generator to generate the word sequence from the weighted content vector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical area

[0001] The invention relates generally to a method and system for describing multimodal data, and more particularly to a method and system for video description. Background to the state of the art

[0002] Automatic video description, known as video captioning, refers to the automatic generation of a natural language description (e.g., a sentence) that narrates an input video. Video description can apply to a wide range of applications, including video retrieval, automatic description of home movies or online uploaded video clips, video descriptions for the visually impaired, alert generation for surveillance systems, and scene understanding for human-machine knowledge sharing.

[0003] Video description systems extract the most relevant features from the video data, which can be multimodal features such as image features representing some objects, motion features representing some actions, and audio features indicating some events, and generate a description that narrates events such that the words in the description are relevant to these extracted features and arranged accordingly as natural language.

[0004] Non-patent literatures 1 and 2 concern possibilities for video description using a recurrent neural network and a multimodal fusion strategy, respectively.

[0005] An inherent problem with video description is that the sequence of video features and the sequence of words in the description are not synchronized. In fact, objects and actions in the video may appear in a different order than they appear in the sentence. When selecting the right words to describe something, only the features that directly correspond to that object or action are relevant, and the other features are a source of interference. Furthermore, some events are not always included in all features. REFERENCE LISTS NON-PATENT LITERATURE Non-patent literature 1: Haonan Yu et.al, Video Paragraph captioning using hierarchical recurrent neural networks. Proceedings, 29 th IEEE Conference on Computer Vision and Pattern Recognition, June 26 - July 1, 2016, Las Vegas, Nevada, Piscataway, NJ: IEEE 2016, pp. 4584-4593. ISBN 978-1-4673-8850-4. Non-patent literature 2: Shizhe Chen, Qin Jin, Multi-modal conditional attention fusion for dimensional emotion prediction. MM'16: Proceedings of the 2016 ACM Multimedia Conference, October 15-19, 2016, Amsterdam, The Netherlands. New York, NY: ACM, 2016, S571-575, ISBN 978-1-4503-3603-1. Summary of the inventionTechnical problem

[0006] Accordingly, there is a need to use different features globally or selectively to derive each word of the description to achieve high-quality video description. Solution to the problem

[0007] Some embodiments of the present disclosure are based on generating content vectors from input data having multiple modalities. In some cases, the modalities may be audio signals, video signals (image signals), and motion signals contained in video signals.

[0008] The present disclosure is based on a multimodal fusion system that generates content vectors from input data comprising multiple modalities. In some cases, the multimodal fusion system receives input signals comprising image (video) signals, motion signals, and audio signals, and generates a description that recounts events relevant to the input signals.

[0009] According to some embodiments of the present invention, a system for generating a word sequence from multimodal input vectors comprises one or more processors and one or more memory devices storing instructions executable when executed by the one or more processors to cause the one or more processors to perform operations including receiving first and second input vectors according to first and second consecutive intervals, extracting first and second feature vectors using first and second feature extractors, respectively, from the first and second inputs; estimating a first set of weights and a second set of weights, respectively, from the first and second feature vectors and a pre-step context vector of a sequence generator;Calculating a first content vector from the first set of weights and the first feature vectors, and calculating a second content vector from the second set of weights and the second feature vectors; transforming the first content vector into a first modal content vector having a predetermined dimension; and transforming the second content vector into a second modal content vector having the predetermined dimension; estimating a set of modal attention weights from the pre-step context vector and the first and second content vectors or the first and second modal content vectors; generating a weighted content vector having the predetermined dimension from the set of modal attention weights and the first and second modal content vectors; and generating a predicted word using the sequence generator to generate the word sequence from the weighted content vector.

[0010] Additionally, some embodiments of the present disclosure provide a non-transitory computer-readable medium storing software comprising instructions executable by one or more processors that, upon such execution, cause the one or more processors to perform operations. The operations include receiving first and second input vectors according to first and second consecutive intervals; extracting first and second feature vectors using first and second feature extractors, respectively, from the first and second input;Estimating a first set of weights and a second set of weights from the first and second feature vectors and a pre-step context vector of a sequence generator, respectively; calculating a first content vector from the first set of weights and the first feature vectors, and calculating a second content vector from the second set of weights and the second feature vectors; transforming the first content vector into a first modal content vector having a predetermined dimension, and transforming the second content vector into a second modal content vector having the predetermined dimension; estimating a set of modal attention weights from the pre-step context vector and the first and second content vectors or the first and second modal content vectors;Generating a weighted content vector having the predetermined dimension from the set of modal attention weights and the first and second modal content vectors; and generating a predicted word using the sequence generator to generate the word sequence from the weighted content vector.

[0011] According to another embodiment of the present disclosure, a method for generating a word sequence from multimodal input vectors comprises receiving first and second input vectors according to first and second consecutive intervals; extracting first and second feature vectors using first and second feature extractors, respectively, from the first and second input; estimating a first set of weights and a second set of weights, respectively, from the first and second feature vectors and a pre-step context vector of a sequence generator; calculating a first content vector from the first set of weights and the first feature vectors, and calculating a second content vector from the second set of weights and the second feature vectors;Transforming the first content vector into a first modal content vector having a predetermined dimension, and transforming the second content vector into a second modal content vector having the predetermined dimension; estimating a set of modal attention weights from the pre-step context vector and the first and second content vectors or the first and second modal content vectors; generating a weighted content vector having the predetermined dimension from the set of modal attention weights and the first and second modal content vectors; and generating a predicted word using the sequence generator to generate the word sequence from the weighted content vector.

[0012] The presently disclosed embodiments are further explained below with reference to the accompanying drawings. The illustrated drawings are not necessarily to scale, but rather are generally intended to illustrate the principles of the presently disclosed embodiments. Brief description of the drawings [ Fig. 1] Fig. 1 is a block diagram illustrating a multimodal fusion system according to some embodiments of the present disclosure. [ Fig. 2A] Fig. 2A is a block diagram illustrating a simple multimodal method according to embodiments of the present disclosure. [ Fig. 28] Fig. 2B is a block diagram illustrating a multimodal attention method according to embodiments of the present disclosure. [ Fig. 3] Fig. 3 is a block diagram illustrating an example of the LSTM-based decoder-encoder architecture according to embodiments of the present disclosure. [ Fig. 4] Fig. 4 is a block diagram illustrating an example of the attention-based sentence generator from video according to embodiments of the present disclosure. [ Fig. 5] Fig. 5 is a block diagram illustrating an extension of the attention-based sentence generator from video according to embodiments of the present disclosure. [ Fig. 6] Fig. 6 is a diagram illustrating a simple fusion approach (simple multimodal method) according to embodiments of the present disclosure. [ Fig. 7] Fig. 7 is a diagram illustrating an architecture of a sequence generator according to embodiments of the present disclosure. [ Fig. 8] Fig. 8 shows comparisons of performance results obtained by conventional methods and the multimodal attention method according to embodiments of the present disclosure. [ Fig. 9A to 9D] Fig. 9A, Fig. 9B and Fig. 9C show comparisons of performance results obtained by conventional methods and the multimodal attention method according to embodiments of the present disclosure. Description of the embodiments

[0013] While the above drawings describe presently disclosed embodiments, other embodiments are also contemplated, as noted in the discussion. The present disclosure presents illustrative embodiments by way of illustration and not limitation. Numerous other modifications and embodiments that fall within the scope and spirit of the principles of the presently disclosed embodiments may be made by those skilled in the art.

[0014] The following description provides only exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the disclosure. Rather, the following description of the exemplary embodiments is intended to provide those skilled in the art with a description enabling them to implement one or more of the exemplary embodiments. Various changes that may be made in the function and arrangement of elements are contemplated without departing from the spirit and scope of the disclosed subject matter as set forth in the appended claims.

[0015] Specific details are provided in the following description to provide a thorough understanding of the embodiments. However, it should be understood by one of ordinary skill in the art that the embodiments may be practiced without these specific details. For example, systems, processes, and other elements in the disclosed subject matter may be presented as components in block diagram form in order not to obscure the embodiments with unnecessary detail. In other instances, well-known processes, structures, and techniques may be presented without unnecessary detail in order not to obscure the embodiments. In addition, like reference numerals and labels indicate like elements throughout the various drawings.

[0016] Additionally, individual embodiments may be described as a process represented as a flowchart, sequence diagram, data flow diagram, structure diagram, or block diagram. Although a flowchart may describe the operations as a sequential process, many of the operations may be performed in parallel or concurrently. Furthermore, the order of the operations may be rearranged. A process may terminate when its operations are complete, but may have additional steps not explained or included in a figure. Furthermore, not all operations in a specifically explained process may occur in all embodiments. A process may correspond to a method, function, act, subroutine, subprogram, etc.If a process corresponds to a function, the termination of the function can correspond to the function returning to the calling function or the main function.

[0017] Furthermore, embodiments of the disclosed subject matter may be implemented, at least in part, either manually or automatically. Manual or automatic implementations may be performed, or at least supported, through the use of machines, hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segments for performing the necessary tasks may be stored in a machine-readable medium. A processor(s) may perform the necessary tasks.

[0018] According to embodiments of the present disclosure, a system for generating a word sequence from multimodal input vectors comprises one or more processors in communication with one or more memories and one or more storage devices storing instructions that are executable. When executed by the one or more processors, the instructions cause the one or more processors to perform operations including: receiving first and second input vectors according to first and second consecutive intervals; extracting first and second feature vectors using first and second feature extractors, respectively, from the first and second input;Estimating a first set of weights and a second set of weights from the first and second feature vectors and a pre-step context vector of a sequence generator, respectively; calculating a first content vector from the first weight and the first feature vector, and calculating a second content vector from the second weight and the second feature vector; transforming the first content vector into a first modal content vector having a predetermined dimension and transforming the second content vector into a second modal content vector having the predetermined dimension; estimating a set of modal attention weights from the pre-step context vector and the first and second modal content vectors; generating a weighted content vector having the predetermined dimension from the set of modal attention weights and the first and second content vectors;and generating a predicted word using the sequence generator to generate the word sequence from the weighted content vector;

[0019] In this case, the first modal content vector, the second modal content vector, and the weighted content vector can have the same predetermined dimension. This enables the system to implement a multimodal fusion model. In other words, by designing or determining the dimensions of the input vectors and the weighted content vectors to have an identical dimension, these vectors can be easily handled in the data processing of the multimodal fusion model, since these vectors are expressed using an identical data format having the same dimension.By simplifying data processing using transformed data to have the identical dimension, the multimodal fusion model method or system according to embodiments of the present disclosure can reduce the usage of a central processing unit and the energy consumption for generating a string of words from the multimodal input vectors.

[0020] Of course, the number of vectors can be changed to a predetermined number of N vectors according to the system design requirements. For example, if the predetermined number of N is set to three, the three input vectors can be image features, motion features, and audio features received from image data, video signals, and audio signals via an input / output interface integrated into the system.

[0021] In some cases, the first and second consecutive intervals may be an identical interval, and the first and second vectors may be different modalities.

[0022] Fig. 1 shows a block diagram illustrating a multimodal fusion system 100 according to some embodiments of the present disclosure.The multimodal fusion system 100 may include a human-machine interface (HMI) with input / output (I / O) interface 110 connectable to a keyboard 111 and a pointing device / media 112, a microphone 113, a receiver 114, a transmitter 115, a 3D sensor 116, a global positioning system (GPS) 117, one or more I / O interfaces 118, a processor 120, a storage device 130, a memory 140, a network interface controller 150 (NIC) connectable to a network 155 including local area networks and an Internet network (not shown), a display interface 160 connected to a display device 165, a imaging interface 170 connectable to an imaging device 175, a printer interface 180, connectable to a printer device 185. The HMI with I / O interface 110 may include analog-to-digital and digital-to-analog converters.The HMI with I / O interface 110 includes a wireless communication interface that can communicate with other 3D point cloud display systems or other computers via wireless internet connections or wireless local area networks, enabling the construction of multiple 3D point clouds. The 3D point cloud system 100 may include a power source 190. The power source 190 may be a battery that can be recharged from an external power source (not shown) via the I / O interface 118. Depending on the application, the power source 190 may optionally be located external to the system 100.

[0023] The HMI and I / O interface 110 and the I / O interface 118 may be adapted to connect to another display device (not shown), including a computer monitor, a camera, a television, a projector, or a mobile device, among others.

[0024] The multimodal fusion system 100 can receive electronic text / image documents 195 comprising speech data via the network 155 connected to the NIC 150. The storage device 130 includes a sequence generation model 131, a feature extraction model 132, and a multimodal fusion model 200, in which algorithms of the sequence generation model 131, the feature extraction model 132, and the multimodal fusion model 200 are stored as program code data in the memory 130. The algorithms of the models 131-132 and 200 can be stored on a computer-readable recording medium (not shown) so that the processor 120 can execute the algorithms of the models 131-132 and 200 by loading the algorithms from the medium. In addition, the pointing device / medium 112 may include modules that read and execute programs stored on a computer-readable recording medium.

[0025] To begin execution of the algorithms of Models 131-132 and 200, instructions may be transmitted to System 100 via keyboard 111, pointing device / media 112, or via the wireless network or network 155 connected to other computers (not shown). The algorithms of Models 131-132 and 200 may be initiated in response to receiving an audible signal from a user through microphone 113 using a pre-installed conventional speech recognition program stored in memory 130. System 100 further includes an on / off switch (not shown) that allows the user to start / stop operation of System 100.

[0026] The HMI and I / O interface 110 may include an analog-to-digital (A / D) converter, a digital-to-analog (D / A) converter, and a wireless signal antenna for connecting to the network 155. Furthermore, the one or more I / O interfaces 118 may be connected to a cable television (TV) network or a conventional television (TV) antenna that receives television signals. The signals received via the interface 118 may be converted into digital image and audio signals, which may be processed according to the algorithms of models 131-132 and 200 in conjunction with the processor 120 and memory 140 to generate video scripts and display them on the display device 165 with image frames of the digital images, while the audio of the TV signals is output via a speaker 119.The speaker may be integrated into the system 100, or an external speaker may be connected via the interface 110 or the I / O interface 118.

[0027] Processor 120 may be a plurality of processors including one or more graphics processing units (GPUs). Memory 130 may include speech recognition algorithms (not shown) that can recognize speech signals received via microphone 113.

[0028] The multimodal fusion system module 200, the sequence generation model 131 and the feature extraction model 132 may be formed by neural networks.

[0029] Fig. 2A is a block diagram illustrating a simple multimodal method according to embodiments of the present disclosure. The simple multimodal method may be performed by the processor 120 executing programs of the sequence generation model 131, the feature extraction model 132, and the multimodal fusion model 200 stored in memory. The sequence generation model 131, the feature extraction model 132, and the multimodal fusion model 200 may be stored in a computer-readable recording medium such that the simple multimodal method may be performed when the processor 120 loads and executes the algorithms of the sequence generation model 131, the feature extraction model 132, and the multimodal fusion model 200. The simple multimodal method is performed in combination with the sequence generation model 131, the feature extraction model 132, and the multimodal fusion model 200.Furthermore, the simple multimodal method uses the feature extractors 211, 221 and 231 (feature extractors 1~K), the attention estimators 212, 222 and 232 (attention estimators 1~K), the weighted sum processors 213, 223 and 233 (weighted sum processors (calculators) 1~K), the feature transformation modules 214, 224 and 234 (feature transformation modules 1~K), a simple sum processor (calculators) 240 and a sequence generator 250.

[0030] Fig. 2B is a block diagram illustrating a multimodal attention method according to embodiments of the present disclosure. In addition to the feature extractors 1~K, the attention estimators 1~K, the weighted sum processors 1~K, the feature transformation modules 1~K, and the sequence generator 250, the multimodal attention method further includes a modal attention estimator 255 and a weighted sum processor 245 instead of using the simple sum processor 240. The multimodal attention method is performed in combination with the sequence generation model 131, the feature extraction model 132, and the multimodal fusion model 200. In both methods, the sequence generation model 131 provides the sequence generator 250, and the feature extraction model 132 provides the feature extractors 1~K.Furthermore, the feature transformation modules 1~K, the modal attention estimator 255 and the weighted sum processors 1~K and the weighted sum processor 245 may be provided by the multimodal fusion model 200.

[0031] Given multimodal video data having K modalities such that K ≥ 2, and some of the modalities may be the same, modal 1 data is converted into a fixed-dimensional content vector using the feature extractor 211, the attention estimator 212, and the weighted sum processor 213 for the data, where the feature extractor 211 extracts multiple feature vectors from the data, the attention estimator 212 estimates each weight for each extracted feature vector, and the weighted sum processor 213 outputs (generates) the content vector calculated as the weighted sum of the extracted feature vectors with the estimated weights. Modal 2 data is converted into a fixed-dimensional content vector using the feature extractor 221, the attention estimator 222, and the weighted sum processor 223 for the data.Up to the modal-K data, K fixed-dimensional content vectors are obtained using the feature extractor 231, the attention estimator 232, and the weighted sum processor 233 for modal-K data. The modal-1, modal-2, ...., modal-K data may each be sequential data in a time-sequential order with an interval or other predetermined orders with predetermined time intervals.

[0032] Each of the K content vectors is then transformed (converted) into an N-dimensional vector by each feature transformation module 214, 224, and 234, and K transformed N-dimensional vectors are obtained, where N is a predefined positive integer.

[0033] The K transformed N-dimensional vectors are transformed in the simple multimodal method of Fig. 2A into a single N-dimensional content vector, while the vectors are combined using the modal attention estimator 255 and the weighted sum processor 245 in the multimodal attention method of Fig. 2B into a single N-dimensional content vector, wherein the modal attention estimator 255 estimates each weight for each transformed N-dimensional vector, and the weighted sum processor 245 outputs (generates) the N-dimensional content vector calculated as a weighted sum of the K transformed N-dimensional vectors with the estimated weights.

[0034] The sequence generator 250 receives the single N-dimensional content vector and predicts a label corresponding to a word of a sentence describing the video data. To predict the next word, the sequence generator 250 provides context information of the sentence, such as a vector representing the previously generated words, to the attention estimators 212, 222, 232 and the modal attention estimator 255 to estimate the attention weights to obtain appropriate content vectors. This vector may be referred to as a pre-step (or pre-step) context vector.

[0035] The sequence generator 250 says the next word starting with the sentence start token “ <sos>" and generates a descriptive sentence or sentences by iteratively predicting the next word (predicted word) until a special symbol " <eos>" corresponding to the "end of sentence". In other words, the sequence generator 250 generates a word sequence from multimodal input vectors. In some cases, the multimodal input vectors may be received via various input / output interfaces, such as the HMI and I / O interface 110, or one or more I / O interfaces 118.

[0036] In each generation process, a predicted word is generated that has the highest probability among all possible words, given by the weighted content vector and the pre-step context vector. Furthermore, the predicted word may be accumulated in memory 140, storage device 130, or multiple storage devices (not shown) for generating the word sequence, and this accumulation process may continue until the specific symbol (end of the sequence) is received. System 100 may transmit the predicted words generated by sequence generator 250 via NIC 150 and network 155, HMI and I / O interface 110, or one or more I / O interfaces 118 so that the predicted word data can be used by other computers 195 or other output devices (not shown).

[0037] If each of the K content vectors originates from specific modality data and / or from a specific feature extractor, modality or feature fusion with the weighted sum of the K transformed vectors enables better prediction of each word by paying attention to different modalities and / or different features according to the context information of the sentence. Thus, this multimodal attention method can utilize different features globally or selectively using attention weights across different modalities or features to derive each word of the description.

[0038] Furthermore, the multimodal fusion model 200 in the system 100 includes a data distribution module (not shown) that receives a plurality of time-sequential data via the I / O interface 110 or 118 and distributes the received data into Modal-1, Modal-2,..., Modal-K data, divides all of the distributed time-sequential data according to a predetermined interval or intervals, and then supplies the Modal-1, Modal-2,..., Modal-K data to the feature extractors 1~K, respectively.

[0039] In some cases, the multiple time-sequential data may be video signals and audio signals contained in a video clip. When the video clip is used for modal data, the system 100 uses the feature extractors 211, 221, and 231 (set K=3) in Fig. 2B. The video clip is provided to the feature extractors 211, 221 and 231 in the system 100 via the I / O interface 110 or 118. The feature extractors 211, 221 and 231 can extract image data, audio data and motion data respectively from the video clip as modal 1 data, modal 2 data and modal 3 data (e.g., K=3 in Fig. 2B). In this case, the feature extractors 211, 221, and 231 receive modal 1 data, modal 2 data, and modal 3 data according to the first, second, and third intervals, respectively, from the data stream of the video clip.

[0040] In some cases, the data distribution module may divide the multiple time-sequential data with predetermined different time intervals if image features, motion features, or audio features with different time intervals can be respectively acquired. Encoder-decoder-based sentence generator

[0041] One approach to video description can be based on sequence-by-sequence learning. The input sequence, i.e., the image sequence, is first encoded into a fixed-dimensional semantic vector. The output sequence, i.e., the word sequence, is then generated from the semantic vector. In this case, both the encoder and decoder (or generator) are typically modeled as long-short-term memory (LSTM) networks.

[0042] Fig. Figure 3 shows an example of the LSTM-based encoder-decoder architecture. For a given sequence of images, X = x1, x2, ...., x L , each image is first fed to a feature extractor, which can be a pre-trained convolutional neural network (CNN) for an image or video classification task, such as GoogLeNet, VGGNet, or C3D. The sequence of image features, X' = x'1, x'2, ...., x' L , is obtained by extracting the activation vector of a fully connected layer of the CNN for each input image. The sequence of feature vectors is then fed to the LSTM encoder, and the hidden state of the LSTM is given by h1=LSTM(ht−1,xt';λE), where the LSTM function of the encoder network λ E is calculated as LSTM(ht−1,xt;λ)=ot tanh(ct), where ot=σ(Wxo(λ)xt+Who(λ)ht−1+bo(λ)) ct=ftct−1+it tanh(Wxc(λ)xt +Whc(λ)ht−1+bc(λ)) ft=σ(Wxf(λ)xt+Whf(λ)ht−1+bf(λ)) it=σ(Wxi(λ)xt+Whi(λ)ht−1+bi(λ)), where σ() is the element-wise sigmoid function, and i t , f t , to and c t are the input gate, forget gate, output gate, and cell activation vectors for the t-th input vector, respectively. The weight matrices W zz (λ) and the bias vectors b Z (λ) are identified by the index z ∈ {x, h, i, f, , o, c}. For example, W hi the Hidden Gate Matrix and W xo is the input-output gate matrix. Peephole connections are not used in this method.

[0043] The decoder predicts the next word iteratively, starting with the sentence start token “ <sos>", until he finishes the sentence " <eos>" predicts. The sentence start token may be referred to as a start tag, and the sentence end token may be referred to as an end tag.

[0044] For a given decoder state s i-1 , the decoder network λ D the probability distribution of the next word as P(y|si−1)=softmax(Ws(λD)si−1+bs(λD)), and produces word y i , which has the highest probability, according to yi=argmaxy∈VP(y|si−1), where V denotes the vocabulary. The decoder state is updated using the decoder's LSTM network as si=LSTM(si−1,yi';λD), where y' i a word embedding vector of y m and the output state s0 from the final encoder state h L and y'0 = Embed( <sos>) as in Fig. 3 is obtained.

[0045] In the training phase, Y = y1, ..., y M as the reference. In the test phase, however, the best word sequence is found based on Y^=argmaxY∈V* P(Y|X)=argmaxy1,…,yM∈V*P(y1|s0)P(y2|s1)⋯ P(yM|sM−1)P( <eos>|sM).

[0046] Accordingly, a beam search can be used in the testing phase to obtain multiple states and hypotheses with the highest cumulative probabilities at each m-th step and select the best hypothesis from those that have reached the sentence end token. Attention-based sentence generator

[0047] Another approach for video description can be an attention-based sequence generator, which allows the network to emphasize features from specific times or spatial regions depending on the current context, allowing the next word to be predicted more accurately. Compared to the basic approach described above, the attention-based generator can selectively exploit input features according to the input and output context. The effectiveness of attention models has been demonstrated in many tasks, such as machine translation.

[0048] Fig. Figure 4 is a block diagram illustrating an example of an attention-based sentence generator from video that includes a temporal attention mechanism over the input image sequence. The input image sequence may be a time-sequential sequence with predetermined time intervals. The input sequence of feature vectors is obtained using one or more feature extractors. In this case, attention-based generators may use an encoder based on a bidirectional LSTM (BLSTM) or gated recurrent unit (GRU) to extract the feature vector sequence, as shown in Fig. 5, to further convert so that each vector contains its context information.

[0049] However, in video description tasks, CNN-based features can be used directly, or another feed-forward layer can be added to reduce dimensionality.

[0050] After feature extraction as in Fig. 5 a BLSTM encoder is used, the activation vectors (i.e. encoder states) can be obtained as ht=[ht(f)ht(b)], where h t (f) and h t (b) the hidden forward and backward activation vectors are: ht(f)=LSTM(ht−1(f),xt';λE(f)) ht(b)=LSTM(ht+1(b),xt';λE(b)),

[0051] When a feedforward layer is used, the activation vector is calculated as ht=tanh(Wpxt'+bp), where W p is a weighting matrix and b p is a bias vector. Furthermore, if the CNN features are used directly, then this is assumed to be h t = x t to be.

[0052] The attention mechanism is implemented by applying attention weights to the hidden activation vectors throughout the input sequence. These weights allow the network to highlight features from the time steps that are most important for predicting the next output word.

[0053] It is assumed that α i,t an attention weighting between the i ten Output word and the t ten Input feature vector. For the i te The output is the vector representing the relevant content of the input sequence, obtained as a weighted sum of the activation vectors of the hidden unit: ci=∑t=1Lαi,tht.

[0054] The decoder network is an attention-based recurrent sequence generator (ARSG) that generates an output label sequence with content vectors c i The network also has an LSTM decoder network in which the decoder state can be updated in the same way as in equation (9).

[0055] Then the probability of the output label is calculated as P(y|si−1,ci)=softmax(Ws(λD)si−1+Wc(λD)ci+bs(λD)), and the word y i is calculated according to yi=argmaxy∈VP(y|si−1,ci).

[0056] In contrast to equations (7) and (8) of the basic encoder-decoder, the probability distribution depends on the content vector ci, which highlights specific features that are most important for predicting each subsequent word. Another feedforward layer can be inserted before the softmax layer. In this case, the probabilities are calculated as follows: gi=tanh(ws(λD)si−1+Wc(λD)ci+bs(λD)), and P(y|si−1,ci)=softmax(Wg(λD)gi+bg(λD)),

[0057] The attention weights can be calculated as αi,t=exp(ei,t)∑τ=1Lexp(ei,τ) and ei,t=wAT tanh(WAsi−1+VAht+bA), where W A and V A Matrices are, w A and b A vectors are, and e i,t is a scalar. Attention-based multimodal fusion

[0058] Embodiments of the present disclosure provide an attention model for handling multimodal fusion, where each modality has its own sequence of feature vectors. Multimodal inputs, such as image features, motion features, and audio features, are available for video description. Furthermore, combining multiple features from different feature extraction methods is often effective for improving description accuracy.

[0059] In some cases, content vectors from VGGNet (image features) and C3D (spatiotemporal motion features) can be combined into a single vector used to predict the next word. This can be done in the fusion layer. Assuming K is the number of modalities, i.e., the number of sequences of input feature vectors, the following activation vector is calculated instead of Equation (19): gi=tanh(Ws(λD)si−1+∑k=1Kdk,i+bs(λD)) where dk,i=Wck(λD)ck,i and c k,i is the k-th content vector corresponding to the k-th feature extractor or modality.

[0060] Fig. Figure 6 shows the simple feature fusion approach (simple multimodal method) assuming K=2, where content vectors with attention weights for individual input sequences x 11 ,..., x 1L and x 21 τ ,..., x 2L τ However, these content vectors are weighted with weight matrices W c1 and W c2 that are typically used in the sentence generation step. Consequently, the content vectors of each feature type (or modality) are always fused using the same weights, regardless of the decoder state. This architecture can introduce the ability to effectively exploit multiple types of features, allowing the relative weights of each feature type (modality) to change depending on the context.

[0061] According to embodiments of the present disclosure, the attention mechanism can be extended to multimodal fusion. Using the multimodal attention mechanism based on the current decoder state, the decoder network can selectively focus on specific input modalities (or specific feature types) to predict the next word. The attention-based feature fusion according to embodiments of the present disclosure can be performed using gi=tanh(Ws(λD)si−1+∑k=1Kβk,idk,i+bs(λD)), where dk,i=Wck(λD)ck,i+bck(λD).

[0062] The multimodal attention weights β k,i are obtained in a similar way to the temporal attention mechanism: βk,i=exp(vk,i)∑κ=1Kexp(vκ,i) where vk,i=wBT tanh(WBsi−1+VBkck,i+bBk), where W B and V Bk Matrices are, w B and b Bk are vectors, and v k,i is a scalar.

[0063] Fig. Figure 7 shows the architecture of the sentence generator according to embodiments of the present disclosure, including the multimodal attention mechanism. In contrast to the simple multimodal fusion method in Fig. 6, can be found in Fig. 7 change the attention weights at the feature level according to the decoder state and the content vectors, allowing the decoder network to pay attention to a different set of features and / or modalities when predicting each subsequent word in the description. Data set for evaluation

[0064] Some experimental results are described below to illustrate feature fusion according to an embodiment of the present disclosure using the YouTube2Text video corpus. This corpus is well suited for training and evaluating automatic models for generating video descriptions. The dataset contains 1,970 video clips with descriptions in multiple natural languages. Each video clip is annotated with several parallel sentences provided by different Mechanical Turkers. There are a total of 80,839 sentences, with approximately 41 annotated sentences per clip. Each sentence contains an average of approximately 8 words. The words contained in all sentences form a vocabulary of 13,010 unique lexical entries. The dataset is open domain and covers a wide range of topics, such as sports, animals, and music. The dataset is divided into a training set with 1.200 video clips, a validation set with 100 clips and a test set with the remaining 670 clips. Video preprocessing

[0065] Image data is extracted from each video clip, which consists of 24 frames per second and is rescaled to 224x224 pixel images. For image feature extraction, a pre-trained GoogLeNet CNN (M. Lin, Q. Chen, and S. Yan. Network within a network. CoRR, abs / 1312.4400, 2013) is used to extract fixed-length representations using the popular implementation in Caffe (Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell. Caffe: Convolutional architecture for fast feature embedding. arXiv Preprint arXiv:1408.5093, 2014). Features are extracted from the 5 / 7x7 s1 hidden layer pool. A frame is selected from each video clip every 16 frames and fed to the CNN to obtain 1024-dimensional, frame-wise feature vectors.

[0066] A VGGNet (K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. CoRR, abs / 1409.1556, 2014) pre-trained on the ImageNet dataset (A. Krizhevsky, I. Sutskever, and G.E. Hinton. Imagenet classification with deep convolutional neural networks. In F. Pereira, C.J.C. Burges, L. Bottou, and K.Q. Weinberger, editors, Advances in Neural Information Processing Systems 25, pages 1097–1105. Curran Associates, Inc., 2012) is also used. For image features, the hidden activation vectors of the fully connected layer fc7 are used, resulting in a sequence of 4096-dimensional feature vectors. In addition, the pre-trained C3D (D. Tran, LD Bourdev, R. Fergus, L. Torresani and M. Paluri. Learning spatiotemporal features with 3d convultional networks.In 2015 IEEE International Conference on Computer Vision, ICCV 2015, Santiago, Chile, December 7-13, 2015, pages 4489-4497, 2015.) was used (which was trained on the Sports-1M dataset (A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei). Large-scale classification with convolutional neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1725-1732, 2014.) was trained). The C3D network reads consecutive frames in the video and outputs a fixed-length feature vector after every 16 frames. The activation vectors were extracted from the fully connected layer fc6-1, which has 4096-dimensional features. Audio processing

[0067] Audio features are integrated for use in the attention-based feature fusion method according to embodiments of the present disclosure. Since the YouTube2Text corpus does not contain an audio track, the audio data was extracted via the original video URLs. Although a subset of the videos was no longer available on YouTube, the audio data for 1,649 video clips, representing 84% of the corpus, could be collected. The 44 kHz sampled audio data is downsampled to 16 kHz, and mel-frequency cepstral coefficients (MFCCs) are extracted from each 50 ms time window with a 25 ms shift. The sequence of 13-dimensional MFCC features is then concatenated into a vector from each group of 20 consecutive frames, resulting in a sequence of 260-dimensional vectors. The MFCC features are normalized so that the mean and variance vectors are 0 and 1 in the training set.The validation and test sets are also fitted with the original mean and variance vectors of the training set. Unlike image features, MFCC features use a BLSTM encoder network, which is trained jointly with the decoder network. If audio data is missing for a video clip, a sequence of dummy MFCC features, which is simply a sequence of zero vectors, is fed in. Configuration for describing multimodal data

[0068] The subtitling generation model, i.e., the decoder network, is trained to minimize the cross-entropy criterion using the training set. Image features are fed to the decoder network through a 512-unit projection layer, while audio features, i.e., MFCCs, are fed to the BLSTM encoder followed by the decoder network. The decoder network comprises a 512-unit projection layer and 512-cell bidirectional LSTM layers. The decoder network comprises a 512-cell LSTM layer. Each word is embedded in a 256-dimensional vector when fed to the LSTM layer. The AdaDelta optimizer (MD: Zeiler. ADADELTA: an adaptive learning rate method. CoRR, abs / 1212.5701, 2012) is used to update the parameters, which is widely used for optimizing attention models. The LSTM and attention models are optimized using Chainer (S. Tokui, K. Oono, S.Hido, and J. Clayton. Chainer: a next generation open source framework for deep learning. Implemented in the Workshop Proceedings on Machine Learning Systems (Learn-7 ingSys) in the Twenty-ninth Annual Conference on Neural Information Processing Systems (NIPS), 2015.

[0069] The similarity between ground truth and automatic video description results is evaluated using machine translation-motivated metrics: BLEU (K. Papineni, S. Roukos, T. Ward, and W. Zhu. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, July 6-12, 2002, Philadelphia, PA, USA., pages 311-318, 2002.), METEOR (M.J. Denkowski and A. Lavie. Meteor universal: Language-specific translation evaluation for any target language. In Proceedings of the Ninth Workshop on Statistical Machine Translation, WMT@ACL 2014, June 26-27, 2014, Baltimore, Maryland, USA, pages 376-380, 2014.), and the other metric for image description, CIDEr (R. Vedantam, C.L. Zitnick, and D. Parikh. Cider: Consensus-based image description evaluation. In IEEE Conference on Computer Vision and Pattern Recognition CVPR 2015, Boston, MA, USA, June 7-12, 2015, pages 4566-4575, 2015.The publicly available evaluation script prepared for the image captioning challenge was used (X. Chen, H. Fang, T. Lin, R. Vedantam, S. Gupta, P. Doll'ar, and CL Zitnick. Microsoft COCO captioning: Data collection and evaluation server. CoRR, abs / 1504.00325, 2015.). Evaluation results

[0070] Fig. Figure 8 shows comparisons of performance results obtained by conventional methods and the multimodal attention method according to embodiments of the present disclosure on the YouTube2text dataset. The conventional methods, which are simple additive multimodal fusion (Simple Multimodal), unimodal models with temporal attention (Unimodal), and baseline systems using temporal attention, are performed.

[0071] The first three rows of the table use temporal attention but only one modality (one feature type). The next two rows perform a multimodal fusion of two modalities (image and space-time), using either Simple Multimodal Fusion (see Fig. 6) or the proposed multimodal attention mechanism (see Fig. 7). The next two rows also perform multimodal fusion, this time using three modalities (image, spatiotemporal, and audio features). In each column, the results of the two best methods are shown in bold.

[0072] The Simple Multimodal Model performed better than the unimodal models. However, the Multimodal Attention Model outperformed the Simple Multimodal Model. The audio feature degrades the baseline performance because some YouTube data contains noise, such as background music, that is unrelated to the video content. The Multimodal Attention Model mitigated the effects of the noise on the audio features. Furthermore, combining the audio features using the proposed method achieved the best performance of CIDEr for all experimental conditions.

[0073] However, the Multimodal Attention Model improved the Simple Multimodal Model.

[0074] Fig. 9A, Fig. 9B, Fig. 9C and Fig. 9D show comparisons of performance results obtained by conventional methods and the multimodal attention method according to embodiments of the present disclosure.

[0075] The Fig. 9A-9C show three exemplary video clips for which the attention-based multimodal fusion procedure (Temporal & Multimodal Attention with VGG and C3D) outperformed the single modal procedure (Temporal Attention with VGG) and the single modal fusion procedure (Temporal Attention with VGG and C3D) on the CIDEr measure. Fig. Figure 9D shows an exemplary video clip for which the attention-based multimodal fusion method (Temporal & Multimodal Attention) with audio features outperformed the unimodal method (Temporal Attention with VGG) and the single-modal fusion method (Temporal Attention with VGG, C3D) with / without audio features. These examples demonstrate the effectiveness of the multimodal attention mechanism.

[0076] In some embodiments of the present disclosure, when the multimodal fusion model described above is installed in a computer system, the video script can be effectively generated with less computing power, so that the use of the multimodal fusion model method or system can reduce the use of central processing units and energy consumption.

[0077] Furthermore, embodiments according to the present disclosure provide an effective method for performing the multimodal fusion model, such that use of a method and system using the multimodal fusion model can reduce central processing unit (CPU) utilization, power consumption, and / or utilized network bandwidth.

[0078] The above-described embodiments of the present disclosure may be implemented in a variety of ways. For example, the embodiments may be realized using hardware, software, or a combination thereof. When implemented in software, the software code may execute on any suitable processor or collection of processors, whether provided in a single computer or distributed across multiple computers. Such processors may be implemented as integrated circuits with one or more processors in an integrated circuit component. However, a processor may also be implemented using circuitry in any suitable format.

[0079] Furthermore, the various methods or processes described herein may be encoded as software executable on one or more processors using any of a variety of operating systems or platforms. Furthermore, this software may be written using any suitable programming language and / or programming or scripting tools, and may also be compiled as executable machine language code or intermediate code running on a framework or virtual machine. Typically, the functionality of the program modules may be arbitrarily combined or distributed in various embodiments.

[0080] Furthermore, the embodiments of the present disclosure may be embodied as a method, for which an example has been provided. The acts performed as part of the method may be arranged in any suitable manner. Accordingly, embodiments may be constructed in which acts are performed in a different order than that illustrated, which may include performing some acts simultaneously, even if they are illustrated as sequential acts in illustrative embodiments.Furthermore, the use of ordering terms such as first, second, in the claims to modify a claim element does not, by itself, imply priority, precedence, or ranking of one claim element over another, or the chronological order in which acts of a method are performed, but serves merely as a label to distinguish one claim element with a particular designation from another element with a like designation (except for the use of the ordering term) to distinguish claim elements from one another.< / eos> < / sos> < / eos> < / sos> < / eos> < / sos>

Claims

[1] System for generating a word sequence from multimodal input vectors, comprising: one or more processors (120) in communication with a memory (140) and one or more storage devices (130) storing instructions executable when executed by the one or more processors to cause the one or more processors to perform operations comprising: Receiving (110, 118) first and second input vectors according to first and second consecutive intervals; Extracting (211-231) first and second feature vectors from the first and second input vectors using first and second feature extractors, respectively; estimating (212-232) a first set of weights and a second set of weights respectively from the first and second feature vectors and a pre-step context vector of a sequence generator; Calculating (213-233) a first content vector from the first set of weights and the first feature vectors, and calculating a second content vector from the second set of weights and the second feature vectors; Transforming (214-234) the first content vector into a first modal content vector having a predetermined dimension, and transforming the second content vector into a second modal content vector having the predetermined dimension; estimating (255) a set of modal attention weights from the pre-step context vector and the first and second content vectors or the first and second modal content vectors; Generating (245) a weighted content vector having the predetermined dimension from the set of modal attention weights and the first and second modal content vectors; and Generating (250) a predicted word using the sequence generator to generate the word sequence from the weighted content vector. [2] The system of claim 1, wherein the first and second consecutive intervals are an identical interval. [3] The system of claim 1, wherein the first and second input vectors are different modalities. [4] The system of claim 1, wherein the operations further comprise: Accumulating the predicted word in the memory (140) or the one or more storage devices (130) to generate the word sequence. [5] The system of claim 4, wherein the accumulation continues until an end flag is received. [6] The system of claim 1, wherein the operations further comprise: Transmitting the predicted word generated from the sequence generator (250). [7] The system of claim 1, wherein the first and second feature extractors (211-231) are pre-trained convolutional neural networks (CNNs) trained for an image or video classification task. [8] The system of claim 1, wherein the first and second feature extractors (211-231) are long short-term memory (LSTM) networks. [9] The system of claim 1, wherein the predicted word with the highest probability is determined among all possible words given the weighted content vector and pre-step context vector. [10] The system of claim 1, wherein the sequence generator (250) employs a long-short-term memory (LSTM) network. [11] The system of claim 1, wherein the first input vector is received via a first input / output (I / O) interface and the second input vector is received via a second I / O interface. [12] A non-transitory computer-readable medium storing software containing instructions executable by one or more processors (120) which, when so executed, cause the one or more processors in conjunction with a memory (140) to perform operations comprising: Receiving (110, 118) first and second input vectors according to first and second consecutive intervals; Extracting (211-231) first and second feature vectors respectively using first and second feature extractors from the first and second input vectors; estimating (212-232) a first set of weights and a second set of weights respectively from the first and second feature vectors and a pre-step context vector of a sequence generator; Calculating (231-233) a first content vector from the first set of weights and the first feature vectors, and calculating a second content vector from the second set of weights and the second feature vectors; Transforming (214, 234) the first content vector into a first modal content vector having a predetermined dimension and transforming the second content vector into a second modal content vector having the predetermined dimension; estimating (255) a set of modal attention weights from the pre-step context vector and the first and second content vectors or the first and second modal content vectors; Generating (245) a weighted content vector having the predetermined dimension from the set of modal attention weights and the first and second modal content vectors; and Generating (250) a predicted word using the sequence generator to generate the word sequence from the weighted content vector. [13] The non-transitory computer-readable medium of claim 12, wherein the first and second consecutive intervals are an identical interval. [14] The non-transitory computer-readable medium of claim 12, wherein the first and second input vectors are different modalities. [15] The non-transitory computer-readable medium of claim 12, wherein the operations further comprise: Accumulating the predicted word in the memory (140) or the one or more storage devices (130) to generate the word sequence. [16] The non-transitory computer-readable medium of claim 15, wherein the accumulation continues until an end indicator is received. [17] The non-transitory computer-readable medium of claim 12, wherein the operations further comprise: Transferring the generated predicted word from the sequence generator (250). [18] The non-transitory computer-readable medium of claim 12, wherein the first and second feature extractors (211-231) are pre-trained convolutional neural networks (CNNs) trained for an image or video classification task. [19] A method for generating a word sequence from a multimodal input, comprising: Receiving (110, 118) first and second input vectors according to first and second consecutive intervals; Extracting (211-231) first and second feature vectors from the first and second input vectors using first and second feature extractors, respectively; estimating (212-232) a first set of weights and a second set of weights from the first and second feature vectors and a pre-step context vector of a sequence generator; Calculating (213-233) a first content vector from the first set of weights and the first feature vectors, and calculating a second content vector from the second set of weights and the second feature vectors; Transforming (214-234) the first content vector into a first modal content vector having a predetermined dimension, and transforming the second content vector into a second modal content vector having the predetermined dimension; estimating (255) a set of modal attention weights from the pre-step context vector and the first and second content vectors or the first and second modal content vectors; Generating (245) a weighted content vector having the predetermined dimension from the set of modal attention weights and the first and second modal content vectors; and Generating (250) a predicted word using the sequence generator to generate the word sequence from the weighted content vector. [20] The method of claim 19, wherein the first and second consecutive intervals are an identical interval.