Translation method and system based on multi-modal interaction
By introducing multimodal interaction technology into machine translation systems, feature extraction and fusion are performed based on the modal types of the data to be translated, the problem of inaccurate translation in the existing system when there is a lack of text modality is solved, and higher translation accuracy and flexibility are achieved.
Patent Information
- Application Number
- CN202510134220.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-07
AI Technical Summary
Existing machine translation systems are difficult to perform effective syntactic structure analysis and contextual correlation when there is a lack of text modality, resulting in inaccurate translation results.
A translation method based on multimodal interaction is adopted to determine whether there is a text modality in the data to be translated, and corresponding feature extraction and fusion processing is performed based on the existing modal types (text, audio, image), to generate more accurate translation results.
It realizes stable modal translation in the absence of text mode, improves the accuracy and comprehensiveness of translation results, and enhances the flexibility and adaptability of machine translation.
Smart Images

Figure BDA0005262626830000071 
Figure BDA0005262626830000072 
Figure BDA0005262626830000121
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine translation, and in particular to a translation method and system based on multimodal interaction. Background Art
[0002] Machine translation is the process of automatically converting one natural language into another natural language using computer technology. It aims to break down language barriers and enable information exchange and communication between different languages. Machine translation is widely used in various fields. In international business, it can be used to translate documents such as contracts and business emails; in the tourism field, it helps tourists understand local signs, menus and other information; in academic research, it can assist researchers in quickly obtaining the general content of foreign literature. In addition, it also plays an important role in social media, online education and other fields, making it easier for people to communicate and learn across language barriers.
[0003] At present, most machine translation systems only perform translation analysis based on text modality, including lexical analysis, syntactic analysis, and semantic analysis, in order to understand the structure and meaning of the text modality, and convert the analysis results into the representation of the target language according to certain translation models and strategies, and finally adjust and optimize the generated target language text to make it conform to the grammar and expression habits of the target language. However, when the text modality is missing, the existing machine translation system cannot perform syntactic structure analysis due to the lack of text semantic foundation, and it is difficult to capture contextual associations, resulting in difficulty in generating target language text and difficulty in meeting actual translation needs. Summary of the invention
[0004] The purpose of the present invention is to provide a translation method and system based on multimodal interaction, which can stably perform modal translation in the presence or absence of text modality and can improve the accuracy of translation results.
[0005] The technical solution of the present invention is as follows:
[0006] A translation method based on multimodal interaction includes the following operations:
[0007] S1, determine whether there is a text mode in the material to be translated; if so, execute S2; if not, execute S3;
[0008] S2. Determine whether there is an audio mode; if so, perform noise reduction and normalization on the audio mode to obtain an audio normalized mode; translate the audio normalized mode into a text vector to obtain a first audio imitation text vector; fuse the first audio imitation text vector with the corresponding vector of the text mode to obtain an audio-text fusion vector, and execute S4; if not, divide the text mode into several text sub-modalities, and sequentially perform neural network translation and vector conversion, and then splice them to obtain an initial translation vector; splice the initial translation vector and the corresponding vector of the text mode to obtain a text time series vector; perform time series semantic analysis on the text time series vector to obtain a text time series feature vector, and execute S4;
[0009] S3. If only the audio modality exists, the audio modality is processed by speech feature extraction to obtain the speech signal feature; the acoustic feature is obtained based on the speech signal feature, and the decoding process is performed to obtain the acoustic hidden state sequence; the acoustic hidden state sequence is converted into a phoneme sequence, and the phoneme sequence is converted into a text vector according to the pronunciation rules and vocabulary of the language to obtain the second audio imitation text vector, and execute S4; if only the image modality exists, the image modality is subjected to image enhancement processing to obtain the image enhancement modality; the image enhancement modality is subjected to text feature extraction to obtain the image imitation text vector, and execute S4; if both the audio modality and the image modality exist, the context features of the audio modality and the image modality are extracted respectively to obtain the audio feature and the image feature; the feature corresponding to the maximum information entropy in the audio feature and the image feature is used as the imitated feature, and after translating another feature into the imitated feature, it is fused with the imitated feature to obtain the audio image fusion feature; the audio image fusion feature is translated into a text vector to obtain a fused imitation text vector, and execute S4;
[0010] S4, the audio-text fusion feature vector, or the text time series feature vector, or the second audio imitation text vector, or the image imitation text vector, or the fusion imitation text vector is processed by machine translation to obtain a translation result.
[0011] The operation of obtaining the first audio imitation text mode in S2 is specifically as follows: the audio normalized mode is processed by frame division, and each frame of the audio normalized mode is linearly predicted to obtain a number of linear prediction sub-vectors; the Euclidean distance of each linear prediction sub-vector with each subspace in the vector space formed by multiple non-intersecting subspaces is obtained, and each linear prediction sub-vector is mapped to the text representative vector in the subspace corresponding to the minimum Euclidean distance, and the respective audio imitation text sub-vectors are obtained, thereby forming a first audio imitation text vector.
[0012] The operation of obtaining the audio-text fusion feature vector in S2 is specifically as follows: the first audio imitation text vector and the text modality corresponding vector are processed by a multi-layer perceptron respectively to obtain an audio imitation text vector sequence and a text vector sequence; the attention weight matrix between the audio imitation text vector sequence and the text vector sequence is obtained, and based on the attention weight matrix, the audio imitation text vector sequence and the text vector sequence are weightedly summed to obtain the audio-text fusion feature vector.
[0013] The operations of speech feature extraction and processing in S3 are specifically as follows: performing high-frequency enhancement processing on the audio mode to obtain an audio high-frequency enhanced mode; dividing the audio high-frequency enhanced mode into several short frames to form a sequence to obtain a discrete audio frame sequence; performing frame feature enhancement processing on the discrete audio frame sequence to obtain an audio frame enhanced sequence; obtaining frequency domain information of the audio frame enhanced sequence, and analyzing it on the Mel frequency scale to obtain a Mel frequency feature vector; compressing the energy dynamic range of the Mel frequency feature vector to the standard energy dynamic range and then converting it to the cepstrum domain to obtain speech signal features.
[0014] The operation of character feature processing in S3 is achieved by concatenating the visual feature vector obtained based on the image enhancement mode, or / and the text word vector feature vector, or / and the text semantic feature vector, and then performing a nonlinear transformation.
[0015] In S2, if there is an image modality, the visual feature vector, or / and the text word vector feature vector, or / and the text semantic feature vector of the image modality are obtained; the audio-text fusion feature vector or the text timing feature vector is respectively concatenated with the visual feature vector, or / and the text word vector feature vector, or / and the text semantic feature vector to obtain an optimized audio-text fusion feature vector or an optimized text timing feature vector for executing the operation in S4.
[0016] In S3, when the image feature is the feature to be imitated, the operation of translating the image feature into the audio feature is specifically as follows: multiplying the image feature with the key matrix and the value matrix respectively to obtain the image key vector and the image value vector; multiplying the audio feature with the query matrix to obtain the audio query vector; after the audio query vector and the image key vector are subjected to dot product operation, scaling processing and linear processing, they are subjected to dot product operation, residual connection processing, normalization and nonlinear processing with the image value vector to obtain the image-to-audio feature for fusion with the audio feature.
[0017] A translation system based on multimodal interaction, used to implement the above-mentioned translation method based on multimodal interaction, comprising:
[0018] The text modality existence judgment module is used to judge whether there is a text modality in the material to be translated; if so, the audio modality existence judgment and modality feature extraction module is executed; if not, the audio and image modality feature extraction module is executed;
[0019] The audio modality existence judgment and modality feature extraction module is used to judge whether the audio modality exists; if it exists, the audio modality is subjected to noise reduction and normalization processing to obtain the audio normalized modality; the audio normalized modality is translated into a text vector to obtain a first audio imitation text vector; the first audio imitation text vector and the corresponding vector of the text modality are fused to obtain an audio-text fusion vector, and the translation result generation module is executed; if it does not exist, the text modality is divided into several text sub-modalities, which are successively translated by a neural network and converted by vectors, and then spliced to obtain an initial translation vector; the initial translation vector and the corresponding vector of the text modality are spliced to obtain a text time series vector; the text time series vector is subjected to time series semantic analysis to obtain a text time series feature vector, and the translation result generation module is executed;
[0020] Audio and image modality feature extraction module: if only audio modality exists, the audio modality is processed by speech feature extraction to obtain speech signal features; based on the speech signal features, acoustic features are obtained and decoded to obtain an acoustic hidden state sequence; the acoustic hidden state sequence is converted into a phoneme sequence, and the phoneme sequence is converted into a text vector according to the pronunciation rules and vocabulary of the language to obtain a second audio imitation text vector, and the translation result generation module is executed; if only image modality exists, the image modality is subjected to image enhancement processing to obtain an image enhancement modality; the image enhancement modality is subjected to text feature extraction to obtain an image imitation text vector, and the translation result generation module is executed; if both audio modality and image modality exist, the context features of the audio modality and image modality are extracted respectively to obtain audio features and image features; the feature corresponding to the maximum information entropy value in the audio features and image features is used as the imitated feature, and after translating another feature into the imitated feature, it is fused with the imitated feature to obtain an audio-image fusion feature; the audio-image fusion feature is translated into a text vector to obtain a fusion imitation text vector, and the translation result generation module is executed;
[0021] The translation result generation module obtains the translation result by processing the audio-text fusion feature vector, or the text time sequence feature vector, or the second audio imitation text vector, or the image imitation text vector, or the fusion imitation text vector through machine translation.
[0022] A translation device based on multimodal interaction comprises a processor and a memory, wherein the processor implements the above-mentioned translation method based on multimodal interaction when executing a computer program stored in the memory.
[0023] A computer-readable storage medium is used to store a computer program, wherein the computer program implements the above-mentioned translation method based on multimodal interaction when executed by a processor.
[0024] The beneficial effects of the present invention are:
[0025] The present invention provides a translation method based on multimodal interaction. Different methods for extracting features to be translated are respectively executed according to whether there is a text modality in the material to be translated. On the one hand, when there is a text modality, different operations are respectively executed according to whether there is an audio modality whose influence on translation accuracy is second only to the text modality. When there is an audio modality, the audio modality is translated into a text vector and then fused with a vector corresponding to the text modality to realize the interaction between the audio modality information and the text modality information, thereby improving the feature expression ability of the features to be translated and obtaining an audio-text fusion vector. When there is no audio modality, an initial translation vector obtained after a preliminary translation of the text modality, which contains a preliminary understanding and conversion of the source language text, is fused with a vector corresponding to the text modality containing feature information of the source language text. After splicing, the temporal semantic analysis is performed to further improve the feature expression ability of the text modality; on the other hand, when there is no text modality, the audio modality and / or the image modality is translated into a text vector with stronger feature expression ability and higher translation efficiency, so as to improve the expression ability of the audio modality and / or the image modality in vocabulary, grammatical structure and semantics, and obtain a second audio imitation text vector, or a second audio imitation text vector, or a fused imitation text vector; finally, the audio-text fusion feature vector, or the text temporal feature vector, or the second audio imitation text vector, or the image imitation text vector, or the fused imitation text vector with stronger vocabulary, grammatical structure and semantic expression ability is processed by machine translation to obtain a more accurate translation result, thereby realizing the comprehensiveness, flexibility, adaptability and accuracy of machine translation. DETAILED DESCRIPTION
[0026] This embodiment provides a translation method based on multimodal interaction, including the following operations:
[0027] S1, determine whether there is a text mode in the material to be translated; if so, execute S2; if not, execute S3;
[0028] S2. Determine whether there is an audio mode; if so, perform noise reduction and normalization on the audio mode to obtain an audio normalized mode; translate the audio normalized mode into a text vector to obtain a first audio imitation text vector; fuse the first audio imitation text vector with the corresponding vector of the text mode to obtain an audio-text fusion vector, and execute S4; if not, divide the text mode into several text sub-modalities, and sequentially perform neural network translation and vector conversion, and then splice them to obtain an initial translation vector; splice the initial translation vector and the corresponding vector of the text mode to obtain a text time series vector; perform time series semantic analysis on the text time series vector to obtain a text time series feature vector, and execute S4;
[0029] S3. If only the audio modality exists, the audio modality is processed by speech feature extraction to obtain the speech signal feature; the acoustic feature is obtained based on the speech signal feature, and the decoding process is performed to obtain the acoustic hidden state sequence; the acoustic hidden state sequence is converted into a phoneme sequence, and the phoneme sequence is converted into a text vector according to the pronunciation rules and vocabulary of the language to obtain the second audio imitation text vector, and execute S4; if only the image modality exists, the image modality is subjected to image enhancement processing to obtain the image enhancement modality; the image enhancement modality is subjected to text feature extraction to obtain the image imitation text vector, and execute S4; if both the audio modality and the image modality exist, the context features of the audio modality and the image modality are extracted respectively to obtain the audio feature and the image feature; the feature corresponding to the maximum information entropy in the audio feature and the image feature is used as the imitated feature, and after translating another feature into the imitated feature, it is fused with the imitated feature to obtain the audio image fusion feature; the audio image fusion feature is translated into a text vector to obtain a fused imitation text vector, and execute S4;
[0030] S4, the audio-text fusion feature vector, or the text time series feature vector, or the second audio imitation text vector, or the image imitation text vector, or the fusion imitation text vector is processed by machine translation to obtain a translation result.
[0031] S1. Determine whether there is a text mode in the material to be translated; if so, execute S2; if not, execute S3.
[0032] Compared with audio modality and image modality, text modality has clearer vocabulary and grammatical structure, clearer semantics, and greater impact on the accuracy and efficiency of translation results. Therefore, this embodiment executes different translation methods according to whether there is text modality in the material to be translated, so as to improve the accuracy, flexibility and comprehensiveness of the translation method.
[0033] S2. Determine whether there is an audio modality; if so, perform noise reduction and normalization on the audio modality to obtain an audio normalized modality; translate the audio normalized modality into a text vector to obtain a first audio imitation text vector; fuse the first audio imitation text vector with the corresponding vector of the text modality to obtain an audio-text fusion vector, and execute S4; if not, divide the text modality into several text sub-modalities, translate them through a neural network and transform them into vectors in turn, and then splice them to obtain an initial translation vector; splice the initial translation vector and the corresponding vector of the text modality to obtain a text time series vector; perform time series semantic analysis on the text time series vector to obtain a text time series feature vector, and execute S4.
[0034] When there is a text modality, different operations are performed according to whether there is an audio modality whose impact on translation accuracy is second only to the text modality; when there is an audio modality, the audio modality is translated into a text vector and then fused with the corresponding vector of the text modality, thereby improving the feature expression ability of the features to be translated and obtaining an audio-text fusion vector; when there is no audio modality, the text modality is divided into several sub-modalities and then translated in sequence through a neural network, so that the latter sub-modality will be translated according to the translation result of the previous sub-modality during translation, thereby improving the accuracy of the initial translation vector, and then the initial translation vector containing a preliminary understanding and conversion of the source language text is spliced with the corresponding vector of the text modality containing the feature information of the source language text, and then a temporal semantic analysis is performed to obtain a text temporal feature vector, thereby further improving the feature expression ability of the text modality and improving the accuracy of subsequent translations.
[0035] Compared with the image modality, the audio modality can obtain language information more directly and is richer in semantic and emotional information, and has a greater impact on the accuracy and efficiency of the translation results. Therefore, when there is a text modality in the material to be translated, in order to further improve the accuracy of the translation results, it is necessary to further determine whether there is an audio modality.
[0036] If there is an audio modality, in order to improve the vocabulary, grammatical structure and semantic expression ability of the audio modality while improving translation efficiency, the audio modality is translated into a text vector and fused with the corresponding vector of the text modality (the vector obtained after the text modality is embedded) to obtain an audio-text fusion vector with stronger feature expression ability. The specific steps are as follows.
[0037] Step 1: Perform denoising (which can be achieved through the Wiener filter method) and normalization on the audio mode to make the energy of different audio signals comparable, so as to facilitate subsequent feature analysis and obtain the audio normalized mode.
[0038] Among them, the normalization operation can be achieved by the following formula:
[0039]
[0040] y(n) is the normalized audio mode, x(n) is the noise-reduced audio mode of length n (obtained after the audio mode is subjected to noise reduction processing), and N is the total length of the noise-reduced audio mode.
[0041] Step 2: Translate the audio normalized modality into a text vector to obtain the first audio imitation text vector. Specifically, the audio normalized modality is processed by frame division, and each frame of the audio normalized modality is linearly predicted to obtain a number of linear prediction sub-vectors; the Euclidean distance of each linear prediction sub-vector with each subspace in a vector space formed by multiple non-intersecting subspaces (each subspace is represented by a text representative vector) is obtained, and each linear prediction sub-vector is mapped to the text representative vector in the subspace corresponding to the minimum Euclidean distance, so that each linear prediction sub-vector is encoded as a corresponding text codeword index to obtain the respective audio imitation text sub-vector, and the vector composed of these text codeword indices can be regarded as a text-like representation, forming the first audio imitation text vector.
[0042] The linear prediction operation can be achieved through the following formula:
[0043]
[0044] S i is the i-th linear predictor vector, α k is the k-th order linear prediction coefficient, K is the total order of linear prediction, s i The normalized mode of the audio of the i-th frame.
[0045] Step 3: The first audio imitation text vector and the text modality corresponding vector are fused to obtain an audio-text fusion vector, and then S4 is executed.
[0046] The specific process is as follows: the first audio imitation text vector and the text modality corresponding vector are processed by a multi-layer perceptron respectively to obtain a first audio imitation text vector sequence and a text vector sequence; the attention weight matrix between the first audio imitation text vector sequence and the text vector sequence is obtained, and based on the attention weight matrix, the first audio imitation text vector sequence and the text vector sequence are weighted and summed to obtain an audio-text fusion feature vector. In the above attention weight matrix, the attention weight is obtained based on the cosine similarity between the u-th subsequence of the first audio imitation text vector sequence and the v-th first subsequence in the text vector sequence, that is, the attention weight is obtained based on the cosine similarity of the two subsequences.
[0047] Furthermore, if there is both text modality and audio modality and image modality, in order to further improve the information richness of the features to be translated, the visual feature vector of the image modality, or / and the text word vector feature vector, or / and the text semantic feature vector are obtained; the audio-text fusion feature vector is spliced with the visual feature vector, or / and the text word vector feature vector, or / and the text semantic feature vector to achieve interactive fusion of text modality information, audio modality information and image modality information, and obtain an optimized audio-text fusion feature vector for executing the operation in S4.
[0048] The visual feature vector of the above image modality is achieved by processing the image modality with a convolutional neural network.
[0049] The operations for obtaining the text word vector feature vector of the above-mentioned image modality are specifically as follows: graying the image modality, simplifying the color information of the image, highlighting the brightness characteristics of the image, and obtaining a grayscale image; subjecting the grayscale image to median filtering to remove salt and pepper noise of the image and Gaussian filtering to remove Gaussian noise of the image, so that the image can maintain good quality after noise removal, providing more reliable image data for subsequent text recognition, and obtaining a denoised image; binarizing the denoised image, so that the text part in the image is represented by black pixels, and the background part can be represented by white pixels, so that the text and the background are clearly separated, and a binary image is obtained; the text information in the binary image is obtained, and after embedding processing, a text word vector feature vector is obtained.
[0050] The operation of obtaining the text semantic feature vector of the above-mentioned image modality can be realized by performing a mapping process based on a neural network on the image modality. The operation of the mapping process based on the neural network is specifically as follows: the image modality is embedded to obtain an image vector; the image vector is multiplied by the weight matrix in the hidden layer, and after adding the bias vector of the hidden layer, a nonlinear transformation is performed through an activation function to obtain the output feature of the image hidden layer; the output feature of the image hidden layer is multiplied by the weight matrix of the output layer, and the bias vector of the output layer is added, so as to map the image modality feature into a pseudo-text vector with text semantic features, and obtain a text semantic feature vector.
[0051] If there is no audio modality and only text modality exists, in order to further enhance the feature expression capability of the text modality, perform the following operations.
[0052] Step 1: The text modality is divided into several text sub-modalities, and the several text sub-modalities are sequentially translated by neural network (so that the subsequent sub-modality will be translated according to the translation result of the previous sub-modality during translation, improving the accuracy of the initial translation) and vector conversion, and then spliced to obtain an initial translation vector. The above-mentioned neural network translation includes but is not limited to being implemented by a recurrent neural network language model (Recurrent Neural Network-Language Model, RNN-LM).
[0053] Step 2: Because some words and expressions in the source text modality often have multiple meanings, their specific meanings may not be accurately determined by the initial translation results alone. Therefore, in this embodiment, the initial translation vector that contains a preliminary understanding and conversion of the source language text is concatenated with the text modality corresponding vector (i.e., the text vector) containing the feature information of the source language text to obtain a text time series vector that contains both historical translation information and richer semantic information, which can improve the accuracy of subsequent translations.
[0054] Step 3: The text time series vector is subjected to time series semantic analysis to further extract text features, obtain a text time series feature vector, and execute S4.
[0055] The operations of temporal semantic analysis can be: the text temporal vector is multiplied with the query item parameter matrix, the key item parameter matrix and the value item parameter matrix respectively to obtain the text temporal query vector, the text temporal key vector and the text temporal value vector; the text temporal query vector, the text temporal key vector and the text temporal value vector are multiplied with different parameter weights respectively to obtain several text temporal query sub-vectors, several text temporal key sub-vectors and several text temporal value sub-vectors; the attention scores of each text temporal query sub-vector and each text temporal key sub-vector are obtained, and the scores are added to the corresponding mask matrices and processed by probability mapping to obtain several mask scores; the weighted sum of several mask scores and each text temporal value sub-vector is processed by splicing, linear processing and nonlinear processing to obtain the text temporal feature vector. The above attention score is obtained based on the product of the text temporal query sub-vector and the text temporal key sub-vector.
[0056] The operation of temporal semantic analysis can also be implemented through the Long Short-Term Memory (LSTM) network.
[0057] Furthermore, if there is text modality and image modality, the visual feature vector of the image modality, or / and the text word vector feature vector, or / and the text semantic feature vector are obtained; the text timing feature vector is spliced with the visual feature vector, or / and the text word vector feature vector, or / and the text semantic feature vector to realize the interaction of text modality information and image modality information, and the optimized text timing feature vector is obtained for executing the operation in S4.
[0058] The methods for obtaining the visual feature vector of the image modality, the text word vector feature vector, and the text semantic feature vector have been described above and will not be repeated here to save space.
[0059] S3. If only the audio modality exists, the audio modality is processed by speech feature extraction to obtain the speech signal feature; the acoustic feature is obtained based on the speech signal feature, and the decoding process is performed to obtain the acoustic hidden state sequence; the acoustic hidden state sequence is converted into a phoneme sequence, and the phoneme sequence is converted into a text vector according to the pronunciation rules and vocabulary of the language to obtain a second audio imitation text vector, and execute S4; if only the image modality exists, the image modality is processed by image enhancement to obtain the image enhancement modality; the image enhancement modality is subjected to text feature extraction to obtain the image imitation text vector, and execute S4; if both the audio modality and the image modality exist, the context features of the audio modality and the image modality are extracted respectively to obtain the audio feature and the image feature; the feature corresponding to the maximum information entropy in the audio feature and the image feature is used as the imitated feature, and after translating another feature into the imitated feature, it is fused with the imitated feature to obtain the audio image fusion feature; the audio image fusion feature is translated into a text vector to obtain a fused imitation text vector, and execute S4.
[0060] When there is no text modality, in order to improve the expression ability of existing modalities (audio modality and / or image modality) in vocabulary, grammatical structure and semantics, the audio modality and / or image modality are translated into a text vector with stronger feature expression ability and higher translation efficiency to obtain a second audio imitation text vector, or a second audio imitation text vector, or a fused imitation text vector, which is used to improve the accuracy of subsequent translation results.
[0061] When there is only audio mode but no text mode, in order to improve the information richness of the features to be translated, the detailed speech signal features in the audio mode are extracted, decoded into an acoustic hidden state sequence and converted into a phoneme sequence. According to the pronunciation rules and vocabulary of the language, the phoneme sequence is converted into a text vector that is rich in details and easy to translate. The specific operations are as follows.
[0062] Step 1: The audio modality is processed by speech feature extraction to obtain speech signal features.
[0063] The operation of speech feature extraction and processing is specifically as follows: during the transmission process of the speech signal in the audio mode, the high-frequency part will be attenuated to a certain extent. Therefore, the audio mode is first subjected to high-frequency enhancement processing (including but not limited to being implemented by a first-order FIR filter) to enhance the high-frequency component of the speech signal and obtain an audio high-frequency enhanced mode, which is conducive to subsequent feature extraction and analysis; then, in order to facilitate digital signal processing, the audio high-frequency enhanced mode is divided into several short frames to form a sequence to obtain a discrete audio frame sequence; subsequently, in order to prevent and reduce spectral leakage of the discrete audio frame sequence, the discrete audio frame sequence is subjected to frame feature enhancement processing (including but not limited to being implemented by applying a preset window function to each frame sequence) so that each frame sequence signal is gradually enhanced at both ends. Gradually attenuate, effectively reduce spectrum leakage, and obtain an audio frame enhancement sequence; then, obtain the frequency domain information of the audio frame enhancement sequence (including but not limited to through fast Fourier transform), and analyze it on the Mel frequency scale to obtain the characteristics of speech signals of different frequencies, so as to improve the robustness and distinguishability of acoustic feature analysis, and obtain the Mel frequency feature vector; compress the energy dynamic range of the Mel frequency feature vector to the standard energy dynamic range (which can be achieved by performing logarithmic operation on the Mel frequency feature vector) and then convert it to the cepstrum domain (including but not limited to through discrete cosine transform method) to achieve the extraction of speech signal details on the basis of highlighting the energy change of the low-frequency part and reducing the noise influence of the high-frequency part, and obtain the speech signal characteristics.
[0064] Step 2: Acoustic features are obtained based on the speech signal features, and decoding is performed to obtain an acoustic hidden state sequence. The above operation of obtaining acoustic features based on the speech signal features can be achieved by training a deep neural network to process the speech signal features, and the decoding operation includes but is not limited to being achieved through a hidden Markov model.
[0065] Step 3: Convert the acoustic hidden state sequence into a phoneme sequence, convert the phoneme sequence into a text vector according to the pronunciation rules and vocabulary of the language, obtain a second audio imitation text vector, and execute S4. The above-mentioned operation of converting the acoustic hidden state sequence into a phoneme sequence includes but is not limited to being implemented by an N-Gram model, a recurrent neural network language model, or a long short-term memory network language model.
[0066] If only the image modality exists, the image modality is subjected to image enhancement processing (including but not limited to being realized by an adaptive equalization method) to obtain the image enhancement modality; the image enhancement modality is subjected to text feature extraction to obtain an image imitation text vector, and S4 is executed. Among them, the text feature processing operation is achieved by splicing the visual feature vector, or / and the text word vector feature vector, or / and the text semantic feature vector obtained based on the image enhancement modality, and then performing a nonlinear transformation. The method for obtaining the visual feature vector, the text word vector feature vector, and the text semantic feature vector has been described in the above content, and will not be repeated here to save space.
[0067] If both audio and image modalities exist, the interaction between the audio and image modal information is translated into a text vector to improve the expressive power of the features to be translated and the translation efficiency. The operation is as follows.
[0068] Step 1: Extract context features of the audio modality and image modality respectively to obtain audio features and image features.
[0069] The specific operations for obtaining audio features are as follows: the audio modality is processed by multi-head attention to obtain audio attention features; the audio attention features are processed by residual connection to obtain audio residual features; the audio residual features and the corresponding audio attention weights are processed by layer normalization and nonlinearity to obtain audio features.
[0070] The operation of obtaining the image features may be the same as the operation of obtaining the audio features described above, and may also be implemented by subjecting the image modality to spatial pyramid pooling to obtain context information at different levels of the image to obtain the image features.
[0071] Step 2: The feature corresponding to the maximum information entropy value among the audio features and the image features is used as the simulated feature, and the other feature is translated into the simulated feature, and then fused with the simulated feature to obtain the audio-image fusion feature.
[0072] Information entropy can be obtained by the following formula:
[0073]
[0074] H is information entropy. The larger the information entropy, the richer the corresponding modal information. i ) is the probability of occurrence of the i-th information feature in the audio feature or image feature, and I is the total number of information feature types in the audio feature or image feature.
[0075] In the process of translating another modality into the imitated modality, when the audio feature is the imitated feature,
[0076] The specific operation of translating image features into audio features is as follows: multiplying the image features with the key matrix and the value matrix respectively to obtain the image key vector and the image value vector; multiplying the audio features with the query matrix to obtain the audio query vector; the audio query vector and the image key vector are dot-producted, scaled, and linearly processed, and then dot-producted, residual-connected, normalized, and nonlinearly processed with the image value vector to obtain the image-to-audio features for fusion with the audio features. The operation of translating audio features into image features is the same as the operation of translating image features into audio features.
[0077] Step 3: Translate the audio image fusion features into a text vector (which can be achieved by training a Transformer pre-trained model or other training neural network), obtain a fused pseudo-text vector, and execute S4.
[0078] S4, the audio-text fusion feature vector, or the text time series feature vector, or the second audio imitation text vector, or the image imitation text vector, or the fusion imitation text vector is processed by machine translation to obtain a translation result.
[0079] The audio-text fusion feature vector, or the text time series feature vector, or the second audio imitation text vector, or the image imitation text vector, or the fusion imitation text vector with stronger vocabulary, grammatical structure and semantic expression capabilities is processed through machine translation (which can be achieved through the N-Gram model, or the recurrent neural network language model, or the long short-term memory network language model, or the training Transformer pre-training model) to obtain a more accurate translation result.
[0080] This embodiment further provides a translation system based on multimodal interaction, which is used to implement the above-mentioned translation method based on multimodal interaction, including:
[0081] The text modality existence judgment module is used to judge whether there is a text modality in the material to be translated; if so, the audio modality existence judgment and modality feature extraction module is executed; if not, the audio and image modality feature extraction module is executed;
[0082] The audio modality existence judgment and modality feature extraction module is used to judge whether the audio modality exists; if it exists, the audio modality is subjected to noise reduction and normalization processing to obtain the audio normalized modality; the audio normalized modality is translated into a text vector to obtain a first audio imitation text vector; the first audio imitation text vector and the corresponding vector of the text modality are fused to obtain an audio-text fusion vector, and the translation result generation module is executed; if it does not exist, the text modality is divided into several text sub-modalities, which are successively translated by a neural network and converted by vectors, and then spliced to obtain an initial translation vector; the initial translation vector and the corresponding vector of the text modality are spliced to obtain a text time series vector; the text time series vector is subjected to time series semantic analysis to obtain a text time series feature vector, and the translation result generation module is executed;
[0083] Audio and image modality feature extraction module: if only audio modality exists, the audio modality is processed by speech feature extraction to obtain speech signal features; based on the speech signal features, acoustic features are obtained and decoded to obtain an acoustic hidden state sequence; the acoustic hidden state sequence is converted into a phoneme sequence, and the phoneme sequence is converted into a text vector according to the pronunciation rules and vocabulary of the language to obtain a second audio imitation text vector, and the translation result generation module is executed; if only image modality exists, the image modality is subjected to image enhancement processing to obtain an image enhancement modality; the image enhancement modality is subjected to text feature extraction to obtain an image imitation text vector, and the translation result generation module is executed; if both audio modality and image modality exist, the context features of the audio modality and image modality are extracted respectively to obtain audio features and image features; the feature corresponding to the maximum information entropy value in the audio features and image features is used as the imitated feature, and after translating another feature into the imitated feature, it is fused with the imitated feature to obtain an audio-image fusion feature; the audio-image fusion feature is translated into a text vector to obtain a fusion imitation text vector, and the translation result generation module is executed;
[0084] The translation result generation module obtains the translation result by processing the audio-text fusion feature vector, or the text time sequence feature vector, or the second audio imitation text vector, or the image imitation text vector, or the fusion imitation text vector through machine translation.
[0085] This embodiment further provides a translation device based on multimodal interaction, including a processor and a memory, wherein the processor implements the above-mentioned translation method based on multimodal interaction when executing a computer program stored in the memory.
[0086] This embodiment further provides a computer-readable storage medium for storing a computer program, wherein the computer program implements the above-mentioned translation method based on multimodal interaction when executed by a processor.
[0087] The present embodiment provides a translation method based on multimodal interaction, which respectively executes different methods for extracting features to be translated according to whether there is a text modality in the material to be translated; on the one hand, when there is a text modality, different operations are respectively performed according to whether there is an audio modality whose influence on translation accuracy is second only to the text modality; wherein, when there is an audio modality, the audio modality is translated into a text vector and then fused with the corresponding vector of the text modality to realize the interaction between the audio modality information and the text modality information, thereby improving the feature expression ability of the features to be translated and obtaining an audio-text fusion vector; when there is no audio modality, the initial translation vector obtained after the preliminary translation of the text modality, which contains a preliminary understanding and conversion of the source language text, is fused with the corresponding vector of the text modality containing the feature information of the source language text After splicing, the temporal semantic analysis is performed to further improve the feature expression ability of the text modality; on the other hand, when there is no text modality, the audio modality and / or the image modality is translated into a text vector with stronger feature expression ability and higher translation efficiency, so as to improve the expression ability of the audio modality and / or the image modality in vocabulary, grammatical structure and semantics, and obtain a second audio imitation text vector, or a second audio imitation text vector, or a fused imitation text vector; finally, the audio-text fusion feature vector, or the text temporal feature vector, or the second audio imitation text vector, or the image imitation text vector, or the fused imitation text vector with stronger vocabulary, grammatical structure and semantic expression ability is processed by machine translation to obtain a more accurate translation result, thereby realizing the comprehensiveness, flexibility, adaptability and accuracy of machine translation.
Claims
1. A translation method based on multimodal interaction, characterized in that: The following operations are included: S1, determine whether there is a text mode in the material to be translated; if so, execute S2; if not, execute S3; S2, determining whether there is an audio mode; If it exists, the audio mode is subjected to denoising and normalization to obtain an audio normalized mode; Translate the audio normalized modality into a text vector to obtain a first audio imitation text vector; The first audio imitation text vector and the text modality corresponding vector are fused to obtain an audio text fusion vector, and S4 is executed; If it does not exist, the text modality is divided into several text sub-modalities, which are successively translated by a neural network and converted into vectors, and then concatenated to obtain an initial translation vector; the initial translation vector and the corresponding vector of the text modality are concatenated to obtain a text time series vector; the text time series vector is subjected to time series semantic analysis to obtain a text time series feature vector, and S4 is executed; S3, if only the audio mode exists, the audio mode is processed by speech feature extraction to obtain speech signal features; based on the speech signal features, acoustic features are obtained, and decoding is performed to obtain an acoustic hidden state sequence; the acoustic hidden state sequence is converted into a phoneme sequence, and according to the pronunciation rules and vocabulary of the language, the phoneme sequence is converted into a text vector to obtain a second audio imitation text vector, and S4 is executed; If only the image modality exists, the image modality is subjected to image enhancement processing to obtain the image enhancement modality; the image enhancement modality is subjected to text feature extraction to obtain the image imitation text vector, and S4 is executed; If both audio mode and image mode exist, extract the context features of the audio mode and image mode respectively to obtain audio features and image features; take the feature corresponding to the maximum information entropy value in the audio features and image features as the imitated feature, translate another feature into the imitated feature, and fuse it with the imitated feature to obtain the audio-image fusion feature; translate the audio-image fusion feature into a text vector to obtain a fused imitation text vector, and execute S4; S4, the audio-text fusion feature vector, or the text time series feature vector, or the second audio imitation text vector, or the image imitation text vector, or the fusion imitation text vector is processed by machine translation to obtain a translation result.
2. The translation method based on multimodal interaction according to claim 1, characterized in that: In S2, the operation of obtaining the first audio imitation text mode is specifically: The audio normalized mode is processed by frame division, and each frame of the audio normalized mode is linearly predicted to obtain a number of linear prediction sub-vectors; Obtain the Euclidean distance of each linear prediction sub-vector with each subspace in the vector space formed by multiple non-intersecting subspaces, map each linear prediction sub-vector to the text representative vector in the subspace corresponding to the minimum Euclidean distance, obtain the respective audio imitation text sub-vector, and form a first audio imitation text vector.
3. The translation method based on multimodal interaction according to claim 1, characterized in that: In S2, the operation of obtaining the audio-text fusion feature vector is specifically as follows: The first audio imitation text vector and the text modality corresponding vector are processed by a multi-layer perceptron respectively to obtain an audio imitation text vector sequence and a text vector sequence; an attention weight matrix between the audio imitation text vector sequence and the text vector sequence is obtained, and based on the attention weight matrix, the audio imitation text vector sequence and the text vector sequence are weightedly summed to obtain an audio-text fusion feature vector.
4. The translation method based on multimodal interaction according to claim 1, characterized in that: In S3, the operation of speech feature extraction processing is specifically as follows: Performing high-frequency enhancement processing on the audio mode to obtain an audio high-frequency enhancement mode; dividing the audio high-frequency enhancement mode into a plurality of short frames to form a sequence to obtain a discrete audio frame sequence; The discrete audio frame sequence is processed by frame feature enhancement to obtain an audio frame enhanced sequence; the frequency domain information of the audio frame enhanced sequence is obtained and analyzed on the Mel frequency scale to obtain a Mel frequency feature vector; the energy dynamic range of the Mel frequency feature vector is compressed to the standard energy dynamic range and then converted to the cepstrum domain to obtain the speech signal characteristics.
5. The translation method based on multimodal interaction according to claim 1, characterized in that: In S3, the text feature processing operation is achieved by concatenating the visual feature vector obtained based on the image enhancement mode, or / and the text word vector feature vector, or / and the text semantic feature vector, and then performing a nonlinear transformation.
6. The translation method based on multimodal interaction according to claim 1, characterized in that: In S2, if there is an image modality, a visual feature vector of the image modality, or / and a text word vector feature vector, or / and a text semantic feature vector are obtained; The audio-text fusion feature vector or the text timing feature vector is concatenated with the visual feature vector, or / and the text word vector feature vector, or / and the text semantic feature vector to obtain an optimized audio-text fusion feature vector or an optimized text timing feature vector for executing the operation in S4.
7. The translation method based on multimodal interaction according to claim 1, characterized in that: In S3, when the image feature is the feature to be imitated, the operation of translating the image feature into the audio feature is specifically as follows: The image features are multiplied with the key matrix and the value matrix respectively to obtain the image key vector and the image value vector; the audio features are multiplied with the query matrix to obtain the audio query vector; the audio query vector and the image key vector are dot-product processed, scaled and linearly processed, and then dot-product processed, residual connected, normalized and nonlinearly processed with the image value vector to obtain the image-to-audio features for fusion with the audio features.
8. A translation system based on multimodal interaction, used to implement the translation method based on multimodal interaction described in claim 1, characterized in that: include: A text modality existence judgment module is used to judge whether there is a text modality in the material to be translated; If it exists, execute the audio mode existence judgment and modal feature extraction module; If it does not exist, execute the audio and image modality feature extraction module; The audio mode existence judgment and modal feature extraction module is used to judge whether the audio mode exists; If it exists, the audio mode is subjected to noise reduction and normalization processing to obtain an audio normalized mode; the audio normalized mode is translated into a text vector to obtain a first audio imitation text vector; the first audio imitation text vector is fused with the corresponding vector of the text mode to obtain an audio-text fusion vector, and the translation result generation module is executed; if it does not exist, the text mode is divided into several text sub-modalities, which are successively translated by a neural network and converted into vectors, and then spliced to obtain an initial translation vector; the initial translation vector and the corresponding vector of the text mode are spliced to obtain a text time series vector; the text time series vector is subjected to time series semantic analysis to obtain a text time series feature vector, and the translation result generation module is executed; Audio and image modality feature extraction module: if only audio modality exists, the audio modality is processed by speech feature extraction to obtain speech signal features; based on the speech signal features, acoustic features are obtained and decoded to obtain an acoustic hidden state sequence; the acoustic hidden state sequence is converted into a phoneme sequence, and according to the pronunciation rules and vocabulary of the language, the phoneme sequence is converted into a text vector to obtain a second audio imitation text vector, and the translation result generation module is executed; If only the image modality exists, the image modality is subjected to image enhancement processing to obtain the image enhancement modality; the image enhancement modality is subjected to text feature extraction to obtain the image imitation text vector, and the translation result generation module is executed; If both audio mode and image mode exist, extract the context features of audio mode and image mode respectively to obtain audio features and image features; take the feature corresponding to the maximum information entropy value in the audio features and image features as the imitated feature, translate another feature into the imitated feature, and fuse it with the imitated feature to obtain the audio-image fusion feature; translate the audio-image fusion feature into a text vector to obtain a fusion imitation text vector, and execute the translation result generation module; The translation result generation module obtains the translation result by processing the audio-text fusion feature vector, or the text time sequence feature vector, or the second audio imitation text vector, or the image imitation text vector, or the fusion imitation text vector through machine translation.
9. A translation device based on multimodal interaction, characterized in that: The invention comprises a processor and a memory, wherein when the processor executes the computer program stored in the memory, the translation method based on multimodal interaction as claimed in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: Used to store a computer program, wherein when the computer program is executed by a processor, the translation method based on multimodal interaction as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Multi-modal data preprocessing method based on machine translation
CN118378029A
Multimodal fusion speech translation method, system and equipment
CN118692446A
Multi-mode sentiment analysis method and system taking audio mode as target mode
CN118965139A
Fused acoustic and text encoding for multimodal bilingual pretraining and speech translation
US20230169281A1