Multi-language real-time translation interaction system and method based on artificial intelligence
Through a multilingual real-time translation interaction system based on artificial intelligence, deep learning technology and generators are used to generate real-time translation results, the problem of insufficient accuracy of speech translation in the existing technology is solved, and higher translation accuracy and quality are achieved.
Patent Information
- Application Number
- CN202510134562.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-06-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing speech translation techniques have shortcomings in terms of translation accuracy, especially when the stages of converting user speech into text may cause errors, affecting the quality of the final translation.
Using a multilingual real-time translation interaction system based on artificial intelligence, we use deep learning technology to perform feature extraction and correlation analysis through the generator, and finally generate real-time translation results to reduce speech recognition errors.
It improves the accuracy of multilingual real-time translation interaction, ensures the quality of the final translation results, and reduces speech recognition errors.
Smart Images

Figure CN120124644A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent translation, and more specifically, to a multi-language real-time translation interaction system and method based on artificial intelligence. Background Art
[0002] Multi-language translation refers to converting the content of one language into another through language processing technology to achieve cross-language communication and understanding.
[0003] With the continuous progress of artificial intelligence and natural language processing technologies, speech translation technology has been widely used in fields such as simultaneous interpretation and foreign language teaching. For example, in the scenario of simultaneous interpretation, speech translation technology can convert the language of the speaker into the target language in real time, greatly facilitating cross-language communication. However, current speech translation technology still has some deficiencies, especially in terms of translation accuracy. During the translation process, especially in the stage of converting the user's speech into text, some errors may occur, thus affecting the quality of the final translation.
[0004] Therefore, a multi-language real-time translation interaction system and method based on artificial intelligence are desired. Summary of the Invention
[0005] To solve the above technical problems, the present application is proposed. Embodiments of the present application provide a multi-language real-time translation interaction system and method based on artificial intelligence, which first acquire the user speech information collected by a sound sensor, then use deep learning technology to perform feature extraction and correlation analysis on it, and finally generate a real-time translation result through a generator, thereby reducing speech recognition errors, improving the accuracy of multi-language real-time translation interaction, and ensuring the quality of the final translation result.
[0006] According to one aspect of the present application, there is provided a multi-language real-time translation interaction system based on artificial intelligence, which includes:
[0007] A user speech information acquisition module for acquiring the user speech information collected by a sound sensor;
[0008] A user speech information extraction module for extracting a user communication text information associated semantic feature vector and a user communication speech log mel spectrogram feature vector from the user speech information collected by the sound sensor;
[0009] A real-time translation result generation module for generating a real-time translation result based on the user communication text information associated semantic feature vector and the user communication speech log mel spectrogram feature vector.
[0010] According to another aspect of the present application, there is provided a multi-language real-time translation interaction method based on artificial intelligence, which includes:
[0011] Obtain the user voice information collected by the voice sensor;
[0012] Extract the semantic feature vector associated with the user communication text information and the logarithmic Mel spectrogram feature vector of the user communication voice from the user voice information collected by the voice sensor;
[0013] Generate a real-time translation result based on the semantic feature vector associated with the user communication text information and the logarithmic Mel spectrogram feature vector of the user communication voice.
[0014] Compared with the prior art, a multi-language real-time translation interaction system and method based on artificial intelligence provided by the present application first obtains the user voice information collected by the voice sensor, then uses deep learning technology to perform feature extraction and correlation analysis on it, and finally passes through a generator to generate a real-time translation result, thereby reducing the speech recognition error, improving the accuracy of multi-language real-time translation interaction, and ensuring the quality of the final translation result. Description of the Drawings
[0015] By describing the embodiments of the present application in more detail in conjunction with the drawings, the above and other objects, features, and advantages of the present application will become more obvious. The drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification, and are used to explain the present application together with the embodiments of the present application, and do not constitute a limitation to the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0016] Figure 1 It is a block diagram of a multi-language real-time translation interaction system based on artificial intelligence according to an embodiment of the present application.
[0017] Figure 2 It is a block diagram of a user voice information extraction module in a multi-language real-time translation interaction system based on artificial intelligence according to an embodiment of the present application.
[0018] Figure 3 It is a block diagram of a real-time translation result generation module in a multi-language real-time translation interaction system based on artificial intelligence according to an embodiment of the present application.
[0019] Figure 4 It is a flowchart of a multi-language real-time translation interaction method based on artificial intelligence according to an embodiment of the present application. Detailed Embodiments
[0020] Next, exemplary embodiments according to the present application will be described in detail with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.
[0021] Figure 1 This is a block diagram of a multi - language real - time translation interaction system based on artificial intelligence according to an embodiment of the present application. As Figure 1 shown, the multi - language real - time translation interaction system 100 based on artificial intelligence according to an embodiment of the present application includes: a user voice information acquisition module 110, configured to acquire user voice information collected by a sound sensor; a user voice information extraction module 120, configured to extract a user communication text information - related semantic feature vector and a user communication voice log mel - spectrum feature vector from the user voice information collected by the sound sensor; and a real - time translation result generation module 130, configured to generate a real - time translation result based on the user communication text information - related semantic feature vector and the user communication voice log mel - spectrum feature vector.
[0022] In the above - mentioned multi - language real - time translation interaction system 100 based on artificial intelligence, the user voice information acquisition module 110 is configured to acquire user voice information collected by a sound sensor. It should be understood that voice is one of the most direct forms of language communication, and many application scenarios (such as simultaneous interpretation, voice assistants, automatic translation, etc.) rely on voice input to interact with users. The user's voice signal can be captured by a sound sensor, and these signals will become the input data for subsequent voice processing, text conversion, translation, or other operations. Specifically, the process of acquiring voice information is usually implemented by a sound sensor (such as a microphone array). The main function of the sound sensor is to convert sound waves in the air into electrical signals. These electrical signals will be further processed and converted into digital signals for analysis and recognition by a computer or a processing system. Among them, the microphone converts sound waves into electrical signals by sensing their vibrations. At this time, the signal is a continuous analog signal. After the sound signal is acquired, the system usually converts it into a digital signal through an analog - to - digital converter (ADC). Once the user's voice information is collected and converted into a digital signal, the next tasks are to process, identify, and analyze these voice signals, and finally implement functions such as text output, semantic understanding, and translation.
[0023] Specifically, considering that multilingual translation is to convert the information of one language into another through language processing technology, so as to achieve communication and understanding between different languages. With the continuous development of artificial intelligence and natural language processing technologies, speech translation technology has been widely used in many fields such as simultaneous interpretation and foreign language learning. Taking simultaneous interpretation as an example, speech translation technology can translate the speaker's language into the target language in real time, greatly facilitating cross-language communication. However, there are still some challenges in existing speech translation technologies, especially in terms of translation accuracy. During the translation process, especially in the stage of converting speech into text, recognition errors often occur, which may affect the final translation quality. Therefore, in the technical solution of this application, by obtaining the user speech information collected by the sound sensor and combining deep learning technology, a real-time translation result is generated, thereby reducing the error in speech recognition and improving the accuracy of multilingual real-time translation interaction to ensure the high quality of the final translation result.
[0024] In the above-mentioned artificial intelligence-based multilingual real-time translation interaction system 100, the user speech information extraction module 120 is used to extract the semantic feature vector associated with the user communication text information and the logarithmic mel spectrogram feature vector of the user communication speech from the user speech information collected by the sound sensor. It should be understood that the speech signal itself is a continuous sound wave waveform, containing rich language information, but this information needs to be converted into a more structured and processable form at the level of machine understanding. The process of extracting the semantic feature vector and the logarithmic mel spectrogram feature vector is exactly to convert the speech signal into a format that the computer can understand, so as to perform subsequent language processing and analysis. By combining these two feature vectors, the system can not only understand the physical characteristics of speech (such as the audio characteristics of speech), but also grasp the semantic content of speech. Such multimodal feature fusion helps to improve the accuracy of speech understanding, especially in complex multi-task speech processing systems, such as application scenarios like speech translation and intelligent voice assistants.
[0025] Figure 2 It is a block diagram of the user speech information extraction module in the artificial intelligence-based multilingual real-time translation interaction system according to an embodiment of the present application. As Figure 2 shown, in a specific embodiment of the present application, the user speech information extraction module 120 includes: a user speech information text feature extraction unit 121, which is used to perform text feature extraction on the user speech information collected by the sound sensor to obtain the semantic feature vector associated with the user communication text information; a user speech information audio feature extraction unit 122, which is used to perform audio feature extraction on the user speech information collected by the sound sensor to obtain the logarithmic mel spectrogram feature vector of the user communication speech.
[0026] It should be understood that text feature extraction can extract meaningful semantic information from speech, facilitating subsequent language understanding and analysis. Specifically, first, the user speech signal collected by a sound sensor (such as a microphone) is preprocessed and then speech recognition technology (such as an end-to-end speech recognition model based on a deep neural network) is used to convert the speech signal into text. This process involves converting phonemes, syllables, words, etc. in the audio signal into corresponding characters. After obtaining the text of the speech transcription, text feature extraction needs to be carried out next to generate semantic feature vectors. Semantic feature vectors are the conversion of text into a high-dimensional vector representation through natural language processing technology, and this vector captures the semantic information in the text. Specifically, modern text feature extraction methods usually rely on word embeddings (such as Word2Vec, GloVe) or pre-trained language models (such as BERT, GPT) to represent the text as a vector. These models can map each word, phrase or sentence to a vector space of a fixed dimension by learning the context relationships in a large amount of text data, and each dimension in the vector reflects a certain semantic feature. These semantic feature vectors can not only express the semantics of words but also retain the context relationships, so they can more accurately understand the user's communication content. In this way, the machine can understand the deep meaning of the speech transcription text, rather than just the literal words.
[0027] Furthermore, considering that the original speech signal is a continuous analog waveform, it is very difficult for a computer to directly process such an original signal. The key information contained in the speech signal is mainly distributed in the frequency and time domains. Through audio feature extraction, spectral features can be extracted from it, and further, the spectral characteristics can be adjusted by the Mel scale to be closer to the auditory perception mode of the human ear. The log Mel spectrogram is a commonly used audio feature representation. It can convert the speech signal into a feature vector of a fixed dimension, effectively reducing the irrelevant noise components in the speech signal while retaining the main information related to the speech content, facilitating subsequent analysis and processing. In a specific embodiment of the present application, first, the original speech signal collected by the sound sensor is preprocessed. This step includes removing noise, removing the DC component, enhancing the clarity of the signal, etc., making the signal more suitable for subsequent analysis. Then, in order to capture the local time-frequency features of the speech, the speech signal is usually segmented into short-time frames (usually 20 - 40 ms), and a window function (such as a Hamming window) is applied to each frame to reduce the distortion at the frame edges. Next, after each frame is subjected to the Fourier transform (or short-time Fourier transform, STFT), the spectral representation of the frame is obtained. The Fourier transform converts the signal in the time domain into a signal in the frequency domain, revealing the frequency components of the signal. Further, the Mel spectrogram is processed by a Mel scale filter bank. The Mel scale is based on the auditory perception characteristics of the human ear, with higher resolution in the low-frequency part and lower resolution in the high-frequency part, which conforms to the auditory characteristics of humans. The Mel filter bank weights the spectrum, making the spectrum more consistent with the auditory perception of the human ear. Finally, after performing a logarithmic transformation on the Mel spectrogram, the dynamic range of the signal can be effectively compressed, and the recognizability of low-amplitude signals can be enhanced. The log Mel spectrogram feature (Log-MelSpectrogram) retains the frequency characteristics of the speech signal in this way and can better represent information such as speech patterns and timbre changes in the speech. In this way, the spectral information of each frame of speech is converted into a log Mel spectrogram feature vector. The feature matrix composed of multiple consecutive frames constitutes the complete audio feature representation. By performing audio feature extraction on the user speech information collected by the sound sensor, the obtained log Mel spectrogram feature vectors not only effectively represent the frequency characteristics of the speech signal but also consider the auditory perception mode of humans. These feature vectors can help the speech processing system extract important information related to the speech content from the speech signal, providing strong support for subsequent tasks such as speech recognition, speech synthesis, and emotion analysis.
[0028] In a specific embodiment of the present application, the user speech information text feature extraction unit 121 includes: extracting user communication text information from the user speech information collected by the sound sensor; passing the user communication text information through a user communication text information bidirectional long short-term memory neural network to obtain a plurality of user communication text information semantic feature vectors; two-dimensionally arranging the plurality of user communication text information semantic feature vectors into a user communication text information two-dimensional semantic feature matrix; passing the user communication text information two-dimensional semantic feature matrix through a user communication text information two-dimensional semantic feature filter to obtain the user communication text information associated semantic feature vector.
[0029] It should be understood that extracting accurate text information from the user speech signal collected by the sound sensor enables the speech recognition system to convert speech into machine-understandable text. Further, passing the user communication text information through a bidirectional long short-term memory (Bi-LSTM) neural network to obtain a plurality of semantic feature vectors aims to deeply understand the semantic information in the text, especially long-distance dependencies and context information. Among them, the bidirectional LSTM network can learn from both the front and back directions of the text simultaneously. Compared with the unidirectional LSTM, it can capture the bidirectional information in the context more comprehensively, thereby improving the accuracy of semantic understanding. Specifically, the LSTM can effectively remember and forget information through its gating mechanism, solving the problem of gradient disappearance in traditional RNNs when processing long sequences. In the framework of the bidirectional LSTM, the text sequence is first encoded in two directions, from left to right and from right to left, respectively, to obtain the forward and backward hidden state vectors. Then, by concatenating or weighted summing these two directions of feature vectors, a more comprehensive text representation is obtained, reflecting the temporal dependence and context relationship in the text. Finally, these feature vectors encoded by the bidirectional LSTM can accurately represent the semantic information in the text, including the relationship between words, syntactic structure, and context features, providing strong support for subsequent tasks such as text classification, sentiment analysis, or dialogue understanding. In this way, the user's communication text information can not only be effectively processed and represented, but also capture the deep semantics and context of the text, thereby enhancing the accuracy of the model's semantic understanding.
[0030] Furthermore, the semantic feature vectors of multiple user communication text messages are arranged two-dimensionally into a two-dimensional semantic feature matrix of user communication text messages, aiming to structure the semantic feature information of each user's communication for further analysis and processing. Among them, the semantic feature vector is usually a high-dimensional vector generated by a neural network model (such as Bi-LSTM), representing the deep semantic information of the user text. In many natural language processing tasks, especially in multi-turn conversations and text analysis, it is necessary to integrate and arrange these individual semantic feature vectors so that they can present an organized structure, facilitating further learning and reasoning by subsequent models. In this way, arranging multiple feature vectors two-dimensionally into a matrix can concentrate the communication content of each user in a unified structure for subsequent processing. Specifically, each row can represent the communication text features of a user, and each column represents different dimensions of the text features. In this way, the semantic information of the text can be clearly presented in the matrix, facilitating tasks such as matrix operations, clustering analysis, or feature extraction. In this way, by taking the semantic feature vector of each user as a row of the matrix and successively combining the feature vectors of all users into a large two-dimensional matrix, the sequential relationship and context information between different communication rounds can be retained. This method can not only effectively capture the communication content of each user but also retain the semantic information of multi-turn conversations, helping the model to perform more accurate semantic analysis.
[0031] Specifically, processing the two-dimensional semantic feature matrix of user communication text information through a two-dimensional semantic feature filter can extract and screen features from the input matrix through the filter, obtaining a semantic feature vector with stronger task relevance, thereby improving the expressiveness and accuracy of the model. Among them, each row in the two-dimensional semantic feature matrix represents a semantic feature vector of user communication, and the columns represent different dimensions of these vectors. In practical applications, the original semantic feature matrix may contain some redundant or irrelevant features. Therefore, processing through a two-dimensional semantic feature filter can effectively remove irrelevant or noisy information and retain the relevant semantic features that are important for the task. Specifically, considering that each round of user communication may involve different topics or contexts, and some features may not be important in the current task. Through the filter, semantic features with a higher degree of relevance to the current task can be screened out, reducing noise and irrelevant information, and enhancing the precision and robustness of the model. Moreover, after removing irrelevant features through the filter, the dimension of the two-dimensional semantic feature matrix will be correspondingly reduced, thereby reducing the complexity of subsequent calculations. This can accelerate the training and inference speed, especially when dealing with large-scale data, significantly improving the computational efficiency. Among them, the filter usually uses methods such as convolution operations, pooling operations, or attention mechanisms to weight or select on each dimension or feature of the matrix. Then, the convolution kernel slides on the matrix to perform weighted summation on local regions in the matrix, thereby extracting more representative features. The parameters of the convolution kernel can be automatically learned from the data, focusing on capturing features that are helpful for the task. Next, the pooling layer reduces the dimension of the data by selecting the maximum or average value in the feature matrix and enhances the expression of important features. The pooling operation can help the model suppress noise and improve the sensitivity to important features. Immediately afterwards, the attention mechanism is used to weight different parts of the two-dimensional feature matrix, enabling the model to "focus" on the parts that are most critical to the current task and ignoring other irrelevant information. This way can more intelligently select features with high relevance. Through these processes, the two-dimensional semantic feature filter can extract refined and task-closely related associated semantic feature vectors from the original two-dimensional semantic feature matrix. Specifically, each layer of the two-dimensional semantic feature filter for user communication text information performs convolution processing, mean pooling processing based on the local feature matrix, and non-linear activation processing on the input data respectively during the forward pass of the layer to output the associated semantic feature vector of the user communication text information by the last layer of the two-dimensional semantic feature filter for user communication text information, where the input of the two-dimensional semantic feature filter for user communication text information is the two-dimensional semantic feature matrix of the user communication text information.
[0032] In a specific embodiment of the present application, the user voice information audio feature extraction unit 122 includes: extracting a user communication speech logarithmic Mel spectrogram from the user voice information collected by the sound sensor; performing feature encoding on the user communication speech logarithmic Mel spectrogram to obtain the user communication speech logarithmic Mel spectral feature vector.
[0033] It should be understood that a speech signal is a time-domain signal, containing various information such as pitch, speech rate, and tone. However, this information is relatively complex in the time domain and difficult to directly use for analysis. Frequency-domain analysis can better reflect the perceptual characteristics of speech. Among them, the Mel spectrogram is a spectral representation method designed based on the perceptual characteristics of the human auditory system. The human ear is more sensitive to low frequencies and less sensitive to high frequencies, and the Mel frequency scale is a way to simulate the frequency response of the human ear. Therefore, the Mel spectrogram can better capture the perceptual characteristics of the speech signal and retain important information of the speech. Moreover, considering that the original Mel spectrogram often shows a very large numerical range, and the amplitude of the high-frequency part is usually small. Through logarithmic transformation, these numerical ranges can be effectively compressed, reducing the dynamic range of the data, so that the energy distribution of the speech signal more conforms to the perceptual law, enabling the logarithmic Mel spectrogram to highlight important speech features and reduce the noise components in the signal. Specifically, since the speech signal is time-varying, in order to capture the spectral features of the signal in a short time, the speech signal needs to be framed according to a certain frame length (such as 25 ms), and there may be a certain overlap (such as 50% overlap) between each frame to ensure the continuity of time-frequency features. Then, perform a short-time Fourier transform on each frame to convert the time-domain signal into a frequency-domain signal. This process converts the speech signal from a time-domain representation to a frequency-domain representation, which can reflect the spectral information of the signal. Then, in the frequency domain, use a Mel filter bank to process the spectrum of each frame. The Mel filter bank consists of a group of filters, and the frequency distribution of these filters follows the Mel scale to simulate the frequency response of the human ear. By passing the spectrum through the Mel filter bank, a Mel spectrogram can be obtained, which reflects the distribution of the signal on the Mel frequency scale. Finally, apply logarithmic transformation to process the Mel spectrogram. Perform logarithmic transformation on the energy values of each Mel band, which can reduce the influence of the high-frequency part and compress the dynamic range of the data to obtain a logarithmic Mel spectrogram. In this way, a logarithmic Mel spectrogram can be extracted from the original speech signal. It is a two-dimensional image, where the horizontal axis usually represents time and the vertical axis represents Mel frequency, and each point in the image represents the energy value corresponding to the time and frequency. The logarithmic Mel spectrogram provides a more intuitive representation of speech features than the original speech signal and can be effectively applied in subsequent speech recognition, sentiment analysis, speaker recognition, and other tasks.
[0034] Furthermore, the log Mel spectrogram, as a common method for representing speech features, can effectively capture the spectral information and temporal variation features of speech signals, and is particularly suitable for speech recognition and speech understanding tasks. In the log Mel spectrogram, the speech signal is transformed into a two-dimensional representation of frequency and time through a series of preprocessing steps (such as short-time Fourier transform, Mel frequency filter bank, logarithmic transform, etc.), which can reflect the pitch, speech rate, tone, and pronunciation features of the speech. However, the original log Mel spectrogram usually has high dimensionality and redundant information, and directly using it for further processing may lead to low computational efficiency and limited impact on subsequent tasks. Therefore, it is very important to transform it into a more concise and representative feature vector through feature encoding. The process of feature encoding is usually carried out through deep learning models (such as convolutional neural network CNN, long short-term memory network LSTM, etc.). In this process, the feature encoder will learn how to extract key information from the Mel spectrogram, eliminate irrelevant noise, and compress it into a more distinguishable and expressive feature vector. These encoded feature vectors contain the most important speech features in the speech signal, such as the timbre, tone, emotion of the speech, and the rhythm of the language, thus being able to effectively represent the speech content.
[0035] In the technical solution of this application, feature encoding the user communication speech log Mel spectrogram to obtain the user communication speech log Mel spectral feature vector includes: passing the user communication speech log Mel spectrogram through a user communication speech log Mel spectral feature extractor based on a convolutional neural network model to obtain a user communication speech log Mel spectral feature map; performing max pooling on the user communication speech log Mel spectral feature map to obtain the user communication speech log Mel spectral feature vector.
[0036] It should be understood that the logarithmic Mel spectrogram has already transformed the original speech signal into a frequency-domain representation that can reflect speech features. However, these features are still at a low level and contain a large amount of redundant and local information. Processing the logarithmic Mel spectrogram of user communication speech through a feature extractor based on a convolutional neural network (CNN) model can further extract richer and higher-level speech features from the spectrogram. Among them, the convolutional neural network can effectively extract more abstract and discriminative high-level features from these low-level spectral features through a series of convolutional layers and pooling layers. This process can help the model better capture the time-frequency relationship in the speech signal, the pitch, intonation, speech rate of the speech, and the features of the speaker, reduce the sensitivity to noise, and enhance the robustness of the speech recognition system. Specifically, the convolutional neural network extracts local features on the spectrogram by sliding the convolutional kernel, and these local features can capture the local time-frequency structure in the speech signal. Subsequently, the pooling layer reduces the dimensionality of the convolutional result, retains important feature information, compresses the computational amount, and reduces redundancy. After these convolutional and pooling operations, the obtained logarithmic Mel spectrogram feature map of user communication speech contains more refined and representative features. Therefore, as a feature extractor, the convolutional neural network can help extract a highly recognizable feature map from the original logarithmic Mel spectrogram and enhance the system's ability to understand complex speech patterns. Specifically, each layer of the logarithmic Mel spectrogram feature extractor of user communication speech based on the convolutional neural network model performs the following operations on the input data during the forward pass of the layer: performing convolutional processing on the input data based on the convolutional kernel to generate a convolutional feature map; performing global average pooling processing on the convolutional feature map based on the feature matrix to generate a pooling feature map; and non-linearly activating the feature values at each position in the pooling feature map to generate an activation feature map; where the output of the last layer of the logarithmic Mel spectrogram feature extractor of user communication speech based on the convolutional neural network model is the logarithmic Mel spectrogram feature map of user communication speech, the input of the second layer to the last layer of the logarithmic Mel spectrogram feature extractor of user communication speech based on the convolutional neural network model is the output of the previous layer, and the input of the logarithmic Mel spectrogram feature extractor of user communication speech based on the convolutional neural network model is the logarithmic Mel spectrogram of user communication speech.
[0037] Furthermore, max pooling is a common dimensionality reduction operation that reduces the spatial dimension of data by selecting the maximum value within a local region while retaining focus on the most important features in the input data. In speech processing tasks, the pooling operation helps improve the robustness of the model, reduce overfitting, and enhance computational efficiency. Specifically, the max pooling operation is typically performed on the time and frequency dimensions of the log Mel spectrogram feature map. First, the feature map is usually a two-dimensional matrix where the horizontal axis represents time, the vertical axis represents Mel frequency, and each element corresponds to the energy value at a specific time and frequency. When performing pooling, a window of a specific size (e.g., 2x2 or 3x3) slides over the feature map, and for each window, the maximum value within it is selected as the representative of that region. In this way, max pooling not only reduces the size of the data but also helps the network focus on important features, thereby enhancing the model's ability to recognize key information. Among them, the pooling window slides over the feature map with a stride. For each 2x2 region, the maximum value within it is selected. In this way, after pooling, the size of the feature map is reduced by half, and the maximum value of each pooling region contains the most prominent features of that region. Finally, the pooled feature map is flattened to obtain a one-dimensional feature vector that represents the core information of the user's communication speech. In this way, by extracting the most prominent local features, the data dimension can be reduced and the stability of the model can be enhanced.
[0038] In the above-mentioned multi-language real-time translation interaction system 100 based on artificial intelligence, the real-time translation result generation module 130 is used to generate a real-time translation result based on the semantic feature vector associated with the user communication text information and the log Mel spectrogram feature vector of the user communication speech. It should be understood that a user's communication usually includes two forms: speech and text, which respectively provide rich language and pronunciation information. However, a single text or speech information may not be able to fully capture the deep meaning and context of the language, which may cause misunderstandings or ambiguities. Therefore, it is necessary to combine these two types of information for comprehensive analysis and translation. Among them, the semantic feature vector of the text provides the basic structure and grammar information of the language, while the log Mel spectrogram feature vector of the speech complements the pitch, rhythm, and emotional information. Especially in cases where the accent, emotion, or context changes significantly, these speech features are crucial for improving the naturalness and accuracy of the translation. By fusing these two vectors, the model can effectively associate the text and speech information to form a comprehensive understanding model. Finally, the fused feature vector can be input into a machine translation model to generate a real-time translation result. Such a multi-modal fusion translation system has higher accuracy and adaptability than traditional single-modal translation systems, especially performing more excellently in complex language communication environments.
[0039] Figure 3The block diagram of the real-time translation result generation module in the multi-language real-time translation interaction system based on artificial intelligence according to an embodiment of the present application. As Figure 3 shown, in a specific embodiment of the present application, the real-time translation result generation module 130 includes: a user communication language information feature fusion unit 131, configured to fuse the user communication text information associated semantic feature vector and the user communication speech log Mel spectrogram feature vector to obtain a user communication language information feature vector; a user communication language information feature optimization unit 132, configured to perform node decomposition optimization based on a topological space constraint structure on the user communication language information feature vector to obtain an optimized user communication language information feature vector; and a user communication language translation result generation unit 133, configured to obtain a real-time translation result by passing the optimized user communication language information feature vector through a generator.
[0040] It should be understood that text information provides the semantic content of language, while voice information contains important features such as pronunciation, tone, and emotion. Relying solely on text or voice information may have limitations, and fusing the features of both can achieve more accurate speech recognition, emotion analysis, and semantic understanding. Especially in application scenarios such as real-time translation and voice assistants, it can significantly improve the performance of the system.
[0041] Specifically, considering that in the process of multi-modal feature fusion, the processing methods of text information and voice information are quite different. Among them, text information extracts semantic feature vectors through a bidirectional long short-term memory neural network. In this process, it focuses on the grammar, semantics, and context understanding of language, emphasizing the logical relationship between words and the semantic structure of sentences. While voice information extracts spectral features through a convolutional neural network, and these features reflect the time-frequency characteristics of the voice signal, mainly capturing physical characteristics such as pitch and timbre of the audio signal. The feature spaces of the two are quite different. After directly fusing these two types of features, although certain comprehensive features of speech and text can be obtained, since they originate from different modalities and the information representation methods are different, the fused user communication language information feature vector may lose the original high-dimensional feature details, resulting in the inability to fully retain the rich information of the original voice and text data. In addition, although fusion can enhance the expression of information to a certain extent, due to the weak internal connection between the two types of information, especially the complex and not fully synchronous speech-semantic conversion between speech and text, the fused user communication language information feature vector may not effectively strengthen the internal connection between different data points. This weakness in the internal connection limits the accuracy and robustness of the fused features in semantic understanding and translation tasks, and cannot fully reflect the joint semantic features of speech and text. Therefore, in the technical solution of the present application, the user communication language information feature vector is subjected to node decomposition optimization based on a topological space constraint structure to obtain an optimized user communication language information feature vector.
[0042] Among them, performing node decomposition optimization on the user communication language information feature vector based on the topological space constraint structure to obtain an optimized user communication language information feature vector includes: extracting the target parameter matrix of the generator; performing node decomposition on the target parameter matrix in units of row vectors to obtain a set of target parameter node coding vectors; using each target parameter node coding vector in the set of target parameter node coding vectors as a wandering topological space, respectively performing topological space constraints on the user communication language information feature vector to obtain a set of constrained user communication language information feature vectors; and calculating the position-wise mean vector of the set of constrained user communication language information feature vectors to obtain the optimized user communication language information feature vector.
[0043] Among them, using each target parameter node coding vector in the set of target parameter node coding vectors as a wandering topological space, respectively performing topological space constraints on the user communication language information feature vector to obtain a set of constrained user communication language information feature vectors includes: multiplying the user communication language information feature vector by the transposed vector of the target parameter node coding vector, and then calculating the natural exponential function value of the multiplication result to obtain a weighted exponential response weight; calculating the Euclidean distance between the user communication language information feature vector and the target parameter node coding vector to obtain a node coding distance value; performing a dot product on the node coding distance value and the user communication language information feature vector, and calculating the natural exponential function value for each eigenvalue of the dot product vector to obtain a user communication language information distance-guided exponential feature vector; multiplying the weighted exponential response weight and the user communication language information distance-guided exponential feature vector to obtain the constrained user communication language information feature vector.
[0044] Specifically, performing node decomposition optimization on the user communication language information feature vector based on the topological space constraint structure to obtain an optimized user communication language information feature vector is expressed by the formula:
[0045] W = [w 1 , w 2 ,..., w i ,..., w n T
[0046]
[0047]
[0048] Among them, W represents the target parameter matrix, w 1 , w 2 , w i , w n The first, second, i-th, and n-th target parameter node encoding vectors representing the set of target parameter node encoding vectors, T represents the transpose of a vector, v 1 represents the user communication language information feature vector, represents matrix multiplication, ⊙ represents element-wise multiplication, D(v 1 , w i ) represents calculating the Euclidean distance between vector v 1 and vector w i , V 1-i represents the i-th constrained user communication language information feature vector in the set of constrained user communication language information feature vectors, n represents the total number of the set of constrained user communication language information feature vectors, V i ’ represents the optimized user communication language information feature vector.
[0049] In the technical solution of this application, node decomposition optimization based on a topological space constraint structure is performed on the user communication language information feature vector. This process first extracts the key parameters for decision-making from the trained generator, and these parameters form a matrix form in a high-dimensional space, where each row represents the weights or influencing factors in different dimensions. Through the target parameter matrix, the position and shape of the model decision boundary can be insight, and then it can be inferred which input features are the most critical for the prediction result.
[0050] Next, the target parameter matrix is node-decomposed with row vectors as units to obtain a set of target parameter node encoding vectors. Here, each row vector is regarded as a node in graph theory, which means that each group of parameters is now regarded as an entity with potential connectivity. This transformation allows the application of methods from graph theory and network science to explore the interactions between features. The node encoding vector not only carries the information of the original parameters but also implies the knowledge about the topological structure of the whole system. The node decomposition further reveals the internal connection pattern or structure of the data, enabling the optimized feature vector to better adapt to the new task requirements.
[0051] Then, using each target parameter node encoding vector in the set of the target parameter node encoding vectors as a random walk topological space, topological space constraints are respectively imposed on the user communication language information feature vectors to obtain a set of constrained user communication language information feature vectors. Using the topological space defined by the node encoding vector for "random walk" is actually simulating an exploratory process, aiming to find feature transformations that can best preserve the characteristics of the original data structure. Each step determines the next position according to the probability distribution of the current state, and the topological space constraint ensures that even in different contexts, the feature representation still retains certain invariance. At the same time, it can also promote cross-domain transfer learning because it emphasizes the generally existing relationships between features rather than domain-specific details. In this way, the reconstruction of the feature space is achieved, making the optimized feature vectors more compact and having better generalization ability.
[0052] Finally, calculate the position-wise mean vector of the set of the constrained user communication language information feature vectors to obtain the optimized user communication language information feature vector. Calculating the mean vector is a statistical aggregation method for integrating the optimal solutions from multiple perspectives. The idea behind this step is to reduce the bias caused by a single estimate by fusing the information provided by different sample points. The averaging process is equivalent to performing a Soft Voting, enhancing the expressiveness of the common features and making the optimized feature vectors more stable and reliable.
[0053] Furthermore, the core purpose of obtaining the real-time translation result by passing the optimized user communication language information feature vector through the generator is to utilize the capabilities of the generative model to generate accurate, natural, and contextually appropriate translations under multi-modal inputs (such as text and speech). Among them, the generator usually refers to a neural network-based model, especially models like the sequence-to-sequence (Seq2Seq) model, the Transformer model, or the autoregressive generative model. These models are very good at processing sequential data (such as language) and can generate corresponding output sequences (i.e., translation results) from the input feature vectors. In the multi-modal translation task, the role of the generator is to generate the translation result of the target language based on the fused language information feature vector. Specifically, in the technical solution of this application, first, the user's language information feature vector (including the semantic features of the text and the log mel spectrogram features of the speech) is used as the input of the generator. These feature vectors can convey comprehensive language information to the model, including semantic content and non-linguistic pronunciation information. Then, the encoder part of the generator receives the input feature vectors and, through multiple levels of processing (such as through self-attention mechanisms, convolutional layers, LSTM, etc.), understands and models the input information. The purpose of the encoder is to extract the potential semantic representation from the input features and construct an abstract representation of the source language. Next, the output of the encoder is passed to the decoder part of the generator. The decoder gradually generates the translation of the target language based on the abstract semantic representation output by the encoder and in combination with the grammar rules of the target language. The generator can dynamically adjust the content of each word or sentence generated according to the context information to make it more natural and fluent. Finally, since the generator is usually autoregressive (each step of generation depends on the output of the previous step), it can gradually construct a complete translated sentence based on the generated part and the input language information. In this way, the generator can process the user input in real time and quickly generate the translation result. In this way, the generator can utilize the multi-modal information contained in the optimized user communication language information feature vector to make the translation more emotional and tonal while ensuring translation accuracy, improving the naturalness and fluency of the translation result.
[0054] In summary, the embodiments of this application first obtain the user voice information collected by the sound sensor, then use deep learning technology to perform feature extraction and correlation analysis on it, and finally generate a real-time translation result through the generator, thereby reducing speech recognition errors, improving the accuracy of multi-language real-time translation interaction, and ensuring the quality of the final translation result.
[0055] As described above, the artificial intelligence-based multilingual real-time translation interaction system 100 according to the embodiments of the present application can be implemented in various terminal devices. In one example, the artificial intelligence-based multilingual real-time translation interaction system 100 can be integrated into the terminal device as a software module and / or a hardware module. For example, the artificial intelligence-based multilingual real-time translation interaction system 100 can be a software module in the operating system of the terminal device, or can be an application program developed for the terminal device; of course, the artificial intelligence-based multilingual real-time translation interaction system 100 can also be one of the many hardware modules of the terminal device.
[0056] Alternatively, in another example, the artificial intelligence-based multilingual real-time translation interaction system 100 and the terminal device can also be separate devices, and the artificial intelligence-based multilingual real-time translation interaction system 100 can be connected to the terminal device through a wired and / or wireless network and transmit interaction information in accordance with a predefined data format.
[0057] Figure 4 FIG. is a flowchart of an artificial intelligence-based multilingual real-time translation interaction method according to an embodiment of the present application. As Figure 4 shown, the artificial intelligence-based multilingual real-time translation interaction method according to an embodiment of the present application includes: S110, obtaining user speech information collected by a sound sensor; S120, extracting a user communication text information associated semantic feature vector and a user communication speech logarithmic mel frequency spectrum feature vector from the user speech information collected by the sound sensor; S130, generating a real-time translation result based on the user communication text information associated semantic feature vector and the user communication speech logarithmic mel frequency spectrum feature vector.
[0058] Here, those skilled in the art can understand that the specific operations of each step in the above artificial intelligence-based multilingual real-time translation interaction method have been described in detail in the description of the artificial intelligence-based multilingual real-time translation interaction system above with reference to Figures 1 to 3 and thus, the repeated description thereof will be omitted.
Claims
1. A multi-language real-time translation interactive system based on artificial intelligence, characterized in that: include: A user voice information acquisition module is used to acquire user voice information collected by a sound sensor; A user voice information extraction module, used to extract a user communication text information associated semantic feature vector and a user communication voice logarithmic Mel spectrum feature vector from the user voice information collected by the sound sensor; The real-time translation result generating module is used to generate a real-time translation result based on the semantic feature vector associated with the user communication text information and the logarithmic Mel spectrum feature vector of the user communication speech.
2. The multi-language real-time translation interactive system based on artificial intelligence according to claim 1 is characterized in that: The user voice information extraction module comprises: A user voice information text feature extraction unit, used for performing text feature extraction on the user voice information collected by the sound sensor to obtain a semantic feature vector associated with the user communication text information; The user voice information audio feature extraction unit is used to extract audio features from the user voice information collected by the sound sensor to obtain a logarithmic Mel frequency spectrum feature vector of the user communication voice.
3. The multi-language real-time translation interactive system based on artificial intelligence according to claim 2 is characterized in that: The user voice information text feature extraction unit comprises: Extracting user communication text information from the user voice information collected by the sound sensor; Passing the user communication text information through a user communication text information bidirectional long short-term memory neural network to obtain a plurality of user communication text information semantic feature vectors; Arranging the plurality of user communication text information semantic feature vectors in two dimensions into a user communication text information two-dimensional semantic feature matrix; The two-dimensional semantic feature matrix of the user communication text information is passed through a two-dimensional semantic feature filter of the user communication text information to obtain a semantic feature vector associated with the user communication text information.
4. The multi-language real-time translation interactive system based on artificial intelligence according to claim 3 is characterized in that: The user voice information audio feature extraction unit comprises: Extracting a user communication voice logarithmic Mel-frequency spectrum graph from the user voice information collected by the sound sensor; The logarithmic Mel-frequency spectrum diagram of the user communication speech is feature encoded to obtain a feature vector of the logarithmic Mel-frequency spectrum of the user communication speech.
5. The multi-language real-time translation interactive system based on artificial intelligence according to claim 4 is characterized in that: Performing feature encoding on the logarithmic Mel spectrum graph of the user communication speech to obtain a feature vector of the logarithmic Mel spectrum of the user communication speech includes: The user communication speech logarithmic Mel spectrum graph is passed through a user communication speech logarithmic Mel spectrum feature extractor based on a convolutional neural network model to obtain a user communication speech logarithmic Mel spectrum feature graph; The maximum value of the Mel-frequency spectrum feature map of the user communication speech is pooled to obtain the Mel-frequency spectrum feature vector of the user communication speech.
6. The multi-language real-time translation interactive system based on artificial intelligence according to claim 5 is characterized in that: The real-time translation result generating module comprises: A user communication language information feature fusion unit, used to fuse the user communication text information associated semantic feature vector and the user communication speech logarithmic Mel spectrum feature vector to obtain a user communication language information feature vector; A user communication language information feature optimization unit, configured to perform node-based decomposition optimization on the user communication language information feature vector based on a topological space constraint structure to obtain an optimized user communication language information feature vector; The user communication language translation result generating unit is used for passing the optimized user communication language information feature vector through a generator to obtain a real-time translation result.
7. The multi-language real-time translation interactive system based on artificial intelligence according to claim 6 is characterized in that: The user communication language information feature optimization unit includes: Extract the target parameter matrix of the generator; Decomposing the target parameter matrix into nodes in units of row vectors to obtain a set of target parameter node encoding vectors; Taking each target parameter node encoding vector in the set of target parameter node encoding vectors as a walking topological space, topological space constraints are respectively performed on the user communication language information feature vectors to obtain a set of constrained user communication language information feature vectors; and The position-wise mean vector of the set of the constrained user communication language information feature vectors is calculated to obtain the optimized user communication language information feature vector.
8. The multi-language real-time translation interactive system based on artificial intelligence according to claim 7 is characterized in that: Taking each target parameter node encoding vector in the set of target parameter node encoding vectors as the walking topological space, topological space constraints are respectively performed on the user communication language information feature vectors to obtain a set of constrained user communication language information feature vectors, including: After multiplying the user communication language information feature vector and the transposed vector of the target parameter node encoding vector, a natural exponential function value of the multiplication result is calculated to obtain a weighted exponential response weight; Calculating the Euclidean distance between the user communication language information feature vector and the target parameter node encoding vector to obtain a node encoding distance value; Performing a dot product of the node coding distance value and the user communication language information feature vector, and calculating a natural exponential function value for each feature value of the vector after the dot product to obtain a user communication language information distance guidance index feature vector; The weighted index response weight and the user communication language information distance guidance index feature vector are multiplied to obtain a constrained user communication language information feature vector.
9. A multi-language real-time translation interactive method based on artificial intelligence, characterized in that: include: Acquire user voice information collected by the sound sensor; Extracting a semantic feature vector associated with user communication text information and a logarithmic Mel spectrum feature vector of user communication voice from the user voice information collected by the sound sensor; Based on the semantic feature vector associated with the user communication text information and the logarithmic Mel-spectrogram feature vector of the user communication speech, a real-time translation result is generated.
10. The multi-language real-time translation interactive method based on artificial intelligence according to claim 9, characterized in that: Extracting a semantic feature vector associated with user communication text information and a logarithmic Mel spectrum feature vector of user communication voice from the user voice information collected by the sound sensor includes: Performing text feature extraction on the user voice information collected by the sound sensor to obtain a semantic feature vector associated with the user communication text information; Audio feature extraction is performed on the user voice information collected by the sound sensor to obtain a logarithmic Mel frequency spectrum feature vector of the user communication voice.
Citation Information
Patent Citations
Conference system for realizing synchronous translation
CN110083847A
Multimodal fusion speech translation method, system and equipment
CN118692446A
Simultaneous interpretation intelligent proofreading method and device and storage medium
CN119167949A