A multilingual intelligent real-time translation system and method
By designing a multilingual intelligent real-time translation system, the problems of background noise interference, data disorder and model calculation burden in the prior art are solved, and more accurate speech recognition and more efficient translation quality evaluation and correction are achieved.
Patent Information
- Application Number
- CN202510177307.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-18
AI Technical Summary
The existing translation system failed to effectively eliminate background noise, resulting in inaccurate speech recognition; failed to standardize the processing of multi-dimensional situational data and translation environment data, resulting in disordered data and poor model training effect; failed to limit the order of MFCC feature coefficients, increasing the burden of model calculation; failed to accurately evaluate feature correlation, resulting in limited model performance.
A multilingual intelligent real-time translation system is designed, including data acquisition module, data processing module, multilingual translation module, quality evaluation module, error type evaluation module and error correction module. Multidimensional contextual data is processed by spectral subtraction, data cleaning and standardization of translation environment data, and feature-weighted fusion and model training are used to predict translation errors and correct translation results.
Effectively eliminate background noise and improve the accuracy of speech recognition; standardize the processing of data, improve the effect of model training; through feature correlation evaluation and dynamic adjustment of resolution coefficients, the model performance is optimized, and the translation quality and system response speed are improved.
Smart Images

Figure CN119647491B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent translation technology, and more specifically, to a multi-language intelligent real-time translation system and method. Background Art
[0002] The patent application publication number CN116522960A discloses a multi-language interactive real-time translation terminal and method supporting multiple platforms. A quartz crystal oscillator is used to perform frequency control and frequency selection. According to the real-time translation terminal driver, an integrated Bluetooth chip unit is used to perform two-way information communication between the real-time translation terminal and the multi-platform communication device terminal; through multi-platform input recognition, the multi-platform input process is tracked to obtain multi-platform input information; multi-platform input information of the input language is standardized, and the target language is searched and translated through a multi-language real-time translation engine, and the target language expression order and the input language expression order are analyzed and compared to obtain input translation analysis and comparison data; according to the input translation analysis and comparison data, the target language expression order and the input language expression order are adjusted and determined, and translation reference interactive selection is performed to realize intelligent multi-language reference interactive real-time translation.
[0003] The existing translation systems and methods have the following main problems:
[0004] Failure to consider eliminating background noise will interfere with speech signals, making it difficult for the speech recognition system to accurately recognize speech content and reducing recognition accuracy; Failure to segment, encode, and adjust the position of context data will cause the data to be disordered or inconsistent, leading to difficulties in subsequent feature extraction and model training, and may result in data redundancy, information loss, or poor model training results; Failure to limit the order of MFCC feature coefficients will result in too many features, increasing the computational burden of the model and affecting the response speed of the real-time speech recognition system;
[0005] The lack of precise quantitative means makes it impossible to accurately evaluate the correlation between different features, and thus it is difficult to determine which features have a significant impact on the speech feature dataset; it affects the effectiveness of subsequent model training and optimization; inaccurate correlation evaluation may lead to the selection of wrong features for model training, thereby limiting the performance of the model; failure to fully utilize all relevant features may also lead to a decrease in the generalization ability of the model; failure to consider the dynamic adjustment of the resolution coefficient makes the correlation calculation too rigid and unable to adapt to the characteristics and requirements of different datasets; it reduces the accuracy and reliability of the evaluation and limits the application effect of the model in different scenarios; failure to accurately quantify the correlation and select key features for model training may result in a large amount of resources and time being invested in irrelevant features, resulting in a waste of resources.
[0006] In view of this, the present invention proposes a multi-language intelligent real-time translation system and method to solve the above problems. Summary of the invention
[0007] In order to overcome the above-mentioned defects of the prior art and to achieve the above-mentioned purpose, the present invention provides the following technical solution: a multilingual intelligent real-time translation system, comprising:
[0008] A data collection module, used to collect voice input data, multi-dimensional context data and translation environment data;
[0009] The data processing module is used to pre-process the collected speech input data, multi-dimensional situational data and translation environment data to obtain a speech feature data set, a situational feature data set and a translation environment feature data set; evaluate the impact of the translation environment feature data set and the situational feature data set on the speech feature data set to obtain an impact feature data set; and perform weighted fusion of the impact feature data set and the speech feature data set to obtain a comprehensive feature data set;
[0010] A multilingual translation module is used to obtain a multilingual translation error prediction model based on comprehensive feature data set training; and to predict multilingual translation errors through the multilingual translation error prediction model;
[0011] The quality assessment module is used to compare the predicted multilingual translation error with the preset multilingual translation error threshold to determine whether the translation quality meets the standard; if the translation quality meets the standard, the real-time translation result is output; if the translation quality does not meet the standard, the multilingual intelligent translation terminal issues a warning message and automatically collects multilingual translation error data;
[0012] An error type evaluation module is used to obtain an error type diagnosis model based on multilingual translation error data training, and predict the translation error type through the error type diagnosis model;
[0013] The error correction module is used to perform error correction on real-time translation results that do not meet the translation quality standards according to the translation error type, and output the real-time translation results after error correction; each module is connected to another via wired or wireless means.
[0014] Furthermore, the voice input data is continuous audio data received by the user; the multi-dimensional situational data includes emotion data, contextual data and cultural background data; and the translation environment data includes noise intensity data and device information data.
[0015] Furthermore, the method of preprocessing the collected speech input data, multi-dimensional context data and translation environment data to obtain a speech feature data set, a context feature data set and a translation environment feature data set includes:
[0016] The method for preprocessing speech input data is: using spectrum subtraction to eliminate background noise in speech input data, extracting audio features from the speech input data after noise elimination; dividing continuous audio data into segments, and converting them into text data through a speech recognition model; the specific steps are:
[0017] Performing short-time Fourier transform on the speech input data to convert from the time domain to the frequency domain; estimating the noise spectrum in the silent segment of the speech input data; performing spectrum subtraction denoising operation on the spectrum of the speech input data through the estimated noise spectrum; performing inverse short-time Fourier transform on the denoised spectrum to restore it to the speech signal in the time domain; the restored speech signal is the denoised speech input data;
[0018] The denoised speech input data is framed and divided into frames; multiply each frame by a window function and perform fast Fourier transform to obtain the frequency domain signal of the denoised speech input data; use a set of Mel filters to perform Mel filtering on the frequency domain signal to obtain the output power of each filter; perform logarithmic transformation on the output power of the Mel filter to obtain the logarithmic power spectrum;
[0019] The obtained logarithmic power spectrum is subjected to discrete cosine transform to obtain the MFCC feature coefficient. The specific mathematical formula is: ;in, For the MFCC feature coefficients of order; is the order of the MFCC feature coefficient currently calculated; is the total number of Mel filters; For the The output power of the Mel filter; is the index of the Mel filter; is pi;
[0020] The order of the currently calculated MFCC feature coefficients is limited by the order limiting formula, and the order limiting formula is: ; is the order of the MFCC feature coefficients after restriction; is the logarithmic sum of all Mel filter output powers; and To control the parameter factor of the degree of restriction on the MFCC order; collect the MFCC feature coefficients corresponding to each frame to obtain a feature coefficient data set; input the feature coefficient data set into the Transformer model to convert it into text data, and then obtain the speech feature data set;
[0021] The method for preprocessing the multidimensional context data is as follows: segmenting and encoding the multidimensional context data, and adjusting it through position encoding; inputting the adjusted multidimensional context data into the BERT model for feature extraction, thereby obtaining a context feature data set;
[0022] The method for preprocessing the translation environment data is as follows: data cleaning is performed on the obtained translation environment data, missing values are filled using interpolation, and outliers are detected and removed using clustering algorithms; Z-score standardization is used to convert the translation environment data into a standard normal distribution with a mean of 0 and a standard deviation of 1 to eliminate the dimension effect;
[0023] After the preprocessing in the above steps, a speech feature dataset, a situational feature dataset and a translation environment feature dataset are obtained.
[0024] Furthermore, the method of evaluating the influence of the translation environment feature dataset and the situational feature dataset on the speech feature dataset to obtain the influenced feature dataset includes:
[0025] The speech feature dataset is used as the reference sequence, denoted as ;in, For the Reference value at a time point; is the index of the time point; each feature in the translation environment feature dataset and the context feature dataset is taken as a comparison sequence, recorded as ;in, For the The comparison sequence is in The value at a point in time; is the index of the comparison sequence; is the index of the time point;
[0026] Calculate reference sequence Compare each sequence The absolute difference at each time point is used to obtain the absolute difference matrix; the maximum difference and the minimum difference among all the absolute differences are selected, and the correlation coefficient of each sequence at each time point is calculated according to the absolute difference matrix: ;in, The reference sequence and The comparison sequence is in The degree of correlation at each time point; For the reference sequence and The comparison sequence is in The difference between the time points; is the maximum difference among all absolute differences; is the minimum difference among all absolute differences; is the resolution coefficient;
[0027] The resolution factor is adjusted by the resolution factor adjustment formula Dynamic adjustment is performed, and the resolution coefficient adjustment formula is: ;in, is the resolution coefficient after dynamic adjustment; is the maximum value of the resolution coefficient; is the total number of time points; is the length of the reference sequence and comparison sequence; is a constant factor that controls the effect of the total number of time points on the resolution coefficient; A constant factor to control the effect of the length of the reference sequence and the comparison sequence on the resolution coefficient;
[0028] The grey correlation degree is obtained by calculating the average value of the correlation coefficient of each comparison sequence; the grey correlation degree of each comparison sequence is sorted and the top The features corresponding to the grey relational degrees constitute the influencing feature data set.
[0029] Furthermore, the training method of the multilingual translation error prediction model includes:
[0030] The dataset is divided into a training set, a validation set, and a test set for training the model and evaluating the model performance. The sample set is a subset of the dataset, and each sample set includes a historical comprehensive feature dataset and the corresponding multilingual translation errors.
[0031] Constructing a multilingual translation error prediction model, including an input layer, an RNN layer, a fully connected layer and an output layer; the input layer is used to input a historical comprehensive feature data set, and the output layer is used to output multilingual translation errors; the multilingual translation error prediction model is a recurrent neural network model;
[0032] The mean square error is selected as the loss function to measure the difference between the model prediction result and the true label; the multilingual translation error prediction model is a recurrent neural network model;
[0033] Use the training set to train the multilingual translation error prediction model and use the Adam optimizer to minimize the loss function. Use the validation set to evaluate the performance of the model and measure the performance of the model by calculating the accuracy index.
[0034] The performance of the model is evaluated based on the validation set, and the hyperparameters of the model are tuned until the preset stopping condition is reached to obtain a trained multilingual translation error prediction model; the current comprehensive feature dataset is input into the trained multilingual translation error prediction model to predict the multilingual translation error.
[0035] Furthermore, the method of comparing the predicted multilingual translation error with a preset multilingual translation error threshold to determine whether the translation quality meets the standard includes:
[0036] If the predicted multilingual translation error is less than the preset multilingual translation error threshold, the translation quality is determined to be up to standard;
[0037] If the predicted multilingual translation error is greater than or equal to a preset multilingual translation error threshold, it is determined that the translation quality does not meet the standard.
[0038] Furthermore, the multilingual translation error data refers to translation segments whose translation quality does not meet the standards.
[0039] Furthermore, the training method of the error type diagnosis model includes:
[0040] The data set is divided into a training set, a validation set and a test set; the data set includes multilingual translation error data and corresponding translation error types; an error type diagnosis model is constructed, the input data of the model is historical multilingual translation error data, and the output label of the model is the translation error type; RBF is selected as the kernel function of the model, and the hyperparameters of the model are initialized; the error type diagnosis model is a support vector machine model;
[0041] Use mean absolute error as the loss function to measure the error between the model's predicted value and the actual value; use the training set data to train the model and minimize the loss function through the Adam optimizer; use the validation set to evaluate the performance of the model and tune the model's hyperparameters until the preset stopping condition is reached; use the test set to evaluate the model's performance in the prediction task and input the current multilingual translation error data into the trained error type diagnosis model to obtain the translation error type.
[0042] Furthermore, the method for performing error correction on real-time translation results that do not meet translation quality standards according to the translation error type includes:
[0043] The specific location where the translation error occurs is located according to the obtained translation error type, and the SMT model is used to re-decode the translation error part to generate more accurate translation candidates; the translation probability of all translation candidates is calculated, and the translation candidate with the highest translation probability is selected to replace the translation segment that does not meet the translation quality standard, so as to obtain the real-time translation result after error correction.
[0044] A multilingual intelligent real-time translation method, comprising:
[0045] S1, collecting speech input data, multi-dimensional context data and translation environment data;
[0046] S2, preprocessing the collected speech input data, multi-dimensional situational data and translation environment data to obtain a speech feature data set, a situational feature data set and a translation environment feature data set; evaluating the impact of the translation environment feature data set and the situational feature data set on the speech feature data set to obtain an impact feature data set;
[0047] S3. A multilingual translation error prediction model is obtained by training according to the influencing feature data set; and multilingual translation errors are predicted by the multilingual translation error prediction model;
[0048] S4, comparing the predicted multilingual translation error with the preset multilingual translation error threshold to determine whether the translation quality meets the standard; if the translation quality meets the standard, outputting the real-time translation result; if the translation quality does not meet the standard, the multilingual intelligent translation terminal issues a warning message and automatically collects multilingual translation error data;
[0049] S5, training an error type diagnosis model based on multilingual translation error data, and predicting translation error type data through the error type diagnosis model;
[0050] S6. Error correction is performed on the real-time translation results that do not meet the translation quality standards according to the translation error type data, and the real-time translation results after error correction are output.
[0051] The technical effects and advantages of the multilingual intelligent real-time translation system and method of the present invention are as follows:
[0052] The present invention effectively eliminates the background noise of speech input data through spectrum subtraction, making subsequent speech recognition more accurate; at the same time, the extraction and order restriction of MFCC feature coefficients further enhance the expressiveness and robustness of speech features; for multi-dimensional contextual data, word segmentation, encoding and position adjustment make the data more standardized and structured, which is convenient for subsequent feature extraction and model processing; for translation environment data, data cleaning, missing value filling, outlier detection and standardization effectively eliminate noise and dimensional effects in the data, and improve the reliability and consistency of the data; order restriction can significantly reduce the number of MFCC feature coefficients, thereby reducing the computational cost of the model. This is particularly important for real-time speech recognition systems, which can shorten recognition time and improve user experience;
[0053] By calculating the absolute difference between the reference sequence (speech feature dataset) and each comparison sequence (features in the translation environment feature dataset and the context feature dataset) at each time point, as well as the subsequent correlation coefficient and grey correlation degree, this method can accurately quantify the correlation between different features; it helps to more accurately understand which features have a significant impact on the speech feature dataset; the dynamic adjustment formula of the resolution coefficient takes into account the total number of time points and the length of the reference sequence and the comparison sequence, making the calculation of the correlation more flexible and accurate; the dynamic adjustment can better adapt to the characteristics and needs of different datasets and improve the accuracy and reliability of the evaluation. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 A schematic diagram of the structure of a multi-language intelligent real-time translation system of the present invention;
[0055] Figure 2 The present invention is a flowchart of a multi-language intelligent real-time translation method. DETAILED DESCRIPTION
[0056] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0057] Embodiment 1
[0058] See also Figure 1 As shown, the multi-language intelligent real-time translation system described in this embodiment is applied to a translator. When voice data is transmitted to the translator via Bluetooth hardware, the system performs translation processing, specifically including:
[0059] A data collection module, used to collect voice input data, multi-dimensional context data and translation environment data;
[0060] The data processing module is used to pre-process the collected speech input data, multi-dimensional situational data and translation environment data to obtain a speech feature data set, a situational feature data set and a translation environment feature data set; evaluate the impact of the translation environment feature data set and the situational feature data set on the speech feature data set to obtain an impact feature data set; and perform weighted fusion of the impact feature data set and the speech feature data set to obtain a comprehensive feature data set;
[0061] A multilingual translation module is used to obtain a multilingual translation error prediction model based on comprehensive feature data set training; and to predict multilingual translation errors through the multilingual translation error prediction model;
[0062] The quality assessment module is used to compare the predicted multilingual translation error with the preset multilingual translation error threshold to determine whether the translation quality meets the standard; if the translation quality meets the standard, the real-time translation result is output; if the translation quality does not meet the standard, the multilingual intelligent translation terminal issues a warning message and automatically collects multilingual translation error data;
[0063] An error type evaluation module is used to obtain an error type diagnosis model based on multilingual translation error data training, and predict the translation error type through the error type diagnosis model;
[0064] The error correction module is used to perform error correction on the real-time translation results that do not meet the translation quality standards according to the translation error type, and output the real-time translation results after error correction; each module is connected by wired or wireless means;
[0065] The data acquisition module is connected to the data processing module through Bluetooth hardware, and transmits the collected voice input data, multi-dimensional situational data and translation environment data to the translator implanted with Bluetooth hardware; the translator can receive and process data from the data acquisition module, and translate multiple languages into Chinese interpretations in real time according to different language contexts through built-in models, while intelligently adjusting the context and style of the translation; the translator sends the translated results to the mobile phone App connected to the translator through Bluetooth hardware, and users can view the translated Chinese interpretation text in real time through the App interface, and adjust the translation settings as needed.
[0066] The voice input data is the continuous audio data received by the user; the multi-dimensional situational data includes emotional data, contextual data and cultural background data; the translation environment data includes noise intensity data and device information data.
[0067] The method for preprocessing the collected speech input data, multi-dimensional context data and translation environment data includes:
[0068] The method for preprocessing speech input data is: using spectrum subtraction to eliminate background noise in speech input data, extracting audio features from the speech input data after noise elimination; dividing continuous audio data into segments, and converting them into text data through a speech recognition model; the specific steps are:
[0069] Perform short-time Fourier transform on the speech input data to convert from the time domain to the frequency domain; estimate the noise spectrum in the silent segment of the speech input data, that is, average the background noise in the part without speech components; perform spectrum subtraction denoising operation on the spectrum of the speech input data through the estimated noise spectrum; perform inverse short-time Fourier transform on the denoised spectrum to restore it to the speech signal in the time domain; the restored speech signal is the denoised speech input data;
[0070] The denoised speech input data is framed and divided into frames; multiply each frame by a window function and perform fast Fourier transform to obtain the frequency domain signal of the denoised speech input data; use a set of Mel filters to perform Mel filtering on the frequency domain signal to obtain the output power of each filter; perform logarithmic transformation on the output power of the Mel filter to obtain the logarithmic power spectrum;
[0071] For example: Suppose there is a piece of speech input data containing noise , the sampling rate is 16kHz, and the duration is 2 seconds; therefore, the speech input data contains a total of 16000×2=32000 samples; the signal of 32000 samples is divided into frames of 400 samples (25 milliseconds) to achieve short-term stability; the overlap between frames is 200 samples (50% overlap) for smooth transition; a Hamming window is applied to each frame, and the length of the window function is the same as the frame length; a fast Fourier transform (FFT) is performed on each frame to obtain a frequency domain representation; assuming that the time domain signal of the first frame is subjected to FFT to obtain the following spectrum (partial value): The first 200 milliseconds (first 8 frames) of the audio are silent segments. Calculate the average spectrum of these 8 frames to get the noise spectrum estimate ; Subtract the noise spectrum from the spectrum of each frame, and the denoised spectrum is ; If the value obtained by subtraction is less than 0, it is usually set to 0 to prevent negative spectral values;
[0072] Perform inverse Fourier transform on the denoised spectrum to restore it to the denoised time domain speech signal, and merge all denoised frames into a complete time domain speech signal; divide the denoised time domain speech signal into frames again, and continue to use a frame length of 400 samples and an overlap of 200 samples;
[0073] Apply a set of Mel filters to the denoised spectrum. Assume that 26 Mel filters are used, and the frequency range of the filters covers the frequency range of 0 to 8kHz. Each Mel filter calculates the power of a frequency band. Assume that the output power of the first filter is 100, the second is 120, and so on. Perform a logarithmic transformation on the output power of each Mel filter. Then the logarithmic power of the first filter is log(100)=2.0; the logarithmic power of the second filter is: log(120)≈2.08; and so on, the vector of the logarithmic power spectrum is obtained [2.0,2.08,…,1.5].
[0074] The obtained logarithmic power spectrum is subjected to discrete cosine transform to obtain the MFCC feature coefficient. The specific mathematical formula is: ;in, For the MFCC feature coefficients of order; is the order of the MFCC feature coefficient currently calculated; is the total number of Mel filters; For the The output power of the Mel filter; is the index of the Mel filter; is the circumference of a circle; collect the MFCC feature coefficients corresponding to each frame to obtain a speech feature data set;
[0075] The order of the currently calculated MFCC feature coefficients is limited by the order limiting formula, and the order limiting formula is: ; is the order of the MFCC feature coefficients after restriction; is the logarithmic sum of all Mel filter output powers; and To control the parameter factor of the degree of restriction on the MFCC order;
[0076] For example, the total number of Mel filters =20, that is, there are 20 Mel filters, and the output power of each Mel filter Typical values are between 1 and 10. =5 for all , parameter factor that controls the degree of restriction on the MFCC order is 0.8, is 0.2; then the order of the restricted MFCC feature coefficients ; According to the calculation results, in the current case, the order of MFCC feature coefficients should not exceed 7;
[0077] The method for preprocessing the multidimensional context data is as follows: segmenting and encoding the multidimensional context data, and adjusting it through position encoding; inputting the adjusted multidimensional context data into the BERT model for feature extraction, thereby obtaining a context feature data set;
[0078] The method for preprocessing the translation environment data is as follows: data cleaning is performed on the obtained translation environment data, missing values are filled using interpolation, and outliers are detected and removed using clustering algorithms; Z-score standardization is used to convert the translation environment data into a standard normal distribution with a mean of 0 and a standard deviation of 1 to eliminate the dimension effect;
[0079] After the preprocessing in the above steps, a speech feature dataset, a situational feature dataset and a translation environment feature dataset are obtained.
[0080] The method for evaluating the influence of the translation environment feature dataset and the situational feature dataset on the speech feature dataset and obtaining the influenced feature dataset includes:
[0081] The speech feature dataset is used as the reference sequence, denoted as ;in, For the Reference value at a time point; is the index of the time point; each feature in the translation environment feature dataset and the context feature dataset is taken as a comparison sequence, recorded as ;in, For the The comparison sequence is in The value at a point in time; is the index of the comparison sequence; is the index of the time point;
[0082] Calculate reference sequence Compare each sequence The absolute difference at each time point is used to obtain the absolute difference matrix; the maximum difference and the minimum difference among all the absolute differences are selected, and the correlation coefficient of each sequence at each time point is calculated according to the absolute difference matrix: ;in, The reference sequence and The comparison sequence is in The larger the value, the stronger the correlation between the two sequences. For the reference sequence and The comparison sequence is in The difference between the time points; is the maximum difference among all absolute differences; is the minimum difference among all absolute differences; is the resolution coefficient;
[0083] The resolution factor is adjusted by the resolution factor adjustment formula Dynamic adjustment is performed, and the resolution coefficient adjustment formula is: ;in, is the resolution coefficient after dynamic adjustment; is the maximum value of the resolution coefficient; is the total number of time points; is the length of the reference sequence and comparison sequence; is a constant factor that controls the effect of the total number of time points on the resolution coefficient; A constant factor to control the effect of the length of the reference sequence and the comparison sequence on the resolution coefficient;
[0084] For example, the maximum value of the resolution coefficient 1, the total number of time points is 100, the length of the reference sequence and the comparison sequence =50, a constant factor that controls the effect of the total number of time points on the resolution coefficient The constant factor is 0.01, which controls the effect of the length of the reference sequence and the comparison sequence on the resolution coefficient. is 0.005; then the resolution coefficient after dynamic adjustment ;
[0085] The grey correlation degree is obtained by calculating the average value of the correlation coefficient of each comparison sequence; the grey correlation degree of each comparison sequence is sorted, and the larger the grey correlation degree, the more significant the influence of the feature on the speech feature; the previous The features corresponding to the grey relational degrees constitute the influencing feature data set.
[0086] The training method of the multilingual translation error prediction model includes:
[0087] The dataset is divided into a training set, a validation set, and a test set for training the model and evaluating the model performance. The sample set is a subset of the dataset, and each sample set includes a historical comprehensive feature dataset and the corresponding multilingual translation errors.
[0088] Constructing a multilingual translation error prediction model, including an input layer, an RNN layer, a fully connected layer and an output layer; the input layer is used to input a historical comprehensive feature data set, and the output layer is used to output multilingual translation errors; the multilingual translation error prediction model is a recurrent neural network model;
[0089] The mean square error is selected as the loss function to measure the difference between the model prediction result and the true label; the multilingual translation error prediction model is a recurrent neural network model;
[0090] Use the training set to train the multilingual translation error prediction model and use the Adam optimizer to minimize the loss function. Use the validation set to evaluate the performance of the model and measure the performance of the model by calculating the accuracy index.
[0091] The performance of the model is evaluated based on the validation set, and the hyperparameters of the model are tuned until the preset stopping condition is reached to obtain a trained multilingual translation error prediction model; the current comprehensive feature dataset is input into the trained multilingual translation error prediction model to predict the multilingual translation error.
[0092] The method of comparing the predicted multilingual translation error with the preset multilingual translation error threshold to determine whether the translation quality meets the standard includes:
[0093] If the predicted multilingual translation error is less than the preset multilingual translation error threshold, the translation quality is determined to be up to standard;
[0094] If the predicted multilingual translation error is greater than or equal to a preset multilingual translation error threshold, it is determined that the translation quality does not meet the standard.
[0095] Multilingual translation error data refers to translation segments whose translation quality does not meet the standards.
[0096] The training method of the error type diagnosis model includes:
[0097] The data set is divided into a training set, a validation set and a test set; the data set includes multilingual translation error data and corresponding translation error types; an error type diagnosis model is constructed, the input data of the model is historical multilingual translation error data, and the output label of the model is the translation error type; RBF is selected as the kernel function of the model, and the hyperparameters of the model are initialized, and the hyperparameters of the model include the regularization parameter C and the γ value of the RBF kernel function; the error type diagnosis model is a support vector machine model;
[0098] Use mean absolute error as the loss function to measure the error between the model's predicted value and the actual value; use the training set data to train the model and minimize the loss function through the Adam optimizer; use the validation set to evaluate the performance of the model and tune the model's hyperparameters until the preset stopping condition is reached; use the test set to evaluate the model's performance in the prediction task and input the current multilingual translation error data into the trained error type diagnosis model to obtain the translation error type.
[0099] Methods for correcting errors in real-time translation results that do not meet translation quality standards based on translation error types include:
[0100] The specific location where the translation error occurs is located according to the obtained translation error type, and the SMT model is used to re-decode the translation error part to generate more accurate translation candidates; the translation probability of all translation candidates is calculated, and the translation candidate with the highest translation probability is selected to replace the translation segment that does not meet the translation quality standard, so as to obtain the real-time translation result after error correction.
[0101] In this embodiment, the background noise of the speech input data is effectively eliminated through spectrum subtraction, making the subsequent speech recognition more accurate; at the same time, the extraction and order restriction of MFCC feature coefficients further enhance the expressiveness and robustness of speech features; for multi-dimensional contextual data, word segmentation, encoding and position adjustment make the data more standardized and structured, which is convenient for subsequent feature extraction and model processing; for translation environment data, data cleaning, missing value filling, outlier detection and standardization effectively eliminate the noise and dimensionality effects in the data, and improve the reliability and consistency of the data; order restriction can significantly reduce the number of MFCC feature coefficients, thereby reducing the computational cost of the model. This is particularly important for real-time speech recognition systems, which can shorten recognition time and improve user experience;
[0102] By calculating the absolute difference between the reference sequence (speech feature dataset) and each comparison sequence (features in the translation environment feature dataset and the context feature dataset) at each time point, as well as the subsequent correlation coefficient and grey correlation degree, this method can accurately quantify the correlation between different features; it helps to more accurately understand which features have a significant impact on the speech feature dataset; the dynamic adjustment formula of the resolution coefficient takes into account the total number of time points and the length of the reference sequence and the comparison sequence, making the calculation of the correlation more flexible and accurate; the dynamic adjustment can better adapt to the characteristics and needs of different datasets and improve the accuracy and reliability of the evaluation.
[0103] Embodiment 2
[0104] See also Figure 2 As shown, the part not described in detail in this embodiment is described in Example 1, which provides a multilingual intelligent real-time translation method, including:
[0105] S1, collecting speech input data, multi-dimensional context data and translation environment data;
[0106] S2, preprocessing the collected speech input data, multi-dimensional situational data and translation environment data to obtain a speech feature data set, a situational feature data set and a translation environment feature data set; evaluating the impact of the translation environment feature data set and the situational feature data set on the speech feature data set to obtain an impact feature data set;
[0107] S3. A multilingual translation error prediction model is obtained by training according to the influencing feature data set; and multilingual translation errors are predicted by the multilingual translation error prediction model;
[0108] S4, comparing the predicted multilingual translation error with the preset multilingual translation error threshold to determine whether the translation quality meets the standard; if the translation quality meets the standard, outputting the real-time translation result; if the translation quality does not meet the standard, the multilingual intelligent translation terminal issues a warning message and automatically collects multilingual translation error data;
[0109] S5, training an error type diagnosis model based on multilingual translation error data, and predicting translation error type data through the error type diagnosis model;
[0110] S6. Error correction is performed on the real-time translation results that do not meet the translation quality standards according to the translation error type data, and the real-time translation results after error correction are output.
[0111] Since the electronic device introduced in this embodiment is an electronic device used to implement a multilingual intelligent real-time translation system and method in the embodiment of the present application, based on the multilingual intelligent real-time translation system and method introduced in the embodiment of the present application, a person skilled in the art can understand the specific implementation of the electronic device of the present embodiment and its various variations, so how the electronic device implements the method in the embodiment of the present application is not described in detail here. As long as a person skilled in the art implements an electronic device used in a multilingual intelligent real-time translation system and method in the embodiment of the present application, it belongs to the scope of protection of this application.
[0112] The above formulas are all dimensionless and numerical calculations. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters and thresholds in the formula are set by technicians in this field according to actual conditions.
[0113] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technical users in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.
Claims
1. A multilingual intelligent real-time translation system, characterized in that: include: A data collection module, used to collect voice input data, multi-dimensional context data and translation environment data; The data processing module is used to pre-process the collected speech input data, multi-dimensional situational data and translation environment data to obtain a speech feature data set, a situational feature data set and a translation environment feature data set; evaluate the impact of the translation environment feature data set and the situational feature data set on the speech feature data set to obtain an impact feature data set; and perform weighted fusion of the impact feature data set and the speech feature data set to obtain a comprehensive feature data set; A multilingual translation module is used to obtain a multilingual translation error prediction model based on comprehensive feature data set training; and to predict multilingual translation errors through the multilingual translation error prediction model; A quality assessment module is used to compare the predicted multilingual translation error with the preset multilingual translation error threshold to determine whether the translation quality meets the standard; If the translation quality meets the requirements, the real-time translation result will be output; If the translation quality does not meet the standards, the multilingual intelligent translation terminal will issue a warning message and automatically collect multilingual translation error data; An error type evaluation module is used to obtain an error type diagnosis model based on multilingual translation error data training, and predict the translation error type through the error type diagnosis model; The error correction module is used to perform error correction on real-time translation results that do not meet the translation quality standards according to the translation error type, and output the real-time translation results after error correction; each module is connected to another via wired or wireless means.
2. A multilingual intelligent real-time translation system according to claim 1, characterized in that: The voice input data is continuous audio data received by the user; the multi-dimensional situational data includes emotion data, context data and cultural background data; and the translation environment data includes noise intensity data and device information data.
3. A multilingual intelligent real-time translation system according to claim 2, characterized in that: The method of preprocessing the collected speech input data, multi-dimensional situation data and translation environment data to obtain a speech feature data set, a situation feature data set and a translation environment feature data set includes: The method for preprocessing speech input data is: using spectrum subtraction to eliminate background noise in speech input data, extracting audio features from the speech input data after noise elimination; dividing continuous audio data into segments, and converting them into text data through a speech recognition model; the specific steps are: Performing short-time Fourier transform on the speech input data to convert from the time domain to the frequency domain; estimating the noise spectrum in the silent segment of the speech input data; performing spectrum subtraction denoising operation on the spectrum of the speech input data through the estimated noise spectrum; performing inverse short-time Fourier transform on the denoised spectrum to restore it to the speech signal in the time domain; the restored speech signal is the denoised speech input data; The denoised speech input data is framed and further divided into υ frames; each frame is multiplied by a window function and fast Fourier transform is performed to obtain a frequency domain signal of the denoised speech input data; a set of Mel filters are used to perform Mel filtering on the frequency domain signal to obtain the output power of each filter; the output power of the Mel filter is logarithmically transformed to obtain a logarithmic power spectrum; The obtained logarithmic power spectrum is subjected to discrete cosine transform to obtain the MFCC feature coefficient. The specific mathematical formula is: Among them, c m is the mth-order MFCC feature coefficient; m is the order of the currently calculated MFCC feature coefficient; K is the total number of Mel filters; S k is the output power of the kth Mel filter; k is the index of the Mel filter; π is the circumference of a circle; The order of the currently calculated MFCC feature coefficients is limited by the order limiting formula, and the order limiting formula is: m′ is the order of the MFCC feature coefficient after restriction; is the logarithm sum of the output powers of all Mel filters; α and β are parameter factors that control the degree of restriction on the MFCC order; collect the MFCC feature coefficients corresponding to each frame to obtain a feature coefficient data set; input the feature coefficient data set into the Transformer model to convert it into text data, and then obtain the speech feature data set; The method for preprocessing the multidimensional context data is as follows: segmenting and encoding the multidimensional context data, and adjusting it through position encoding; inputting the adjusted multidimensional context data into the BERT model for feature extraction, thereby obtaining a context feature data set; The method for preprocessing the translation environment data is as follows: data cleaning is performed on the obtained translation environment data, missing values are filled using interpolation, and outliers are detected and removed using clustering algorithms; Z-score standardization is used to convert the translation environment data into a standard normal distribution with a mean of 0 and a standard deviation of 1 to eliminate the dimension effect; After the preprocessing in the above steps, a speech feature dataset, a situational feature dataset and a translation environment feature dataset are obtained.
4. A multilingual intelligent real-time translation system according to claim 3, characterized in that: The method of evaluating the influence of the translation environment feature dataset and the situational feature dataset on the speech feature dataset to obtain the influenced feature dataset comprises: The speech feature dataset is used as a reference sequence, denoted as X0 = {x0(1), x0(2), ..., x0(n)}; where x0(n) is the reference value at the nth time point; n is the index of the time point; each feature in the translation environment feature dataset and the context feature dataset is used as a comparison sequence, denoted as X i ={x i (1),x i (2),...,x i (a)}; where x i (a) is the value of the i-th comparison sequence at the a-th time point; i is the index of the comparison sequence; a is the index of the time point; Calculate the value of the reference sequence X0 and each comparison sequence X i The absolute difference at each time point is used to obtain the absolute difference matrix; the maximum difference and the minimum difference among all the absolute differences are selected, and the correlation coefficient of each sequence at each time point is calculated according to the absolute difference matrix: Among them, ξ i (v) is the degree of association between the reference sequence and the ith comparison sequence at the vth time point; Δ i (v) is the difference between the reference sequence and the ith comparison sequence at the vth time point; Δ max is the maximum difference among all absolute differences; Δ min is the minimum difference among all absolute differences; ρ is the resolution coefficient; The resolution coefficient ρ is dynamically adjusted by the resolution coefficient adjustment formula, and the resolution coefficient adjustment formula is: Among them, ρ′ is the resolution coefficient after dynamic adjustment; ρ max is the maximum value of the resolution coefficient; b is the total number of time points; L is the length of the reference sequence and the comparison sequence; γ is the constant factor that controls the influence of the total number of time points on the resolution coefficient; η is the constant factor that controls the influence of the length of the reference sequence and the comparison sequence on the resolution coefficient; The grey correlation degree is obtained by calculating the average value of the correlation coefficient of each comparison sequence; the grey correlation degree of each comparison sequence is sorted, and the features corresponding to the first ω grey correlation degrees are selected to form the influencing feature data set.
5. A multilingual intelligent real-time translation system according to claim 4, characterized in that: The training method of the multilingual translation error prediction model includes: The dataset is divided into a training set, a validation set, and a test set for training the model and evaluating the model performance. The sample set is a subset of the dataset, and each sample set includes a historical comprehensive feature dataset and the corresponding multilingual translation errors. Constructing a multilingual translation error prediction model, including an input layer, an RNN layer, a fully connected layer and an output layer; the input layer is used to input a historical comprehensive feature data set, and the output layer is used to output multilingual translation errors; the multilingual translation error prediction model is a recurrent neural network model; The mean square error is selected as the loss function to measure the difference between the model prediction result and the true label; the multilingual translation error prediction model is a recurrent neural network model; Use the training set to train the multilingual translation error prediction model and use the Adam optimizer to minimize the loss function. Use the validation set to evaluate the performance of the model and measure the performance of the model by calculating the accuracy index. The performance of the model is evaluated based on the validation set, and the hyperparameters of the model are tuned until the preset stopping condition is reached to obtain a trained multilingual translation error prediction model; the current comprehensive feature dataset is input into the trained multilingual translation error prediction model to predict the multilingual translation error.
6. A multilingual intelligent real-time translation system according to claim 5, characterized in that: The method of comparing the predicted multilingual translation error with a preset multilingual translation error threshold to determine whether the translation quality meets the standard includes: If the predicted multilingual translation error is less than the preset multilingual translation error threshold, the translation quality is determined to be up to standard; If the predicted multilingual translation error is greater than or equal to a preset multilingual translation error threshold, it is determined that the translation quality does not meet the standard.
7. A multilingual intelligent real-time translation system according to claim 6, characterized in that: The multilingual translation error data refers to translation segments whose translation quality does not meet the standards.
8. A multilingual intelligent real-time translation system according to claim 7, characterized in that: The training method of the error type diagnosis model includes: The data set is divided into a training set, a validation set and a test set; the data set includes multilingual translation error data and corresponding translation error types; an error type diagnosis model is constructed, the input data of the model is historical multilingual translation error data, and the output label of the model is the translation error type; RBF is selected as the kernel function of the model, and the hyperparameters of the model are initialized; the error type diagnosis model is a support vector machine model; Use mean absolute error as the loss function to measure the error between the model's predicted value and the actual value; use the training set data to train the model and minimize the loss function through the Adam optimizer; use the validation set to evaluate the performance of the model and tune the model's hyperparameters until the preset stopping condition is reached; use the test set to evaluate the model's performance in the prediction task and input the current multilingual translation error data into the trained error type diagnosis model to obtain the translation error type.
9. A multilingual intelligent real-time translation system according to claim 8, characterized in that: The method for performing error correction on a real-time translation result whose translation quality does not meet the standard according to the translation error type comprises: The specific location where the translation error occurs is located according to the obtained translation error type, and the SMT model is used to re-decode the translation error part to generate more accurate translation candidates; the translation probability of all translation candidates is calculated, and the translation candidate with the highest translation probability is selected to replace the translation segment that does not meet the translation quality standard, so as to obtain the real-time translation result after error correction.
10. A multilingual intelligent real-time translation method, used to implement a multilingual intelligent real-time translation system according to any one of claims 1 to 9, characterized in that: include: S1, collecting speech input data, multi-dimensional context data and translation environment data; S2, preprocessing the collected speech input data, multi-dimensional situational data and translation environment data to obtain a speech feature data set, a situational feature data set and a translation environment feature data set; evaluating the impact of the translation environment feature data set and the situational feature data set on the speech feature data set to obtain an impact feature data set; S3. A multilingual translation error prediction model is obtained by training according to the influencing feature data set; and multilingual translation errors are predicted by the multilingual translation error prediction model; S4, comparing the predicted multilingual translation error with a preset multilingual translation error threshold to determine whether the translation quality meets the standard; If the translation quality meets the requirements, the real-time translation result will be output; If the translation quality does not meet the standards, the multilingual intelligent translation terminal will issue a warning message and automatically collect multilingual translation error data; S5, training an error type diagnosis model based on multilingual translation error data, and predicting translation error type data through the error type diagnosis model; S6. Error correction is performed on the real-time translation results that do not meet the translation quality standards according to the translation error type data, and the real-time translation results after error correction are output.
Citation Information
Patent Citations
Multi-language interactive real-time translation terminal and method supporting multiple platforms
CN116522960A
Speech recognition system based on digital twinning
CN116741148A
Machine translation system based on deep learning
CN116861929A