Intelligent voice automatic translation system based on AI recognition

By integrating speech recognition, language translation and emotion analysis modules in the intelligent speech automatic translation system, the problems of inaccurate translation and difficulty in emotional expression caused by different language habits are solved, and more accurate and vivid emotional speech translation is achieved.

CN119964573AInactive Publication Date: 2025-05-09ZHONGKE YIGE (SHENZHEN) TECHNOLOGY CO LTD
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510161318.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Different countries and regions have different language habits. Simple word-by-word translation will lead to the translation results that are contrary to the true meaning expressed by the narrator, and it is difficult to accurately express the narrator's emotional state.

Method used

It adopts an intelligent automatic speech translation system based on AI recognition, including a speech recognition module, a language translation module, a sentiment analysis module and a speech synthesis module. Speech signal preprocessing and feature extraction are performed through the speech recognition module, and acoustic models and language models are used for grammatical analysis and semantic understanding to generate the optimal text sequence. The emotion analysis module analyzes the narrator's emotions through the basic tone frequency, speech speed and rhythm characteristics, and adjusts the parameters of speech synthesis according to emotions.

Benefits of technology

It realizes that on the premise of ensuring the accuracy of pronunciation translation, the narrator's pronunciation meaning is more vivid and vivid, and the narrator's true semantics and emotional state are accurately conveyed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964573A_ABST
    Figure CN119964573A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice translation, in particular to an intelligent voice automatic translation system based on AI recognition. The system comprises a speech recognition module, a language translation module, an emotion analysis module and a speech synthesis module. The speech recognition module converts the speech signals into the text sequence, the text sequence is adjusted and optimized by using the context perception capability of the language model, and the language translation module translates the text sequence into the translated text, so that the accuracy of speech recognition is ensured, and the real meaning of a speaker is accurately expressed; meanwhile, the sentiment analysis module is used for carrying out sentiment analysis on a speaker, parameters of synthesized voice are automatically adjusted according to a mapping rule base between sentiment and voice parameters in the voice synthesis module, and a final translation result is broadcasted by taking a final translation as a text and voice; therefore, the voice meaning of the speaker can be expressed more vividly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech translation, and in particular to an intelligent speech automatic translation system based on AI recognition. Background Art

[0002] The intelligent voice automatic translation system is a system that uses artificial intelligence technology to directly convert the speech of one language into the speech of another language through natural language processing, speech recognition and speech synthesis. This system can realize real-time cross-language communication, greatly facilitating communication in the context of globalization.

[0003] At present, intelligent speech automatic translation systems can accurately recognize language and text. Due to the different living environments and customs of people in different countries and regions, the meaning, semantics, and grammar of the same text expressions are also different. At the same time, the emotional state of the user when speaking language and text will also affect the meaning of the language expression. In order to accurately translate the language information through grammatical analysis and semantic understanding instead of simply translating the speech word by word during speech translation, the parameters of the synthesized speech are automatically adjusted according to the narrator's emotional state. Under the premise of ensuring the accuracy of speech translation, the narrator's speech meaning can be expressed more vividly and vividly. Therefore, we propose an intelligent speech automatic translation system based on AI recognition. Summary of the Invention

[0004] The purpose of the present invention is to solve the problem of different language habits in different countries and regions. If a simple word-for-word translation is performed, the translation result will be contrary to the actual meaning expressed by the narrator. In order to increase grammatical analysis and semantic understanding during translation, the narrator's language information is expressed more accurately. At the same time, the parameters of the synthesized speech are automatically adjusted according to the narrator's emotional state. Under the premise of ensuring the accuracy of the speech translation, the narrator's speech meaning can be expressed more vividly and vividly.

[0005] To achieve the above-mentioned purpose, the present invention provides an intelligent speech automatic translation system based on AI recognition, which includes a speech recognition module, a language translation module, a sentiment analysis module and a speech synthesis module;

[0006] The speech recognition module preprocesses and extracts features from the speech signal to be recognized, inputs the speech feature signal into the acoustic model, calculates and outputs the probability distribution of each phoneme, obtains a phoneme sequence through the Viterbi algorithm, converts the phoneme sequence into a text sequence, inputs the text sequence into the language model, adjusts and optimizes the text sequence using the context-awareness of the language model, generates multiple candidate text sequences, and selects the optimal text sequence by weighted fusion, taking into account the confidence of the acoustic model and the language model.

[0007] The language translation module inputs the text sequence recognized by the speech recognition module into the statistical machine translation model, performs lexical and syntactic analysis on the text sequence, identifies words and phrases, generates multiple target language translation candidates using statistical probability, evaluates the candidate translations using the language model, and selects the translation with the highest score and the most consistent with the language expression habits as the final translation;

[0008] The sentiment analysis module receives the speech feature data extracted by the speech recognition module, takes pitch frequency, speech rate and prosodic features as input, and uses the sentiment categories of positive, negative and neutral as output to establish a recurrent neural network model for sentiment analysis. The cross-entropy loss function is used to measure the difference between the probability distribution of the sentiment category output by the model and the true sentiment label, the model is trained, and the Adam optimizer is used to optimize and adjust the model parameters to output the sentiment category of the speech;

[0009] The speech synthesis module receives the final translation text of the language translation module and the speech emotion category output by the emotion analysis module, automatically adjusts the parameters of the synthesized speech according to the mapping rule library between emotion and speech parameters, uses the final translation text as text, and voice broadcasts the final translation result.

[0010] As a further improvement of this technical solution, the speech recognition module includes an acoustic recognition unit and a language recognition unit;

[0011] The acoustic recognition unit extracts the Mel-frequency cepstral coefficients of each frame of speech signal, inputs the Mel-frequency cepstral coefficients into the acoustic model, outputs phoneme labels, calculates the probability distribution of each phoneme through forward propagation, obtains a phoneme sequence through the Viterbi algorithm, and converts the phoneme sequence into a text sequence;

[0012] The language recognition unit takes the text sequence output by the acoustic model as input, evaluates the text sequence based on the grammar, semantics and pragmatics within the language model, uses the context-awareness of the language model to correct and adjust the text sequence, and selects the optimal text sequence through weighted fusion.

[0013] As a further improvement of the present technical solution, the acoustic recognition unit frames and windows each frame of speech signal, converts the time domain signal into a frequency domain signal using short-time Fourier transform, and then extracts the Mel-frequency cepstral coefficients of each frame of speech signal.

[0014] As a further improvement of this technical solution, in the acoustic recognition unit, the short-time Fourier transform formula is:

[0015]

[0016] Where X(k) is the transformed spectral coefficient, x(n) is the discrete-time speech signal, w(n) is the window function, N is the frame length, k is the frequency index, and j is the imaginary unit.

[0017] As a further improvement of the present technical solution, the language recognition unit performs weighted summation of the acoustic confidence and language confidence of the candidate text sequence to obtain a fused confidence score. The formula for the fused confidence score is:

[0018] P = αPa + (1-α) Pl;

[0019] Among them, P is the confidence score after fusion, α is the weighting coefficient of the acoustic model, 1-α is the weighting coefficient of the language model, and P a is the acoustic confidence, P l is the language confidence.

[0020] As a further improvement of this technical solution, the language translation module uses a gradient descent algorithm to update the parameters of the statistical machine translation model. The formula is:

[0021]

[0022] Among them, θ is the model parameter, θ new is the updated model parameter, θ old is the model parameter before updating, η is the learning rate, is the gradient of the loss function with respect to θ.

[0023] As a further improvement of this technical solution, the sentiment analysis module includes a model building unit and an optimization and adjustment unit;

[0024] The model building unit takes pitch frequency, speech rate and prosodic features as input and takes the emotional categories of positive, negative and neutral as output to build a recurrent neural network model for sentiment analysis, and trains the model using a cross entropy loss function;

[0025] The optimization and adjustment unit optimizes and adjusts the model parameters using the Adam optimizer, and uses the model's accuracy, recall rate, and F1 value as evaluation indicators to measure the model's prediction performance on different emotion categories.

[0026] As a further improvement of the present technical solution, the model building unit utilizes a recurrent neural network layer to capture long-term dependencies in the input feature sequence, models the dynamic changes of emotions, and utilizes a fully connected layer to map the output of the recurrent neural network layer to the emotion category space.

[0027] As a further improvement of the present technical solution, the optimization adjustment unit uses grid search to adjust the hyperparameters of the model according to the evaluation results.

[0028] As a further improvement of the present technical solution, the speech synthesis module uses a support vector machine as a classification model to establish a mapping rule library between emotions and speech parameters.

[0029] Compared with the prior art, the present invention has the following beneficial effects:

[0030] 1. This AI-based intelligent automatic speech translation system uses the speech recognition module to preprocess and extract features from the speech signal to be recognized. It then uses the acoustic model to calculate and output the probability distribution of each phoneme. It then uses the Viterbi algorithm to obtain a phoneme sequence, which is then converted into a text sequence. The text sequence is then evaluated based on the syntax, semantics, and pragmatics of the language model. The language model's contextual awareness is then used to correct and adjust the text sequence. The optimal text sequence is selected through weighted fusion. During speech recognition, the system uses grammatical analysis and semantic understanding to parse the speaker's true meaning into text form. The language translation module then translates the text sequence into a target text, ensuring the accuracy of speech recognition and accurately conveying the speaker's true meaning.

[0031] 2. The sentiment analysis module uses pitch frequency, speaking rate, and prosodic features as input and positive, negative, and neutral sentiment categories as output to establish a recurrent neural network model for sentiment analysis. The model analyzes the narrator's emotions and automatically adjusts the parameters of the synthesized speech based on the mapping rule library between emotions and speech parameters. The final translation result is broadcasted as text, expressing the narrator's emotional state during translation, thereby more vividly expressing the meaning of the narrator's speech. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 It is a schematic diagram of the overall process of the present invention;

[0033] Figure 2 It is a schematic diagram of the overall details of the present invention.

[0034] The meaning of each number in the figure is:

[0035] 100. Speech recognition module; 110. Acoustic recognition unit; 120. Language recognition unit; 200. Language translation module; 300. Sentiment analysis module; 310. Model building unit; 320. Optimization and adjustment unit; 400. Speech synthesis module. DETAILED DESCRIPTION

[0036] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0037] At present, different countries and regions have different language habits. If a simple word-for-word translation is performed, the translation result will be contrary to the true meaning expressed by the narrator. In order to increase grammatical analysis and semantic understanding during translation, and express the narrator's language information more accurately, the parameters of the synthesized speech are automatically adjusted according to the narrator's emotional state. Under the premise of ensuring the accuracy of the speech translation, the narrator's voice meaning can be expressed more vividly.

[0038] Therefore, the present invention proposes to convert the voice signal into a text sequence through the voice recognition module 100, and adjust and optimize the text sequence by using the context perception ability of the language model, translate the text sequence into a translation through the language translation module 200, and at the same time use the sentiment analysis module to perform sentiment analysis on the narrator, and automatically adjust the parameters of the synthesized speech according to the mapping rule library between the emotion and the speech parameters in the speech synthesis module, so as to finally translate the translation into text and announce the final translation result by voice.

[0039] The details are as follows:

[0040] See also Figure 1 As shown, the present invention provides an intelligent speech automatic translation system based on AI recognition, including a speech recognition module 100, a language translation module 200, a sentiment analysis module 300 and a speech synthesis module 400;

[0041] The speech recognition module 100 preprocesses and extracts features from the speech signal to be recognized, inputs the speech feature signal into the acoustic model, calculates and outputs the probability distribution of each phoneme, obtains a phoneme sequence using the Viterbi algorithm, converts the phoneme sequence into a text sequence, inputs the text sequence into the language model, uses the context-awareness of the language model to adjust and optimize the text sequence, generates multiple candidate text sequences, and selects the optimal text sequence through weighted fusion, taking into account the confidence of the acoustic model and language model.

[0042] The language translation module 200 inputs the text sequence recognized by the speech recognition module 100 into the statistical machine translation model, performs lexical and syntactic analysis on the text sequence, identifies words and phrases, and uses statistical probability to generate multiple target language translation candidates. The candidate translations are evaluated using the language model and the translation with the highest score and the most consistent with the language expression habits is selected as the final translation;

[0043] The sentiment analysis module 300 receives the speech feature data extracted by the speech recognition module 100, takes pitch frequency, speech rate, and prosodic features as input, and uses the sentiment categories of positive, negative, and neutral as output to establish a recurrent neural network model for sentiment analysis. The model is trained by measuring the difference between the probability distribution of the sentiment categories output by the model and the true sentiment labels using a cross-entropy loss function, and the model parameters are optimized and adjusted using the Adam optimizer to output the sentiment category of the speech.

[0044] The speech synthesis module 400 receives the final translation text of the language translation module 200 and the speech emotion category output by the sentiment analysis module 300, and automatically adjusts the parameters of the synthesized speech according to the mapping rule library between emotion and speech parameters, and uses the final translation text as text to voice broadcast the final translation result.

[0045] like Figure 2 As shown, the speech recognition module 100 includes an acoustic recognition unit 110 and a language recognition unit 120;

[0046] The acoustic recognition unit 110 extracts the Mel-frequency cepstral coefficients of each frame of speech signal, inputs the Mel-frequency cepstral coefficients into the acoustic model, outputs the phoneme labels, calculates the probability distribution of each phoneme through forward propagation, obtains the phoneme sequence through the Viterbi algorithm, and converts the phoneme sequence into a text sequence;

[0047] The language recognition unit 120 takes the text sequence output by the acoustic model as input, evaluates the text sequence based on the grammar, semantics, and pragmatics of the language model, uses the context-awareness of the language model to correct and adjust the text sequence, and selects the optimal text sequence through weighted fusion.

[0048] Collect a large number of speech samples from different speakers and environments, covering various speech features and pronunciation methods, to form a rich speech data set. The collected speech data is preprocessed, including noise removal, volume normalization, and pre-emphasis, to improve the quality and stability of the speech signal. At the same time, the speech signal is segmented into frames of fixed length, typically 20-30 milliseconds per frame, with a frame shift of about 10 milliseconds.

[0049] By performing a short-time Fourier transform on the speech signal, it is converted to the frequency domain, and then the Mel-frequency cepstral coefficients are calculated. The Mel-frequency cepstral coefficients can effectively capture the spectral characteristics of the speech and have good discrimination between different speech sounds.

[0050] In order to more quickly extract the Mel-frequency cepstral coefficients of each frame of speech signal, the acoustic recognition unit 110 performs framing and windowing on each frame of speech signal, converts the time domain signal into a frequency domain signal using a short-time Fourier transform, and then extracts the Mel-frequency cepstral coefficients of each frame of speech signal;

[0051] The input speech signal is sampled, converting the continuous analog signal into a discrete digital signal. The appropriate quantization accuracy is determined, using 16-bit or higher quantization accuracy to ensure signal quality. The speech signal is then divided into short frames, typically 20 to 30 milliseconds long with a frame shift of 10 to 15 milliseconds. To reduce spectral leakage, a Hamming window function is applied to each frame, and the Mel-frequency cepstral coefficient feature of each frame of speech signal is calculated. This is a widely used feature in speech recognition. It simulates the human ear's perception of sounds of different frequencies. Typically, a 12- to 13-dimensional Mel-frequency cepstral coefficient feature vector is extracted, and then the first- and second-order differences are added to form a feature vector sequence of approximately 39 dimensions.

[0052] In order to more accurately convert the time domain signal into the frequency domain signal, in the acoustic recognition unit 110, the short-time Fourier transform formula is:

[0053]

[0054] Where X(k) is the transformed spectral coefficient, x(n) is the discrete-time speech signal, w(n) is the window function, N is the frame length, k is the frequency index, and j is the imaginary unit.

[0055] The speech signal is framed and the frame length and frame shift parameters are determined. According to the short-time Fourier transform formula, each frame is looped through, a window function is applied, a discrete Fourier transform is performed, and the appropriate portion of the result is taken to complete the short-time Fourier transform calculation for each frame of speech. The frequency domain information is obtained, and the amplitude spectrum is calculated and visualized. The frequency domain changes of the speech signal over time are presented in the form of a spectrogram, helping us better analyze the frequency characteristics of speech.

[0056] After converting it to the frequency domain through short-time Fourier transform, the different frequency components contained in the speech and their changes over time can be clearly seen. For example, in a mixed audio containing music and speech, the frequency range of different instruments and information such as the fundamental frequency and harmonics of the speech can be easily distinguished through frequency domain analysis. The frequency domain representation can intuitively present the spectral characteristics of the speech, such as formants. Formants are peaks in the speech spectrum that reflect the resonance characteristics of the vocal tract and have a significant impact on the timbre and sound quality of the speech. Frequency domain analysis can accurately determine the position and intensity of formants, which is helpful for applications such as speech recognition and synthesis.

[0057] Speech signals are typically non-stationary signals, with frequency characteristics that vary over time. The Short-Time Fourier Transform (STFT) uses a sliding window approach to analyze speech within a specific timeframe, achieving time-frequency localization. This approach not only reflects the frequency characteristics of speech signals at different moments, but also preserves temporal information to a certain extent. For transient phenomena in speech, such as plosives, the STFT can accurately capture their frequency variations within the corresponding time window. These transient signals can be difficult to analyze in the time domain due to their short duration.

[0058] The phoneme sequence output by the acoustic model is converted into a corresponding text sequence as input to the language model. The language model analyzes the input text sequence and, based on its learned language knowledge, evaluates the text's grammatical correctness, semantic rationality, and coherence. For parts that do not conform to grammatical rules, such as subject-verb inconsistency or incorrect verb tense, the language model corrects them according to language rules. For semantic incoherence or irrationality, such as illogical context or unclear reference, the text is optimized by adjusting the choice of words or phrases. The language model's contextual awareness is used to integrate the current phoneme sequence with the surrounding context, making the adjusted text more natural and fluent in the overall context.

[0059] In order to better achieve weighted fusion, the language recognition unit 120 performs weighted summation on the acoustic confidence and language confidence of the candidate text sequence to obtain a fused confidence score. The formula for the fused confidence score is:

[0060] P = αPa + (1-α) Pl;

[0061] Among them, P is the confidence score after fusion, α is the weighting coefficient of the acoustic model, 1-α is the weighting coefficient of the language model, and P a is the acoustic confidence, P l is the language confidence.

[0062] The speech signal is converted into an intermediate representation such as phonemes or subwords, and a probability value is assigned to each possible phoneme or subword sequence to indicate how well the sequence matches the input speech signal, i.e., the acoustic confidence. Based on the knowledge of the language's grammar and semantics, a probability value is assigned to each possible text sequence to indicate the linguistic rationality of the sequence, i.e., the language confidence.

[0063] The weighting coefficients of the acoustic and language models are manually set based on previous experimental data and experience. For example, in speech recognition tasks, if the speech signal quality is high, the weight of the acoustic model can be appropriately increased; if the standardization and logic of the language are high, the weight of the language model can be appropriately increased. Using a large amount of labeled data for training, the optimization algorithm automatically learns the optimal weighting coefficients. For example, the outputs of the acoustic and language models can be used as features, and the correct text sequences as labels. A linear regression model or a neural network model can be trained to learn the weighting coefficients. Based on the fused confidence score, the text sequence with the highest score is selected as the final optimal result.

[0064] In order to better update the parameters of the statistical machine translation model, the language translation module 200 uses a gradient descent algorithm to update the parameters of the statistical machine translation model. The formula is:

[0065]

[0066] Among them, θ is the model parameter, θ new is the updated model parameter, θ old is the model parameter before updating, η is the learning rate, is the gradient of the loss function with respect to θ.

[0067] Based on the loss function, the chain rule is used to calculate the gradient of the loss function with respect to each parameter of the model and determine the learning rate η, which determines the step size of the parameter update. If the learning rate is too large, the model may not converge or converge to a local optimal solution; if the learning rate is too small, the model converges very slowly.

[0068] The gradient of the loss function obtained by the chain rule is substituted into the formula to obtain the updated model parameter θ new , performing multiple iterations, updating the model parameters with each iteration. As the number of iterations increases, the model's translation probability on the training corpus gradually increases, and the loss function gradually decreases. Using a learning rate decay strategy, such as gradually reducing the learning rate as training progresses, can improve the model's convergence speed and stability.

[0069] The sentiment analysis module 300 includes a model building unit 310 and an optimization and adjustment unit 320;

[0070] The model building unit 310 uses the pitch frequency, speech rate and prosodic features as inputs and the emotional categories of positive, negative and neutral as outputs to build a recurrent neural network model for sentiment analysis, and trains the model using a cross entropy loss function;

[0071] The optimization and adjustment unit 320 optimizes and adjusts the model parameters using the Adam optimizer, and uses the model's accuracy, recall rate, and F1 value as evaluation indicators to measure the model's prediction performance on different emotion categories;

[0072] Collect speech data containing different emotional tendencies and label the corresponding emotion categories, i.e., positive, negative, or neutral. This speech data should cover a wide range of emotional variations and scenarios. Process the collected speech data to extract fundamental frequency, speech rate, and prosodic features. Use the Praat tool to extract fundamental frequency, calculate speech rate by analyzing speech duration and pauses, and extract prosodic features using speech signal energy and pitch variations.

[0073] The labeled data is divided into training set, validation set and test set. The training set is used for model training, the validation set is used to adjust the model's hyperparameters, and the test set is used to evaluate the model's final performance.

[0074] Create an input layer that accepts pitch frequency, speech rate, and prosodic features as input. The number of neurons in the input layer should correspond to the dimensionality of the input features. Create an output layer that uses the Softmax activation function to convert the output of the fully connected layer into a probability distribution of emotion categories. The Softmax function normalizes the output values ​​into probabilities, so that each emotion category has a corresponding probability value, indicating the likelihood that the speech belongs to that emotion category.

[0075] In order to better establish a recurrent neural network model for sentiment analysis, the model building unit 310 uses a recurrent neural network layer to capture long-term dependencies in the input feature sequence, models the dynamic changes of sentiment, and uses a fully connected layer to map the output of the recurrent neural network layer to the sentiment category space;

[0076] Add one or more recurrent neural network layers, such as long short-term memory (LSTM) or gated recurrent units (GRU). These layers can capture long-term dependencies in the input feature sequence and model the dynamic changes in emotion. The number of recurrent neural network layers and the number of neurons can be adjusted according to the specific situation. After the recurrent neural network layer, add one or more fully connected layers to map the output of the recurrent neural network layer to the emotion category space. The number of neurons in the fully connected layer should correspond to the number of emotion categories, that is, three neurons for positive, negative, and neutral emotion categories respectively.

[0077] In order to better adjust the hyperparameters of the recurrent neural network model for sentiment analysis, the optimization and adjustment unit 320 uses grid search to adjust the hyperparameters of the model according to the evaluation results;

[0078] Use the test set data to evaluate the trained model and calculate evaluation metrics such as accuracy, recall, and F1 value to measure the model's predictive performance for different sentiment categories. At the same time, you can also draw visualizations such as confusion matrices to intuitively analyze the model's prediction results. Based on the results of the model evaluation, adjust the model's hyperparameters, such as the number of recurrent neural network layers, the number of neurons, the learning rate, and the batch size. You can use methods such as grid search, random search, or Bayesian optimization to find the optimal hyperparameter combination. Based on the results of the hyperparameter adjustment, further optimize and improve the model, such as adding more training data, adjusting the model structure, and adopting more advanced training methods, to improve the model's performance and generalization ability.

[0079] In order to automatically adjust the parameters of the synthesized speech according to the emotion analysis, the speech synthesis module 400 uses the support vector machine as the classification model and establishes a mapping rule library between emotions and speech parameters;

[0080] A support vector machine (SVM) was selected as the classification model, as it performs well with small sample sizes and high-dimensional features. The emotion-labeled speech data was divided into a training set and a test set in a 7:3 ratio. The SVM model was trained using the training set, and its parameters were adjusted using grid search. The model's performance was evaluated on the test set, achieving an accuracy of 75%. The recall rate also reached acceptable levels across different emotion categories, demonstrating the model's effectiveness.

[0081] Analyzing the trained SVM model, for pitch frequency, when the mean is greater than 200Hz, the emotion tends to be positive; when the mean is less than 100Hz, the emotion is negative or neutral. For speech rate, when the number of words per second is greater than 5, it is often associated with positive and excited emotions; when it is less than 3, it is negative or calm emotions. In terms of prosodic features, if the pitch rising slope is greater than 0.5 and the energy fluctuation is large, the emotion is positive; if the pitch is relatively stable and the energy is low, the emotion is neutral, as shown in the following table:

[0082] Emotional categories Fundamental frequency Speech speed range Prosodic features positive >200Hz >5 >0.5 negative <100Hz <3 <0 neutral 100Hz-200HZ 3-5 Near 0

[0083] We continuously collect new voice data for verification and optimization. For example, we discovered that the rules need to be fine-tuned for voice in certain specific areas, such as customer service conversations. This is because customer service representatives may speak at a relatively moderate speed when expressing positive emotions, rather than simply at a speed greater than 5. Based on these new situations, we adjust the rule base to improve its accuracy in practical applications.

[0084] In summary, the working principle of this solution is as follows:

[0085] The intelligent automatic speech translation system based on AI recognition pre-processes and extracts features of the speech signal to be recognized through the speech recognition module 100, calculates and outputs the probability distribution of each phoneme using the acoustic model, obtains the phoneme sequence through the Viterbi algorithm, and converts the phoneme sequence into a text sequence. The text sequence is evaluated according to the grammar, semantics and pragmatics in the language model, and the context perception ability of the language model is used to correct and adjust the text sequence. The optimal text sequence is selected through weighted fusion, so that the real semantics of the narrator are parsed into text form through grammatical analysis and semantic understanding during speech recognition, and then the text sequence is converted into a text form through language recognition. The translation module 200 translates the text sequence into a translation to ensure the accuracy of speech recognition and accurately express the narrator's true meaning. The sentiment analysis module 300 uses the fundamental frequency, speech rate and rhythmic features as input and the sentiment categories of positive, negative and neutral as output to establish a recurrent neural network model for sentiment analysis, analyze the narrator's emotions, and automatically adjust the parameters of the synthesized speech based on the mapping rule library between emotions and speech parameters. The final translation is used as text and the final translation result is broadcast out in voice. During the translation, the narrator's emotional state is expressed, thereby being able to express the narrator's speech meaning more vividly.

[0086] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely preferred examples of the present invention and are not intended to limit the present invention. Various changes and improvements may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and improvements fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.

Claims

1. An intelligent voice automatic translation system based on AI recognition, characterized by: It includes a speech recognition module (100), a language translation module (200), a sentiment analysis module (300) and a speech synthesis module (400); The speech recognition module (100) performs preprocessing and feature extraction on the speech signal to be recognized, inputs the speech feature signal into the acoustic model, calculates and outputs the probability distribution of each phoneme, obtains a phoneme sequence through the Viterbi algorithm, converts the phoneme sequence into a text sequence, inputs the text sequence into the language model, adjusts and optimizes the text sequence using the context perception capability of the language model, generates multiple candidate text sequences, and selects the optimal text sequence by weighted fusion, taking into account the confidence of the acoustic model and the language model; The language translation module (200) inputs the text sequence recognized by the speech recognition module (100) into the statistical machine translation model, performs lexical and syntactic analysis on the text sequence, recognizes words and phrases, generates multiple target language translation candidates using statistical probability, evaluates the candidate translations using the language model, and selects the translation with the highest score and the most consistent with the language expression habits as the final translation; The sentiment analysis module (300) receives the speech feature data extracted by the speech recognition module (100), takes the fundamental frequency, speech speed and prosodic features as input, takes the sentiment categories of positive, negative and neutral as output to establish a recurrent neural network model for sentiment analysis, measures the difference between the probability distribution of the sentiment category output by the model and the real sentiment label through a cross entropy loss function, trains the model, optimizes and adjusts the model parameters using an Adam optimizer, and outputs the sentiment category of the speech; The speech synthesis module (400) receives the final translation text of the language translation module (200) and the speech emotion category output by the emotion analysis module (300), automatically adjusts the parameters of the synthesized speech according to the mapping rule library between emotion and speech parameters, uses the final translation text as text, and voice broadcasts the final translation result.

2. The intelligent voice automatic translation system based on AI recognition according to claim 1 is characterized in that: The speech recognition module (100) comprises an acoustic recognition unit (110) and a language recognition unit (120); The acoustic recognition unit (110) extracts the Mel-frequency cepstral coefficients of each frame of speech signal, inputs the Mel-frequency cepstral coefficients into the acoustic model, outputs phoneme labels, obtains the probability distribution of each phoneme through forward propagation calculation, obtains the phoneme sequence through the Viterbi algorithm, and converts the phoneme sequence into a text sequence; The language recognition unit (120) takes the text sequence output by the acoustic model as input, evaluates the text sequence according to the grammar, semantics and pragmatics in the language model, uses the context perception ability of the language model to correct and adjust the text sequence, and selects the optimal text sequence through weighted fusion.

3. The intelligent voice automatic translation system based on AI recognition according to claim 2 is characterized in that: The acoustic recognition unit (110) performs framing and windowing on each frame of speech signal, converts the time domain signal into a frequency domain signal using short-time Fourier transform, and then extracts the Mel-frequency cepstral coefficients of each frame of speech signal.

4. The intelligent voice automatic translation system based on AI recognition according to claim 3 is characterized in that: In the acoustic recognition unit (110), the short-time Fourier transform formula is: Among them, X(k) is the transformed spectral coefficient, x(n) is the discrete time speech signal, w(n) is the window function, N is the frame length, k is the frequency index, and j is the imaginary unit.

5. The intelligent voice automatic translation system based on AI recognition according to claim 2 is characterized in that: The language recognition unit (120) performs weighted summation of the acoustic confidence and the language confidence of the candidate text sequence to obtain a fused confidence score. The fused confidence score formula is: P = αPa + (1-α) Pl; Among them, P is the confidence score after fusion, α is the weighting coefficient of the acoustic model, 1-α is the weighting coefficient of the language model, and P a is the acoustic confidence, P l is the language confidence.

6. The intelligent voice automatic translation system based on AI recognition according to claim 1, characterized in that: The language translation module (200) uses a gradient descent algorithm to update the parameters of the statistical machine translation model, and the formula is: Among them, θ is the model parameter, θ new is the updated model parameter, θ old is the model parameter before updating, η is the learning rate, is the gradient of the loss function with respect to θ.

7. The intelligent voice automatic translation system based on AI recognition according to claim 1, characterized in that: The sentiment analysis module (300) includes a model building unit (310) and an optimization and adjustment unit (320); The model building unit (310) uses the fundamental frequency, speech speed and prosodic features as inputs and uses the emotional categories of positive, negative and neutral as outputs to build a recurrent neural network model for sentiment analysis, and uses a cross entropy loss function to train the model; The optimization and adjustment unit (320) uses the Adam optimizer to optimize and adjust the model parameters, and uses the accuracy, recall rate, and F1 value of the model as evaluation indicators to measure the prediction performance of the model in different emotion categories.

8. The intelligent voice automatic translation system based on AI recognition according to claim 7 is characterized in that: The model building unit (310) uses a recurrent neural network layer to capture long-term dependencies in an input feature sequence, models dynamic changes in emotions, and uses a fully connected layer to map the output of the recurrent neural network layer to an emotion category space.

9. The intelligent voice automatic translation system based on AI recognition according to claim 7, characterized in that: The optimization and adjustment unit (320) uses grid search to adjust the hyperparameters of the model according to the evaluation results.

10. The intelligent voice automatic translation system based on AI recognition according to claim 1, characterized in that: The speech synthesis module (400) uses a support vector machine as a classification model to establish a mapping rule base between emotions and speech parameters.

Citation Information

Cited By

  • Artificial intelligence assisted cross-language automatic note taking and term marking system

    CN120183408A

  • An AI-Assisted Cross-Language Automatic Note-Taking and Term Annotation System

    CN120183408B

  • Alarm receiving and handling foreign language virtual simultaneous transmission method and system based on intelligent voice interaction service

    CN120708616A

  • System and method for evaluating reading quality of voice model formula

    CN121075370A

  • An evaluation system and method for formula reading quality of a speech model

    CN121075370B