Multi-language real-time translation system based on artificial intelligence

Through a multilingual real-time translation system based on artificial intelligence, the problems of high latency and unstable translation quality of the existing translation system are solved, efficient and accurate multilingual real-time translation is achieved, and high-demand translation is met.

CN120278167APending Publication Date: 2025-07-08HUNAN AUTOMOTIVE ENG VOCATIONAL COLLEGE +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510327048.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing translation systems have high latency, unstable translation quality, poor voice and text compatibility, and the traditional translation methods are inefficient, which cannot meet the needs of efficient and accurate multilingual real-time translation.

Method used

A multilingual real-time translation system based on artificial intelligence is adopted, including speech acquisition, speech analysis, intelligent prediction and translation output modules, and language analysis is determined through speech analysis, text conversion and translation output, and combining probability prediction and timing judgment modules to achieve efficient translation.

Benefits of technology

On the premise of ensuring the quality of translation, it effectively reduces the translation delay and meets high-demand translations, achieving efficient and accurate multilingual real-time translation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278167A_ABST
    Figure CN120278167A_ABST
Patent Text Reader

Abstract

The invention provides a multilingual real-time translation system based on artificial intelligence, and relates to the field of electric digital data processing, the multilingual real-time translation system comprises a voice acquisition module, a voice analysis module, an intelligent prediction module and a translation output module, the voice acquisition module is used for acquiring voice information, the voice analysis module is used for analyzing voice content, and the intelligent prediction module is used for predicting the voice content; the intelligent prediction module is used for predicting and analyzing subsequent voice, and the translation output module is used for translating and outputting voice content; the system predicts and analyzes the subsequent information in the translation process, adjusts the translation rhythm according to the prediction result, and can effectively reduce translation errors and translate the content in time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of electrical digital data processing, and more particularly to a multi-language real-time translation system based on artificial intelligence. Background Art

[0002] With the deepening of globalization, the demand for communication between different languages is increasing. Traditional manual translation methods are inefficient, and existing translation methods based on rules and statistical models still have deficiencies in accuracy and fluency. However, existing real-time translation systems still have problems such as high latency, unstable translation quality, and poor voice and text compatibility. Therefore, there is an urgent need for an efficient, accurate, and low-latency multi-language real-time translation system.

[0003] The foregoing discussion of the background art is only intended to facilitate an understanding of the present invention. This discussion does not recognize or admit that any of the materials mentioned is part of common general knowledge.

[0004] Many translation systems have now been developed. After a large amount of retrieval and reference, it is found that existing translation systems are like the system disclosed in the publication number CN111241853B. These systems generally include obtaining first session information; determining the target user who receives the first session information, and obtaining the target receiving language set by the target user; adding a target language label corresponding to the target receiving language to the first session information; inputting the first session information with the target language label added into a pre-trained multi-language neural network machine translation model to obtain second session information output by the multi-language neural network machine translation model, where the second session information is the session information after the first session information is translated into the target receiving language; and sending the second session information to the target user to achieve real-time translation in a multi-language session, solving the problem of difficult communication and exchange between different languages, so as to meet the needs of people's daily work and life. However, the system uses a traditional translation method and requires sufficient voice information to ensure accurate translation, with a high latency. Summary of the Invention

[0005] The object of the present invention is to propose a multi-language real-time translation system based on artificial intelligence for the existing deficiencies.

[0006] The present invention adopts the following technical solutions:

[0007] A multi-language real-time translation system based on artificial intelligence includes a voice collection module, a voice parsing module, an intelligent prediction module, and a translation output module;

[0008] The voice collection module is used to collect voice information, the voice parsing module is used to parse the voice content, the intelligent prediction module is used to perform predictive analysis on subsequent voices, and the translation output module is used to translate and output the voice content;

[0009] The voice acquisition module includes a signal acquisition unit, a noise filtering unit, and a signal caching unit. The signal acquisition unit is used to acquire sound signals, the noise filtering unit is used to filter noise signals, and the signal caching unit is used to store signal data;

[0010] The voice parsing module includes a language parsing unit, a text conversion unit, and an output control unit. The language parsing unit is used to determine the language type of the collected voice, the text conversion unit is used to convert the voice signal into text information, and the output control unit is used to control the output process of the text information;

[0011] The intelligent prediction module includes a dialogue information database, a probability prediction unit, and a timing judgment unit. The dialogue information database is used to store dialogue data, the probability prediction unit is used to predict the probability information of subsequent voices, and the timing judgment unit is used to judge whether it is in the translation timing;

[0012] The translation output module includes a target setting unit, a text translation unit, and a voice output unit. The target setting unit is used to set translation information, the text translation unit is used to translate the original language text into the target language text, and the voice output unit is used to play the target language voice.

[0013] Further, the probability prediction unit includes a task generation processor, a result classification processor, and a probability calculation processor. The task generation processor generates a retrieval task based on the text information, the result classification processor is used to classify and process the retrieval results, and the probability calculation processor is used to calculate the probability information of each translation classification;

[0014] The probability calculation processor calculates the probability value P(i) of the i-th translation classification according to the following formula:

[0015]

[0016] where n(i) is the statistical quantity of the i-th translation classification, and m is the number of translation classifications.

[0017] Further, the timing judgment unit includes a semantic breakpoint detector, a delay tolerance analyzer, and a probability judgment processor. The semantic breakpoint detector is used to identify pauses and logical demarcation points, the delay tolerance analyzer is used to manage and analyze the time when translation has not been performed, and the probability judgment processor is used to judge whether the current translation classification with the highest probability meets the translation timing;

[0018] The delay tolerance analyzer calculates the tolerance strength R according to the following formula:

[0019]

[0020] Among them, t is the cumulative time that has not been translated, t0 is the mutation time, α is the base coefficient, and β is the mutation coefficient.

[0021] Furthermore, the probability judgment processor calculates the critical probability value P0 according to the following formula:

[0022] P0 = 1 - P'·[log2R];

[0023] Among them, P' is the jump probability value;

[0024] When the maximum probability value in the translation classification exceeds the critical probability value, it is determined that the translation timing is met.

[0025] Furthermore, the text translation unit includes a translation information database, a classification summary processor, and a text conversion processor. The translation information database is used to store translation information content. The classification summary processor is used to receive text information and summarize the corresponding translation types. The text conversion processor is used to convert the text information into text information in the target language.

[0026] The beneficial effects achieved by the present invention are:

[0027] This system intelligently analyzes the current translation content through big data, predicts the possibilities of different translation results, and performs translation in a timely manner at the appropriate time. It can effectively reduce the translation latency on the premise of ensuring translation quality and meet high - requirement translation occasions.

[0028] To further understand the features and technical content of the present invention, please refer to the following detailed description of the present invention and the attached drawings. However, the attached drawings are only provided for reference and illustration, and are not used to limit the present invention. Brief Description of the Drawings

[0029] Figure 1 It is a schematic diagram of the overall structural framework of the present invention;

[0030] Figure 2 It is a schematic diagram of the composition of the voice acquisition module of the present invention;

[0031] Figure 3 It is a schematic diagram of the composition of the voice parsing module of the present invention;

[0032] Figure 4 It is a schematic diagram of the composition of the intelligent prediction module of the present invention;

[0033] Figure 5 It is a schematic diagram of the composition of the translation output module of the present invention;

[0034] Figure 6 It is a data table of the actual test effect of the present invention. Detailed Implementation Modes

[0035] The following are specific embodiments to illustrate the implementation modes of the present invention. Those skilled in the art can understand the advantages and effects of the present invention from the content disclosed in this specification. The present invention can be implemented or applied through other different specific embodiments, and various details in this specification can also be modified and changed based on different viewpoints and applications without departing from the spirit of the present invention. Additionally, the drawings of the present invention are only for simple schematic illustration and are not drawn according to actual dimensions. This is stated in advance. The following implementation modes will further detail the related technical content of the present invention, but the disclosed content is not used to limit the protection scope of the present invention.

[0036] Embodiment 1

[0037] This embodiment provides a multi - language real - time translation system based on artificial intelligence, combined with Figure 1 , including a voice acquisition module, a voice parsing module, an intelligent prediction module, and a translation output module;

[0038] The voice acquisition module is used to acquire voice information, the voice parsing module is used to parse the voice content, the intelligent prediction module is used to perform predictive analysis on subsequent voices, and the translation output module is used to translate and output the voice content;

[0039] The voice acquisition module includes a signal acquisition unit, a noise filtering unit, and a signal caching unit. The signal acquisition unit is used to acquire sound signals, the noise filtering unit is used to filter noise signals, and the signal caching unit is used to store signal data;

[0040] The voice parsing module includes a language parsing unit, a text conversion unit, and an output control unit. The language parsing unit is used to determine the language type of the acquired voice, the text conversion unit is used to convert the voice signal into text information, and the output control unit is used to control the output process of the text information;

[0041] The intelligent prediction module includes a conversation information database, a probability prediction unit, and a timing judgment unit. The conversation information database is used to store conversation data, the probability prediction unit is used to predict the probability information of subsequent voices, and the timing judgment unit is used to judge whether it is in the translation timing;

[0042] The translation output module includes a target setting unit, a text translation unit, and a voice output unit. The target setting unit is used to set translation information, the text translation unit is used to translate the original language text into the target language text, and the voice output unit is used to play the target language voice.

[0043] The probability prediction unit includes a task generation processor, a result classification processor, and a probability calculation processor. The task generation processor generates a retrieval task based on text information. The result classification processor is used to classify the retrieval results. The probability calculation processor is used to calculate the probability information for each translation classification.

[0044] The probability calculation processor calculates the probability value P(i) of the i-th translation classification according to the following formula:

[0045]

[0046] Where n(i) is the statistical quantity of the i-th translation classification, and m is the number of translation classifications.

[0047] The timing judgment unit includes a semantic breakpoint detector, a delay tolerance analyzer, and a probability judgment processor. The semantic breakpoint detector is used to identify pauses and logical demarcation points. The delay tolerance analyzer is used to manage and analyze the time when translation has not been performed. The probability judgment processor is used to judge whether the current translation classification with the maximum probability meets the translation timing.

[0048] The delay tolerance analyzer calculates the tolerance strength R according to the following formula:

[0049]

[0050] Where t is the cumulative time when translation has not been performed, t0 is the mutation time, α is the base coefficient, and β is the mutation coefficient.

[0051] The probability judgment processor calculates the critical probability value P0 according to the following formula:

[0052] P0 = 1 - P'·[log2R];

[0053] Where P' is the jump probability value;

[0054] When the maximum probability value in the translation classification exceeds the critical probability value, it is determined that the translation timing is met.

[0055] The text translation unit includes a translation information database, a classification summary processor, and a text conversion processor. The translation information database is used to store translation information content. The classification summary processor is used to receive text information and summarize the corresponding translation types. The text conversion processor is used to convert the text information into text information in the target language.

[0056] Embodiment 2.

[0057] This embodiment includes all the contents of Embodiment 1 and provides a multi-language real-time translation system based on artificial intelligence, including a voice acquisition module, a voice parsing module, an intelligent prediction module, and a translation output module;

[0058] The voice acquisition module is used to acquire voice information, the voice parsing module is used to parse the voice content, the intelligent prediction module is used to predict and analyze subsequent voices, and the translation output module is used to translate and output the voice content;

[0059] Combined with Figure 2 , the voice acquisition module includes a signal acquisition unit, a noise filtering unit, and a signal caching unit. The signal acquisition unit is used to acquire sound signals, the noise filtering unit is used to filter out noise signals, and the signal caching unit is used to store signal data;

[0060] Combined with Figure 3 , the voice parsing module includes a language parsing unit, a text conversion unit, and an output control unit. The language parsing unit is used to determine the language type of the acquired voice, the text conversion unit is used to convert the voice signal into text information, and the output control unit is used to control the output process of the text information;

[0061] Combined with Figure 4 , the intelligent prediction module includes a conversation information database, a probability prediction unit, and a timing judgment unit. The conversation information database is used to store conversation data, the probability prediction unit is used to predict the probability information of subsequent voices, and the timing judgment unit is used to judge whether it is in the translation timing;

[0062] Combined with Figure 5 , the translation output module includes a target setting unit, a text translation unit, and a voice output unit. The target setting unit is used to set translation information, the text translation unit is used to translate the original language text into the target language text, and the voice output unit is used to play the target language voice;

[0063] The signal acquisition unit includes a microphone array collector, a signal amplification processor, and an analog-to-digital conversion processor. The microphone array collector is used to receive sound wave signals, the signal amplification processor is used to dynamically increase the signal intensity, and the analog-to-digital conversion processor is used to convert analog signals into digital signals;

[0064] The noise filtering unit includes a background noise processor, a burst noise processor, and a Gaussian filtering processor. The background noise processor is used to eliminate background noise in the voice signal, the burst noise processor is used to eliminate burst noise in the voice signal, and the Gaussian filtering processor is used to perform Gaussian filtering processing on the voice signal;

[0065] The signal buffer unit includes a signal data register, a signal output controller, and a signal deletion controller. The signal data register is used to store the filtered voice signal. The signal output controller is used to control the output process of the voice signal. The signal deletion controller is used to control the deletion process of the voice signal;

[0066] The language parsing unit includes an acoustic feature extractor, a language classification processor, and a confidence evaluation processor. The acoustic feature extractor is used to extract the acoustic features in the voice signal. The language classification processor is used to classify the language of the voice signal. The confidence evaluation processor is used to process and obtain the confidence level of the language classification;

[0067] The text conversion unit includes a word extraction processor, a feature parsing processor, and a text mapping processor. The word extraction processor is used to extract the word voice. The feature parsing processor is used to parse and obtain the feature information of the word voice. The text mapping processor is used to map the word voice to the word text;

[0068] The output control unit includes a text transmission processor, a feedback receiving processor, and a text marking processor. The text transmission processor is used to send text information outward. The feedback receiving processor is used to receive prediction information. The text marking processor is used to mark the text status information;

[0069] The conversation information database includes a conversation data register, a conversation information retrieval processor, and a retrieval statistics processor. The conversation data register is used to store conversation data. The conversation information retrieval processor is used to retrieve conversation data. The retrieval statistics processor is used to count the retrieval results;

[0070] The probability prediction unit includes a task generation processor, a result classification processor, and a probability calculation processor. The task generation processor generates a retrieval task based on the text information. The result classification processor is used to classify the retrieval results. The probability calculation processor is used to calculate the probability information for each translation classification;

[0071] The probability calculation processor calculates the probability value P(i) of the i-th translation classification according to the following formula:

[0072]

[0073] where n(i) is the statistical quantity of the i-th translation classification, and m is the number of translation classifications;

[0074] The timing judgment unit includes a semantic breakpoint detector, a delay tolerance analyzer, and a probability judgment processor. The semantic breakpoint detector is used to identify pauses and logical demarcation points. The delay tolerance analyzer is used to manage and analyze the time when translation has not been performed. The probability judgment processor is used to judge whether the current translation classification with the highest probability meets the translation timing.

[0075] The delay tolerance analyzer calculates the tolerance strength R according to the following formula:

[0076]

[0077] where t is the cumulative time when translation has not been performed, t0 is the mutation time, α is the base coefficient, and β is the mutation coefficient.

[0078] The probability judgment processor calculates the critical probability value P0 according to the following formula:

[0079] P0 = 1 - P'·[log2R];

[0080] where P' is the jump probability value.

[0081] When the maximum probability value in the translation classification exceeds the critical probability value, it is determined that the translation timing is met.

[0082] The process by which the timing judgment unit judges the translation timing includes the following steps:

[0083] S21. The semantic breakpoint detector detects whether it is at the breakpoint timing. If so, it sends a translation signal. If not, it proceeds to step S22.

[0084] S22. The delay tolerance analyzer calculates the current tolerance.

[0085] S23. The probability judgment processor calculates the current critical probability value and judges whether to send a translation signal.

[0086] The target setting unit includes a language selection processor, a start control processor, and a parameter setting processor. The language selection processor is used to set the target language to be translated. The start control processor is used to control the start and stop of translation. The parameter setting processor is used to set the playback parameters of the translated voice.

[0087] The text translation unit includes a translation information database, a classification summary processor, and a text conversion processor. The translation information database is used to store translation information content. The classification summary processor is used to receive text information and summarize the corresponding translation types. The text conversion processor is used to convert the text information into text information in the target language.

[0088] The voice output unit includes a speech synthesis processor, a parameter adjustment processor, and an audio playback processor. The speech synthesis processor synthesizes corresponding audio information based on the text information. The parameter adjustment processor adjusts the audio information based on the set parameters. The audio playback processor is used to play the translation content;

[0089] The working process of this system for translation includes the following steps:

[0090] S1. The text transmission processor sends the text information to be translated to the text translation unit;

[0091] S2. The classification and summarization processor processes the text information to obtain the translation type, and sends the translation type information to the probability prediction unit;

[0092] S3. The probability calculation processor processes to obtain the probability information of each translation type;

[0093] S4. The timing judgment unit judges whether to perform translation. If so, it sends the locked translation type to the text translation unit and enters step S5. If not, it returns to step S1;

[0094] S5. The text conversion processor converts the text information based on the locked translation type;

[0095] S6. The voice output unit plays the translated voice information;

[0096] S7. Feed back the original text information corresponding to the translated voice to the output control unit, and return to step S1;

[0097] Both i and j appearing above are ordinal numbers used to represent serial numbers and have no actual meaning.

[0098] Some code information of this system is as follows:

[0099] class SpeechTranslationSystem:

[0100] def __init__(self):

[0101] """Initialize the speech recognition, synthesis engine and translator"""

[0102] self.recognizer = sr.Recognizer()

[0103] self.microphone = sr.Microphone()

[0104] self.translator = Translator()

[0105] self.text_to_speech = pyttsx3.init()

[0106] self.dialogue_database = []

[0107] # Speech capture module

[0108] def capture_speech(self):

[0109] """Capture speech data and perform noise filtering"""

[0110] with self.microphone as source:

[0111] print("Listening, please speak...")

[0112] self.recognizer.adjust_for_ambient_noise(source)

[0113] audio = self.recognizer.listen(source)

[0114] return audio

[0115] # Speech parsing module

[0116] def process_speech(self, audio):

[0117] """Convert speech to text and parse the language"""

[0118] try:

[0119] text = self.recognizer.recognize_google(audio)

[0120] detected_language = langdetect.detect(text)

[0121] print(f"Detected language: {detected_language}|Recognized text: {text}")

[0122] return text, detected_language

[0123] except Exception as e:

[0124] print(f"Speech parsing failed: {e}")

[0125] return None, None

[0126] # Intelligent prediction module

[0127] def predict_next_speech(self, text):

[0128] """Predict the user's possible subsequent statements based on the context"""

[0129] self.dialogue_database.append(text)

[0130] if len(self.dialogue_database) > 1:

[0131] last_text = self.dialogue_database[-1]

[0132] probabilities = np.random.dirichlet(np.ones(2), size = 1)[0]

[0133] predicted_response = f"Predicted possible answer based on {last_text}"

[0134] print(f"Prediction result: {predicted_response} (Confidence:

[0135] {probabilities[0]:.2f})")

[0136] return predicted_response

[0137] return None

[0138] # Translation output module

[0139] def translate_text(self, text, target_lang = "en"):

[0140] """Translate the text and read the translation result aloud"""

[0141] try:

[0142] translated = self.translator.translate(text, dest=target_lang)

[0143] print(f"Translation result ({translated.dest}): {translated.text}")

[0144] return translated.text

[0145] except Exception as e:

[0146] print(f"Translation failed: {e}")

[0147] return None

[0148] def text_to_speech_output(self, text):

[0149] """Convert text to speech and play it"""

[0150] self.text_to_speech.say(text)

[0151] self.text_to_speech.runAndWait().

[0152] Now, 10 segments of speech are used for testing, and the pause times in each time period during the translation and playback process are counted respectively to obtain Figure 6 the data table shown below.

[0153] The content disclosed above is only the preferred and feasible embodiment of the present invention, and does not limit the protection scope of the present invention. Therefore, all equivalent technical changes made by using the content of the specification and drawings of the present invention are included in the protection scope of the present invention. In addition, with the development of technology, the elements therein can be updated.

Claims

1. A multi - language real - time translation system based on artificial intelligence, characterized in that, It includes a voice acquisition module, a voice parsing module, an intelligent prediction module, and a translation output module; The voice acquisition module is used to acquire voice information, the voice parsing module is used to parse the voice content, the intelligent prediction module is used to perform prediction analysis on subsequent voices, and the translation output module is used to translate and output the voice content; The voice acquisition module includes a signal acquisition unit, a noise filtering unit, and a signal caching unit. The signal acquisition unit is used to acquire sound signals, the noise filtering unit is used to filter noise signals, and the signal caching unit is used to store signal data; The voice parsing module includes a language parsing unit, a text conversion unit, and an output control unit. The language parsing unit is used to determine the language type to which the acquired voice belongs, the text conversion unit is used to convert the voice signal into text information, and the output control unit is used to control the output process of the text information; The intelligent prediction module includes a conversation information database, a probability prediction unit, and a timing judgment unit. The conversation information database is used to store conversation data, the probability prediction unit is used to predict the probability information of subsequent voices, and the timing judgment unit is used to judge whether it is in the translation timing; The translation output module includes a target setting unit, a text translation unit, and a voice output unit. The target setting unit is used to set translation information, the text translation unit is used to translate the original language text into the target language text, and the voice output unit is used to play the target language voice.

2. The multilingual real-time translation system based on artificial intelligence according to claim 1, wherein The probability prediction unit includes a task generation processor, a result classification processor, and a probability calculation processor. The task generation processor generates a retrieval task based on the text information, the result classification processor is used to classify and process the retrieval results, and the probability calculation processor is used to calculate the probability information of each translation classification; The probability calculation processor calculates the probability value P(i) of the i-th translation classification according to the following formula: where n(i) is the statistical quantity of the i-th translation classification, and m is the number of translation classifications.

3. The multilingual real-time translation system based on artificial intelligence as claimed in claim 2, wherein The timing judgment unit includes a semantic break point detector, a delay tolerance analyzer, and a probability judgment processor. The semantic break point detector is used to identify pauses and logical demarcation points, the delay tolerance analyzer is used to manage and analyze the time when translation has not been performed, and the probability judgment processor is used to judge whether the translation classification with the current maximum probability meets the translation timing; The delay tolerance analyzer calculates the tolerance strength R according to the following formula: where t is the accumulated time when translation has not been performed, t0 is the mutation time, α is the base coefficient, and β is the mutation coefficient.

4. A multilingual real-time translation system based on artificial intelligence according to claim 3, characterized in that, The probability judgment processor calculates the critical probability value P0 according to the following formula: P0 = 1 - P'·[log2R]; where P' is the jump probability value; When the maximum probability value in the translation classification exceeds the critical probability value, it is determined that the translation timing is met.

5. The multilingual real-time translation system based on artificial intelligence according to claim 4, characterized in that The text translation unit includes a translation information database, a classification and summarization processor, and a text conversion processor. The translation information database is used to store translation information content. The classification and summarization processor is used to receive text information and summarize the corresponding translation types. The text conversion processor is used to convert the text information into text information in the target language.

Citation Information

Patent Citations

  • A conversation translation method, apparatus, storage medium, and terminal device

    CN111241853B