A multifunctional intelligent translation method
By collecting and analyzing voice data in real time through intelligent translation devices, generating and playing translation results, the problem of low efficiency and poor accuracy of traditional translation methods is solved, and a highly efficient, convenient, multi-functional translation experience and accurate translation are achieved.
Patent Information
- Application Number
- CN202510253593.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-03-05
AI Technical Summary
Traditional translation methods are inefficient and inaccurate, and cannot adapt to changing language environments in real time. Existing intelligent translation devices cannot process real-time voice data.
The system collects voice data in real time through intelligent translation devices, analyzes and determines the data to be translated and the target language, generates translated voice and text results using a network server, and then plays and displays them through the device.
It provides an efficient and convenient multi-functional translation experience, improves translation accuracy and comprehension efficiency, enhances device flexibility and user convenience, and improves the interactive experience.
Smart Images

Figure CN120106100B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech analysis technology, and in particular to a multifunctional intelligent translation method. Background Technology
[0002] With the acceleration of globalization and international exchange, the demand for cross-language communication is gradually increasing. Traditional translation methods rely on human translators or fixed translation equipment, but these methods are often inefficient, inaccurate, and unable to adapt to changing language environments in real time. To address this issue, automatic translation technology has made significant progress in recent years, especially technologies based on speech recognition and machine translation.
[0003] Early translation devices primarily focused on text translation, relying on text input for results and unable to handle real-time voice data. However, with the development of deep learning and natural language processing technologies, speech recognition technology has matured, giving rise to intelligent translation devices capable of automatically recognizing and translating voice input into text of the target language.
[0004] Therefore, the present invention provides a multifunctional intelligent translation method. Summary of the Invention
[0005] This invention provides a multifunctional intelligent translation method. It determines the data to be translated and the target language by analyzing real-time collected voice data. Based on the data to be translated and the target language, it generates a translated voice result and a translated text result. The intelligent translation device plays the translated voice result and displays the translated text result. This provides an efficient and convenient multifunctional translation experience, improves comprehension efficiency and translation accuracy, enhances the flexibility and accuracy of intelligent translation devices, meets user needs, improves user understanding and ease of use, and enhances translation accuracy and user interaction experience.
[0006] This invention provides a multifunctional intelligent translation method, comprising:
[0007] 101: Intelligent translation devices collect voice data in real time and determine the data to be translated and the target language based on the voice data;
[0008] 102: The web server generates both spoken and text translations based on the data to be translated and the target language.
[0009] 103: The intelligent translation device plays the translated audio result and displays the translated text result.
[0010] According to the present invention, a multifunctional intelligent translation method is provided, wherein the intelligent translation device collects voice data in real time, including:
[0011] The first voice acquisition sub-device in the intelligent translation device collects user voice data emitted by the user wearing the intelligent translation device;
[0012] The second voice acquisition sub-device in the intelligent translation device collects voice data to be translated, in addition to the user's voice data emitted by the user wearing the intelligent translation device.
[0013] The real-time voice data collected by the intelligent translation device includes: user voice data collected by the first voice acquisition sub-device and voice data to be translated collected by the second voice acquisition sub-device.
[0014] According to the present invention, a multifunctional intelligent translation method is provided, which determines the data to be translated and the target language based on speech data, including:
[0015] Collect speech sample data in multiple languages. The speech sample data includes speech sample sub-data for each language, and each speech sample sub-data includes multiple speech samples.
[0016] The speech sample data is preprocessed, and features are extracted from each speech sample in the preprocessed speech sample sub-data of each language to determine the feature vector of each speech sample in the speech sample sub-data of each language.
[0017] The language feature matrix is determined based on the feature vectors of all speech samples in the speech sample sub-data for each language.
[0018] The language feature matrix of each language is input into the corresponding language model. The corresponding language model is trained based on the speech sample sub-data of each language. The language model outputs the corresponding language feature vector.
[0019] Preprocessing is performed on the user voice data and the voice data to be translated in the voice data to be translated.
[0020] Feature extraction is performed on the preprocessed user speech data to determine the user feature vector based on the user speech data. At the same time, feature extraction is performed on the preprocessed speech data to be translated to determine the translation feature vector based on the speech data to be translated.
[0021] Determine the similarity value between the user feature vector and the language feature vector of each language, and determine the language corresponding to the language feature vector with the highest similarity value as the target language;
[0022] Based on the feature vector to be translated and the language feature vectors of all languages, the second language similar sequence is determined;
[0023] Determine the set of languages to be translated based on similar sequences in second languages;
[0024] The speech data to be translated is determined based on the set of languages to be translated and the preprocessed speech data to be translated.
[0025] According to the present invention, a multifunctional intelligent translation method determines a second language similarity sequence based on the feature vector to be translated and the language feature vectors of all languages, including:
[0026] Compare the number of features in the feature vector to be translated with the number of features in the feature vector of each language;
[0027] If the number of features in the feature vector to be translated and the number of features in the language feature vector are different, embed the feature vector to be translated or the language feature vector with fewer features to make the number of features in the feature vector to be translated and the number of features in the language feature vector the same.
[0028] Calculate the similarity value between the embedded feature vector to be translated and the language feature vector of each language, sort the similarity values of the embedded feature vector to be translated and the language feature vector of all languages from largest to smallest, and determine the second language similarity sequence.
[0029]
[0030]
[0031] Where S1 represents the first language similarity sequence, a represents the feature vector to be translated after embedding the speech data to be translated, and L1, L i L N1 Let represent the language feature vector after embedding the speech sample sub-data of the 1st language, the i-th language, and the N1th language in the speech sample data. Let a represent the feature vector to be translated, L1 of the first language, and L2 of the i-th language, respectively. i The language feature vector L of the N1th language N1 The first similarity value, N1 represents the number of languages, a T This represents the transpose of the feature vector to be translated after embedding the speech data to be translated, (aL) i (aL) represents the difference vector between the embedded feature vector to be translated and the language feature vector of the i-th language. i ) T The transpose of the difference vector between the embedded feature vector to be translated and the language feature vector of the i-th language is ∑(aL i S1 represents the covariance matrix of the embedded feature vector to be translated and the language feature vector of the i-th language, α1 represents the first similarity weight, α2 represents the second similarity weight, and S2 represents the second language similarity sequence. This represents the first similarity value in the sorted first language similarity sequence. This represents the j-th first similarity value in the sorted first language similarity sequence. This represents the first similarity value of the N1th element in the sorted first language similarity sequence.
[0032] According to the present invention, a multifunctional intelligent translation method determines a set of languages to be translated based on a second language similarity sequence, including:
[0033] Based on the second language similarity sequence, determine the number of languages in the set of languages to be translated;
[0034]
[0035]
[0036] Where Nu represents the number of languages in the set of languages to be translated, CN1 j CN2 represents the first cliff value of the j-th first similarity value in the second language sequence. j δ1 represents the second cliff value of the j-th first similarity value in the second language sequence, δ2 represents the preset first cliff threshold, and δ2 represents the preset second cliff threshold. β1 represents the (j+1)th first similarity value in the sorted first language similarity sequence, β2 represents the first adjustment factor, and β2 represents the second adjustment factor;
[0037] The set of languages to be translated is determined based on the second language similarity sequence and the number of languages in the set of languages to be translated.
[0038] According to the present invention, a multifunctional intelligent translation method is provided, in which a network server generates translated speech results and translated text results based on the data to be translated and the target language, including:
[0039] If the number of languages in the set of languages to be translated in the data to be translated is greater than 1, data recognition is performed on the preprocessed speech data to be translated in the data to be translated to determine the speech sub-data to be translated for each language in the set of languages to be translated;
[0040] All languages in the set of languages to be translated, as well as the phonetic data of each language to be translated, are input into the translation model. Based on the output of the translation model, preliminary translation data is determined, which includes preliminary translated text data for each language in the set of languages to be translated.
[0041] The initial translation data is corrected to determine the corrected translation data, and the translated speech and translated text results are generated based on the corrected translation data.
[0042] According to the present invention, a multifunctional intelligent translation method is provided, which corrects preliminary translation data to determine corrected translation data, and generates translated speech results and translated text results based on the corrected translation data, including:
[0043] Based on the user feature vector of user voice data, the preliminary translation text data of each language in the preliminary translation data is corrected to determine the corrected translation text data of each language in the set of languages to be translated. The correction process includes grammatical correction, context correction and semantic correction.
[0044] The translation result is determined based on the corrected translation data for all languages;
[0045] Convert the corrected translated text data for each language in the translated text results into the corresponding translated speech data for that language;
[0046] The translation speech result is determined based on the translation speech data of all languages.
[0047] According to a multifunctional intelligent translation method provided by the present invention, an intelligent translation device plays the translated audio result and displays the translated text result, including:
[0048] The voice output sub-device of the intelligent translation device plays the translated voice result, and the display sub-device of the intelligent translation device displays the translated text result.
[0049] Compared with the prior art, the beneficial effects of this application are as follows:
[0050] By analyzing real-time collected voice data, the system determines the data to be translated and the target language. Based on these factors, it generates both a translated voice result and a translated text result. The intelligent translation device then plays the translated voice result and displays the translated text result. This provides an efficient and convenient multi-functional translation experience, improving comprehension efficiency and translation accuracy. It enhances the flexibility and precision of intelligent translation devices, aligns with user needs, improves user understanding and ease of use, and enhances translation accuracy and user interaction. Attached Figure Description
[0051] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0052] Figure 1 This is a flowchart illustrating a multifunctional intelligent translation method provided in an embodiment of the present invention. Detailed Implementation
[0053] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0054] Example 1:
[0055] This invention provides a multifunctional intelligent translation method, such as... Figure 1 As shown, including:
[0056] 101: Intelligent translation devices collect voice data in real time and determine the data to be translated and the target language based on the voice data;
[0057] 102: The web server generates both spoken and text translations based on the data to be translated and the target language.
[0058] 103: The intelligent translation device plays the translated audio result and displays the translated text result.
[0059] In this embodiment, the intelligent translation device collects the user's voice data in real time. This voice data includes the user's spoken content. Based on the collected voice data, the device analyzes and determines the data content to be translated and infers the target language the user wants to translate into.
[0060] In this embodiment, the web server uses a translation model to generate translation data based on the data to be translated and the target language.
[0061] In this embodiment, the intelligent translation device plays the translated audio result through the voice output sub-device, and displays the translated text content through the display sub-device.
[0062] In this embodiment, the multifunctional intelligent translation method can be applied to an intelligent translation earphone device, including a 4G communication module, a Bluetooth module, a WIFI module, a boost module, a battery, a MIC microphone 1, a MIC microphone 2, an RGB gas indicator, a speaker, buttons, an ESIM card, a TP module, a display module, a storage module, and a Type-C USB interface; the TWS earphone includes: a Bluetooth chip (with internal charging function), a microphone, a speaker, touch buttons, and a battery.
[0063] In this embodiment, the smart translation device can be used by placing TWS earphones inside the smart translation box and turning on the smart translation earphones by pressing a button. At this time, Person A wears the smart translation earphones and speaks in their native language A, while Person B speaks in their native language B. During the conversation, Person A presses a button on the smart translation earphones to start the translation; no Bluetooth pairing is required during the process. When Person A speaks in language A, microphones 1 and 2 collect the voice data and send it to the network server via the 4G communication module. The data is then played back through the speaker on the smart translation earphones, and the translated content is displayed on the OLED screen. Similarly, when Person B speaks in language B, Person A receives the translated content through the speaker and on the OLED screen.
[0064] In this embodiment, the smart translation device can also be used when the TWS earbuds are taken out of the smart translation box. At this time, the TWS earbuds and the Bluetooth module in the smart translation box automatically pair. The translation mode is activated by short-pressing the touch button on the TWS earbuds. At this time, person A is wearing the TWS earbuds and speaks their native language A, while person B speaks their native language B. During the conversation, person A short-presses the touch button on the TWS earbuds to start the translation. When person A speaks in language A, the microphone on the TWS earbuds captures the voice data, transmits it to the smart translation box via Bluetooth, and then sends it to the network server via the 4G communication module. The translation is then played back through the speakers on the smart translation earbuds and the TWS earbuds, while the translated content is displayed on the OLED screen. Similarly, when person B speaks in language B, person A sees the translated content through the speaker and on the OLED screen.
[0065] In this embodiment, the TWS earbuds are placed inside the smart translation box. The smart translation earbuds are powered on by pressing a button. At this point, the smart translation box contains an eSIM, an OLED screen, a TP (Transfer Provider Interface), and a 4G communication module. The smart translation box can be used like a smartphone, allowing users to make calls, listen to music, and download apps.
[0066] The beneficial effects of the above technical solution are as follows: By analyzing real-time collected voice data, the data to be translated and the target language are determined. Based on the data to be translated and the target language, a translated voice result and a translated text result are generated. The intelligent translation device plays the translated voice result and displays the translated text result. This provides an efficient, convenient, and multifunctional translation experience, improves comprehension efficiency and translation accuracy, enhances the flexibility and accuracy of intelligent translation devices, meets user needs, improves user understanding and ease of use, and enhances translation accuracy and user interaction experience.
[0067] Example 2:
[0068] This invention provides a multifunctional intelligent translation method, in which an intelligent translation device collects voice data in real time, including:
[0069] The first voice acquisition sub-device in the intelligent translation device collects user voice data emitted by the user wearing the intelligent translation device;
[0070] The second voice acquisition sub-device in the intelligent translation device collects voice data to be translated, in addition to the user's voice data emitted by the user wearing the intelligent translation device.
[0071] The real-time voice data collected by the intelligent translation device includes: user voice data collected by the first voice acquisition sub-device and voice data to be translated collected by the second voice acquisition sub-device.
[0072] In this embodiment, the first voice acquisition sub-device (such as MIC microphone 1) equipped with the intelligent translation device is used to collect the voice data of the user wearing the device in real time. This voice data is usually the voice commands or conversation content issued by the user.
[0073] In this embodiment, the intelligent translation device also includes a second voice acquisition sub-device (such as a microphone MIC 2) for acquiring other voice data to be translated besides the user's voice, which is the voice spoken by other people in the environment. Through this sub-device, the translation device can receive the voice content of other speakers.
[0074] In this embodiment, the voice data acquisition function of the intelligent translation device integrates the inputs of the first and second voice acquisition sub-devices. The device collects the wearer's voice data and the voice data to be translated in the surrounding environment in real time, forming a complete voice dataset.
[0075] The beneficial effects of the above technical solution are: intelligent translation devices collect voice data in real time, which can provide data basis for determining the data to be translated and the target language.
[0076] Example 3:
[0077] This invention provides a multifunctional intelligent translation method that determines the data to be translated and the target language based on voice data, including:
[0078] Collect speech sample data in multiple languages. The speech sample data includes speech sample sub-data for each language, and each speech sample sub-data includes multiple speech samples.
[0079] The speech sample data is preprocessed, and features are extracted from each speech sample in the preprocessed speech sample sub-data of each language to determine the feature vector of each speech sample in the speech sample sub-data of each language.
[0080] The language feature matrix is determined based on the feature vectors of all speech samples in the speech sample sub-data for each language.
[0081] The language feature matrix of each language is input into the corresponding language model. The corresponding language model is trained based on the speech sample sub-data of each language. The language model outputs the corresponding language feature vector.
[0082] Preprocessing is performed on the user voice data and the voice data to be translated in the voice data to be translated.
[0083] Feature extraction is performed on the preprocessed user speech data to determine the user feature vector based on the user speech data. At the same time, feature extraction is performed on the preprocessed speech data to be translated to determine the translation feature vector based on the speech data to be translated.
[0084] Determine the similarity value between the user feature vector and the language feature vector of each language, and determine the language corresponding to the language feature vector with the highest similarity value as the target language;
[0085] Based on the feature vector to be translated and the language feature vectors of all languages, the second language similar sequence is determined;
[0086] Determine the set of languages to be translated based on similar sequences in second languages;
[0087] The speech data to be translated is determined based on the set of languages to be translated and the preprocessed speech data to be translated.
[0088] In this embodiment, speech sample data in multiple languages is collected. This data includes sub-data in different languages, and each sub-data in each language contains multiple speech samples. These speech samples are preprocessed, such as denoising and normalization, to ensure data quality.
[0089] In this embodiment, feature extraction is performed on each speech sample in the preprocessed speech sample data to extract the feature vector of each sample. Based on the feature vectors of all speech samples, a language feature matrix is generated for each language, representing the typical features of that language.
[0090] In this embodiment, the language feature matrix of each language is input into the corresponding language model, and the language model is trained based on these data. The language model will output the language feature vector of that language.
[0091] In this embodiment, the user's speech and the speech data to be translated are preprocessed, including noise removal and segmentation. Then, features are extracted from these two parts of data to generate user feature vectors and translation feature vectors.
[0092] In this embodiment, the similarity between the user feature vector and the feature vectors of each language is calculated, and the language with the highest similarity is found as the target language.
[0093] The beneficial effects of the above technical solution are as follows: by determining the data to be translated and the target language based on voice data, accurate voice recognition can be achieved, the target language can be intelligently determined, the accuracy and efficiency of translation can be improved, and the multilingual adaptability and intelligence level of intelligent translation equipment can be enhanced.
[0094] Example 4:
[0095] This invention provides a multifunctional intelligent translation method that determines a second language similarity sequence based on the feature vector to be translated and the language feature vectors of all languages, including:
[0096] Compare the number of features in the feature vector to be translated with the number of features in the feature vector of each language;
[0097] If the number of features in the feature vector to be translated and the number of features in the language feature vector are different, embed the feature vector to be translated or the language feature vector with fewer features to make the number of features in the feature vector to be translated and the number of features in the language feature vector the same.
[0098] Calculate the similarity value between the embedded feature vector to be translated and the language feature vector of each language, sort the similarity values of the embedded feature vector to be translated and the language feature vector of all languages from largest to smallest, and determine the second language similarity sequence.
[0099]
[0100]
[0101] Where S1 represents the first language similarity sequence, a represents the feature vector to be translated after embedding the speech data to be translated, and L1, L i L N1 Let represent the language feature vector after embedding the speech sample sub-data of the 1st language, the i-th language, and the N1th language in the speech sample data. Let a represent the feature vector to be translated, L1 of the first language, and L2 of the i-th language, respectively. i The language feature vector L of the N1th language N1 The first similarity value, N1 represents the number of languages, a T This represents the transpose of the feature vector to be translated after embedding the speech data to be translated, (aL) i (aL) represents the difference vector between the embedded feature vector to be translated and the language feature vector of the i-th language. i ) T The transpose of the difference vector between the embedded feature vector to be translated and the language feature vector of the i-th language is ∑(aL iS1 represents the covariance matrix of the embedded feature vector to be translated and the language feature vector of the i-th language, α1 represents the first similarity weight, α2 represents the second similarity weight, and S2 represents the second language similarity sequence. This represents the first similarity value in the sorted first language similarity sequence. This represents the j-th first similarity value in the sorted first language similarity sequence. This represents the first similarity value of the N1th element in the sorted first language similarity sequence.
[0102] In this embodiment, since speech data from different languages may have different feature dimensions, the number of features in the feature vector to be translated is compared with the number of features in the feature vector of each language.
[0103] In this embodiment, if the number of features in the feature vector to be translated is different from the number of features in the language feature vector, the vector with fewer features is embedded. The embedding process is carried out in a certain way (such as padding with zeros) to make the number of features in the feature vector to be translated and the number of features in the language feature vector consistent.
[0104] In this embodiment, once the dimensions of the feature vector to be translated and the language feature vector are consistent, the similarity between them is calculated, and these similarity values are sorted from largest to smallest to generate a second language similarity sequence, that is, the similarity ranking of all languages with the data to be translated.
[0105] In this embodiment, This represents the covariance similarity value between the embedded feature vector to be translated and the language feature vector of the i-th language.
[0106] In this embodiment, This represents the cosine similarity value between the embedded feature vector to be translated and the language feature vector of the i-th language.
[0107] In this embodiment, the first weight α1 represents the weight of the similarity between the embedded feature vector to be translated and the language feature vector based on the cosine similarity value, and the second weight α2 represents the weight of the similarity between the embedded feature vector to be translated and the language feature vector based on the covariance similarity value.
[0108] The beneficial effects of the above technical solution are as follows: Based on the feature vector to be translated and the language feature vectors of all languages, the similarity sequence of the second language can be determined, which can provide data basis for the set of languages to be translated, improve the accuracy of language matching in the translation process, and provide more efficient and accurate translation results.
[0109] Example 5:
[0110] This invention provides a multifunctional intelligent translation method that determines a set of languages to be translated based on similar sequences of second languages, including:
[0111] Based on the second language similarity sequence, determine the number of languages in the set of languages to be translated;
[0112]
[0113] Where Nu represents the number of languages in the set of languages to be translated, CN1 j CN2 represents the first cliff value of the j-th first similarity value in the second language sequence. j δ1 represents the second cliff value of the j-th first similarity value in the second language sequence, δ2 represents the preset first cliff threshold, and δ2 represents the preset second cliff threshold. β1 represents the (j+1)th first similarity value in the sorted first language similarity sequence, β2 represents the first adjustment factor, and β2 represents the second adjustment factor;
[0114] The set of languages to be translated is determined based on the second language similarity sequence and the number of languages in the set of languages to be translated.
[0115] In this embodiment, This represents the difference between the j-th largest first similarity value and the (j+1)-th largest first similarity value in the first language similarity sequence.
[0116] The first adjustment factor β1 adjusts the difference between two adjacent first similarity values in the second language similarity sequence in the first cliff value.
[0117] The second adjustment factor β2 adjusts the average of the j-th first similarity value and the two first similarity values adjacent to the j-th first similarity value in the second language similarity sequence in the first cliff value.
[0118] The beneficial effects of the above technical solution are as follows: determining the set of languages to be translated based on the similarity sequence of the second language can provide a data basis for determining the speech data to be translated, thereby improving the accuracy and efficiency of translation and enhancing the multilingual adaptability and intelligence level of intelligent translation devices.
[0119] Example 6:
[0120] This invention provides a multifunctional intelligent translation method, in which a web server generates translated speech and translated text results based on the data to be translated and the target language, including:
[0121] If the number of languages in the set of languages to be translated in the data to be translated is greater than 1, data recognition is performed on the preprocessed speech data to be translated in the data to be translated to determine the speech sub-data to be translated for each language in the set of languages to be translated;
[0122] All languages in the set of languages to be translated, as well as the phonetic data of each language to be translated, are input into the translation model. Based on the output of the translation model, preliminary translation data is determined, which includes preliminary translated text data for each language in the set of languages to be translated.
[0123] The initial translation data is corrected to determine the corrected translation data, and the translated speech and translated text results are generated based on the corrected translation data.
[0124] In this embodiment, if the data to be translated contains multiple languages (i.e., the number of languages in the set of languages to be translated is greater than 1), the speech data to be translated is first identified to identify the speech content of each language in the data to be translated, and the speech data of each language is divided into corresponding speech sub-data to be translated.
[0125] In this embodiment, all identified languages and their corresponding speech data to be translated are input into the translation model for translation. The translation model generates preliminary translation results based on the input speech data. The preliminary translation data includes the translated text data corresponding to each language.
[0126] The beneficial effects of the above technical solution are as follows: the network server generates translated speech and translated text results based on the data to be translated and the target language, which can improve the accuracy of the translation results, ensure that the generated speech and text translation results meet the user's needs, enhance the flexibility and accuracy of intelligent translation devices, and improve the user experience.
[0127] Example 7:
[0128] This invention provides a multifunctional intelligent translation method that corrects preliminary translation data to determine corrected translation data, and generates translated speech and translated text results based on the corrected translation data, including:
[0129] Based on the user feature vector of user voice data, the preliminary translation text data of each language in the preliminary translation data is corrected to determine the corrected translation text data of each language in the set of languages to be translated. The correction process includes grammatical correction, context correction and semantic correction.
[0130] The translation result is determined based on the corrected translation data for all languages;
[0131] Convert the corrected translated text data for each language in the translated text results into the corresponding translated speech data for that language;
[0132] The translation speech result is determined based on the translation speech data of all languages.
[0133] In this embodiment, a user feature vector is generated using the user's voice data, and the preliminary translated text for each language in the preliminary translation data is corrected based on the feature vector. These corrections include: grammatical correction: correcting grammatical errors to make the sentences conform to the grammatical structure of the user's voice data; context correction: adjusting the translation results according to the context of the user's voice data to ensure that the translation is fluent and natural; and semantic correction: modifying the word selection according to the context of the user's voice data to avoid mistranslation or inaccurate word meanings.
[0134] In this embodiment, after correcting the initial translated text data for each language, corrected translated text data for each language is generated. These corrected translated text data are more accurate and conform to language habits.
[0135] In this embodiment, the corrected translation data for all languages are aggregated to generate the final translated text result.
[0136] In this embodiment, the corrected translated text data for each language in the translated text results is converted into translated speech data for the corresponding language.
[0137] In this embodiment, the translation speech data of all languages are merged to finally generate a complete translation speech result.
[0138] The beneficial effects of the above technical solution are as follows: by correcting the preliminary translation data and determining the corrected translation data, and generating the translated speech and translated text results based on the corrected translation data, the accuracy and fluency of the translation quality can be improved, making it more natural and in line with user needs, thereby enhancing user experience and translation accuracy.
[0139] Example 8:
[0140] This invention provides a multifunctional intelligent translation method, in which an intelligent translation device plays the translated audio result and displays the translated text result, including:
[0141] The voice output sub-device of the intelligent translation device plays the translated voice result, and the display sub-device of the intelligent translation device displays the translated text result.
[0142] In this embodiment, the intelligent translation device is equipped with a voice output sub-device, such as a speaker, for playing the translated voice result. The device generates audio from the translated voice data using speech synthesis technology and plays it out through the output sub-device, allowing the user to hear the translated content.
[0143] In this embodiment, the intelligent translation device also includes a display sub-device, such as a screen or monitor, for displaying the translated text result. The translated text is presented to the user in a visual form, making it convenient for the user to view and understand the translated content.
[0144] The beneficial effects of the above technical solution are: intelligent translation devices can play the translated audio results and display the translated text results, which can enhance the user's understanding and ease of use, improve translation accuracy and user interaction experience.
[0145] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0146] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0147] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multifunctional intelligent translation method, characterized in that, include: 101: Intelligent translation devices collect voice data in real time and determine the data to be translated and the target language based on the voice data; 102: The web server generates both spoken and text translations based on the data to be translated and the target language. 103: The intelligent translation device plays the translated audio result and displays the translated text result; The voice data collected in real time by the intelligent translation device includes: user voice data collected by the first voice acquisition sub-device and voice data to be translated collected by the second voice acquisition sub-device. The data to be translated and the target language are determined based on the speech data, including: Preprocess the user's voice data and the voice data to be translated; Feature extraction is performed on the preprocessed user speech data to determine the user feature vector based on the user speech data. At the same time, feature extraction is performed on the preprocessed speech data to be translated to determine the translation feature vector based on the speech data to be translated. Determine the similarity value between the user feature vector and the language feature vector of each language, and determine the language corresponding to the language feature vector with the highest similarity value as the target language; Based on the feature vector to be translated and the language feature vectors of all languages, the second language similarity sequence is determined, including: Calculate the similarity value between the embedded feature vector to be translated and the language feature vector of each language, sort the similarity values of the embedded feature vector to be translated and the language feature vector of all languages from largest to smallest, and determine the second language similarity sequence. ; Where S1 represents the first language similarity sequence, and a represents the feature vector to be translated after embedding the speech data to be translated. Let represent the language feature vector after embedding the speech sample sub-data of the 1st language, the i-th language, and the N1th language in the speech sample data. Let a represent the feature vector to be translated and the language feature vector of the first language, respectively. The language feature vector of the i-th language The language feature vector of the N1th language The first similarity value, where N1 represents the number of languages. This represents the transpose of the feature vector to be translated after embedding the speech data. This represents the difference vector between the embedded feature vector to be translated and the language feature vector of the i-th language. This represents the transpose of the difference vector between the embedded feature vector to be translated and the language feature vector of the i-th language. Let represent the covariance matrix of the embedded feature vector to be translated and the language feature vector of the i-th language. Indicates the first similarity weight. S1 represents the second similarity weight, and S2 represents the second language similarity sequence. This represents the first similarity value in the sorted first language similarity sequence. This represents the j-th first similarity value in the sorted first language similarity sequence. This represents the first similarity value of the N1th element in the sorted first language similarity sequence; the set of languages to be translated is determined based on the second language similarity sequence, including: Based on the second language similarity sequence, determine the number of languages in the set of languages to be translated; in, This indicates the number of languages in the set to be translated. This represents the first cliff value of the j-th first similarity value in the second language sequence. This indicates the preset first cliff threshold. This represents the (j+1)th first similarity value in the sorted first language similarity sequence. Indicates the first adjustment factor. Indicates the second adjustment factor; The set of languages to be translated is determined based on the second language similarity sequence and the number of languages in the set of languages to be translated.
2. The multifunctional intelligent translation method according to claim 1, characterized in that, Intelligent translation devices collect voice data in real time, including: The first voice acquisition sub-device in the intelligent translation device collects user voice data emitted by the user wearing the intelligent translation device; The second voice acquisition sub-device in the intelligent translation device collects voice data to be translated, in addition to the user's voice data emitted by the user wearing the intelligent translation device.
3. The multifunctional intelligent translation method according to claim 2, characterized in that, Determining the data to be translated and the target language based on voice data also includes: Collect speech sample data in multiple languages. The speech sample data includes speech sample sub-data for each language, and each speech sample sub-data includes multiple speech samples. The speech sample data is preprocessed, and features are extracted from each speech sample in the preprocessed speech sample sub-data of each language to determine the feature vector of each speech sample in the speech sample sub-data of each language. The language feature matrix is determined based on the feature vectors of all speech samples in the speech sample sub-data for each language. The language feature matrix of each language is input into the corresponding language model. The corresponding language model is trained based on the speech sample sub-data of each language. The language model outputs the corresponding language feature vector. Based on the feature vector to be translated and the language feature vectors of all languages, the second language similar sequence is determined; Determine the set of languages to be translated based on similar sequences in second languages; The speech data to be translated is determined based on the set of languages to be translated and the preprocessed speech data to be translated.
4. The multifunctional intelligent translation method according to claim 3, characterized in that, Based on the feature vector to be translated and the language feature vectors of all languages, the second language similarity sequence is determined, which also includes: Compare the number of features in the feature vector to be translated with the number of features in the feature vector of each language; If the number of features in the feature vector to be translated and the number of features in the language feature vector are not the same, embed the feature vector to be translated or the language feature vector with fewer features to make the number of features in the feature vector to be translated and the number of features in the language feature vector the same.
5. The multifunctional intelligent translation method according to claim 3, characterized in that, The web server generates translated speech and translated text results based on the data to be translated and the target language, including: If the number of languages in the set of languages to be translated in the data to be translated is greater than 1, data recognition is performed on the preprocessed speech data to be translated in the data to be translated to determine the speech sub-data to be translated for each language in the set of languages to be translated; All languages in the set of languages to be translated, as well as the phonetic data of each language to be translated, are input into the translation model. Based on the output of the translation model, preliminary translation data is determined, which includes preliminary translated text data for each language in the set of languages to be translated. The initial translation data is corrected to determine the corrected translation data, and the translated speech and translated text results are generated based on the corrected translation data.
6. The multifunctional intelligent translation method according to claim 5, characterized in that, The initial translation data is corrected to determine the corrected translation data. Based on the corrected translation data, the translated speech result and the translated text result are generated, including: The preliminary translation text data for each language in the preliminary translation data is corrected to determine the corrected translation text data for each language in the set of languages to be translated. The correction process includes grammatical correction, contextual correction, and semantic correction. The translation result is determined based on the corrected translation data for all languages; Convert the corrected translated text data for each language in the translated text results into the corresponding translated speech data for that language; The translation speech result is determined based on the translation speech data of all languages.
7. The multifunctional intelligent translation method according to claim 1, characterized in that, The intelligent translation device plays the translated audio result and displays the translated text result, including: The voice output sub-device of the intelligent translation device plays the translated voice result, and the display sub-device of the intelligent translation device displays the translated text result.
Citation Information
Patent Citations
Pre-training method and device of intelligent translation model and storage medium
CN111460838A
Translation method and device, equipment and storage medium
CN116702801A