Multifunctional intelligent translation method

Through intelligent translation devices, the real-time collection and analysis of voice data and generation and playback of translation results are solved, and the problems of low efficiency and poor accuracy of traditional translation methods are achieved, achieving an efficient and convenient multi-function translation experience.

CN120106100AActive Publication Date: 2025-06-06SHENZHEN K FREE WIRELESS INFORMATION TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510253593.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-06
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

Traditional translation methods are inefficient and have poor accuracy, and cannot adapt to variable locale environments in real time, and cannot process real-time voice data.

Method used

The intelligent translation device collects voice data in real time, analyzes and determines the data to be translated and the target language, uses the network server to generate translated voice and text results, and plays and displays them on the device.

Benefits of technology

Provide an efficient and convenient multi-functional translation experience, improve understanding efficiency and translation accuracy, enhance the flexibility and accuracy of the equipment, meet user needs, and improve translation accuracy and user interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106100A_ABST
    Figure CN120106100A_ABST
Patent Text Reader

Abstract

The invention provides a multifunctional intelligent translation method, and belongs to the technical field of voice analysis, and the method comprises the steps: intelligent translation equipment collects voice data in real time, and determines to-be-translated data and a target language based on the voice data; the network server generates a translation voice result and a translation character result based on the to-be-translated data and the target language; and the intelligent translation device plays the translation voice result and displays the translation character result. Efficient and convenient multifunctional translation experience can be provided, the understanding efficiency and translation precision are improved, the flexibility and precision of intelligent translation equipment are enhanced, user requirements are met, user understanding and use convenience is enhanced, and translation accuracy and user interaction experience are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech analysis, and in particular to a multifunctional intelligent translation method. Background Art

[0002] With the acceleration of globalization and international communication, the demand for cross-language communication is gradually increasing. Traditional translation methods rely on manual translation or fixed translation equipment, but these methods are often inefficient, inaccurate, and cannot adapt to changing language environments in real time. To solve this problem, automatic translation technology has made significant progress in recent years, especially technology based on speech recognition and machine translation.

[0003] Early translation devices mainly focused on text translation, and the translation results relied on text input and could not process real-time voice data. However, with the development of deep learning and natural language processing technology, voice recognition technology has gradually matured, and intelligent translation devices have emerged, which can automatically recognize and translate text into the target language through voice input.

[0004] Therefore, the present invention provides a multifunctional intelligent translation method. Summary of the invention

[0005] The present invention provides a multifunctional intelligent translation method, which determines the data to be translated and the target language by analyzing the voice data collected in real time, generates the translation voice result and the translation text result according to the data to be translated and the target language, and the intelligent translation device plays the translation voice result and displays the translation text result. It can provide an efficient and convenient multifunctional translation experience, improve the understanding efficiency and translation accuracy, enhance the flexibility and accuracy of the intelligent translation device, meet the needs of users, enhance the convenience of user understanding and use, and improve the translation accuracy and user interaction experience.

[0006] The present invention provides a multifunctional intelligent translation method, comprising:

[0007] 101: The intelligent translation device collects voice data in real time and determines the data to be translated and the target language based on the voice data;

[0008] 102: The network server generates a translation voice result and a translation text result based on the data to be translated and the target language;

[0009] 103: The intelligent translation device plays the translation voice result and displays the translation text result.

[0010] According to a multifunctional intelligent translation method provided by the present invention, an intelligent translation device collects voice data in real time, including:

[0011] The first voice collection sub-device in the intelligent translation device collects user voice data emitted by a user wearing the intelligent translation device;

[0012] The second voice collection sub-device in the intelligent translation device collects the voice data to be translated in addition to the user voice data emitted by the user wearing the intelligent translation device;

[0013] The voice data collected in real time by the intelligent translation device includes: user voice data collected by the first voice collection sub-device and voice data to be translated collected by the second voice collection sub-device.

[0014] According to a multifunctional intelligent translation method provided by the present invention, the data to be translated and the target language are determined based on the speech data, including:

[0015] Collecting speech sample data in multiple languages, wherein the speech sample data includes speech sample sub-data in each language, and the speech sample sub-data includes multiple speech samples;

[0016] Preprocessing the speech sample data, extracting features from each speech sample in the preprocessed speech sample sub-data of each language, and determining a feature vector of each speech sample in the speech sample sub-data of each language;

[0017] Determine a language feature matrix based on the feature vectors of all speech samples in the speech sample sub-data of each language;

[0018] Inputting the language feature matrix of each language into the corresponding language model, training the corresponding language model based on the speech sample sub-data of each language, and the language model outputting the corresponding language feature vector;

[0019] Preprocessing the user voice data and the voice data to be translated in the voice data to be translated;

[0020] Performing feature extraction on the preprocessed user voice data to determine a user feature vector based on the user voice data, and at the same time, performing feature extraction on the preprocessed voice data to be translated to determine a feature vector to be translated based on the voice data to be translated;

[0021] Determine the similarity value between the user feature vector and the language feature vector of each language, and determine the language corresponding to the language feature vector with the highest similarity value as the target language;

[0022] Determine a similar sequence in the second language based on the feature vector to be translated and the language feature vectors of all languages;

[0023] Determine a set of languages ​​to be translated based on similar sequences in the second language;

[0024] The speech data to be translated is determined based on the language set to be translated and the preprocessed speech data to be translated.

[0025] According to a multifunctional intelligent translation method provided by the present invention, based on the feature vector to be translated and the language feature vectors of all languages, a similar sequence in the second language is determined, including:

[0026] Comparing the number of features of the feature vector to be translated and the number of features of the language feature vector of each language;

[0027] If the feature numbers of the feature vector to be translated and the feature numbers of the language feature vector are different, embedding the feature vector to be translated or the language feature vector with fewer features so that the feature numbers of the feature vector to be translated and the language feature vector are the same;

[0028] Calculating the similarity values ​​of the embedded feature vector to be translated and the language feature vector of each language, sorting the similarity values ​​of the embedded feature vector to be translated and the language feature vectors of all languages ​​from large to small, and determining a similar sequence of the second language;

[0029]

[0030]

[0031] Among them, S1 represents the similar sequence of the first language, a represents the feature vector to be translated after the speech data to be translated is embedded, and L 1 , L i , L N1 represents the language feature vector after the speech sample sub-data of the 1st language, the i-th language, and the N1-th language in the speech sample data are embedded, Represent the feature vector a to be translated and the language feature vector L of the first language respectively 1 , the language feature vector L of the i-th language i , the language feature vector L of the N1th language N1 The first similarity value, N1 represents the number of languages, a T represents the transposition of the feature vector to be translated after the voice data to be translated is embedded, (aL i ) represents the difference vector between the embedded feature vector to be translated and the language feature vector of the i-th language, (aL i ) T represents the transposition of the difference vector between the embedded feature vector to be translated and the language feature vector of the i-th language, ∑(aL i ) represents the covariance matrix of the embedded feature vector to be translated and the language feature vector of the i-th language, α1 represents the first similarity weight, α2 represents the second similarity weight, S2 represents the second language similarity sequence, Indicates the first similarity value in the sorted first language similarity sequence. represents the jth first similarity value in the sorted first language similarity sequence, Indicates the first similarity value of the N1th similarity sequence in the sorted first language.

[0032] According to a multifunctional intelligent translation method provided by the present invention, a set of languages ​​to be translated is determined based on a similar sequence of a second language, including:

[0033] Determine the number of languages ​​in the language set to be translated based on similar sequences in the second language;

[0034]

[0035]

[0036] Among them, Nu represents the number of languages ​​in the language set to be translated, CN1 j represents the first cliff value of the jth first similarity value in the second language sequence, CN2 j represents the second cliff value of the j-th first similarity value in the second language sequence, δ1 represents the preset first cliff threshold, δ2 represents the preset second cliff threshold, represents the j+1th first similarity value in the sorted first language similarity sequence, β1 represents the first adjustment factor, and β2 represents the second adjustment factor;

[0037] The language set to be translated is determined based on the second language similar sequence and the number of languages ​​in the language set to be translated.

[0038] According to a multifunctional intelligent translation method provided by the present invention, a network server generates a translation speech result and a translation text result based on the data to be translated and the target language, including:

[0039] If the number of languages ​​in the language set to be translated in the data to be translated is greater than 1, performing data recognition on the preprocessed speech data to be translated in the data to be translated to determine speech sub-data to be translated of each language in the language set to be translated;

[0040] Inputting all languages ​​in the language set to be translated and the speech sub-data to be translated of each language into the translation model, and determining preliminary translation data based on the output result of the translation model, wherein the preliminary translation data includes preliminary translation text data of each language in the language set to be translated;

[0041] The preliminary translation data is corrected to determine the corrected translation data, and a translation speech result and a translation text result are generated based on the corrected translation data.

[0042] According to a multifunctional intelligent translation method provided by the present invention, data correction is performed on preliminary translation data to determine corrected translation data, and a translation speech result and a translation text result are generated based on the corrected translation data, including:

[0043] Based on the user feature vector of the user voice data, the preliminary translation text data of each language in the preliminary translation data is corrected to determine the corrected translation text data of each language in the set of languages ​​to be translated, wherein the correction processing includes grammar correction, context correction and word meaning correction;

[0044] Determine the translation text result based on the revised translation text data in all languages;

[0045] Converting the corrected translation text data of each language in the translation text result into translation voice data of the corresponding language;

[0046] The translated speech result is determined based on the translated speech data of all languages.

[0047] According to a multifunctional intelligent translation method provided by the present invention, an intelligent translation device plays a translation voice result and displays a translation text result, including:

[0048] The voice output sub-device of the intelligent translation device plays the translated voice result, and the display sub-device of the intelligent translation device displays the translated text result.

[0049] Compared with the prior art, the present invention has the following beneficial effects:

[0050] By analyzing the real-time collected voice data to determine the data to be translated and the target language, the translated voice results and translated text results are generated according to the data to be translated and the target language. The intelligent translation device plays the translated voice results and displays the translated text results. It can provide an efficient and convenient multi-functional translation experience, improve understanding efficiency and translation accuracy, enhance the flexibility and accuracy of intelligent translation devices, meet user needs, enhance user understanding and convenience of use, and improve translation accuracy and user interaction experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0052] Figure 1 It is a flowchart of a multifunctional intelligent translation method provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0054] Embodiment 1:

[0055] The embodiment of the present invention provides a multifunctional intelligent translation method, such as Figure 1 As shown, including:

[0056] 101: The intelligent translation device collects voice data in real time and determines the data to be translated and the target language based on the voice data;

[0057] 102: The network server generates a translation voice result and a translation text result based on the data to be translated and the target language;

[0058] 103: The intelligent translation device plays the translation voice result and displays the translation text result.

[0059] In this embodiment, the intelligent translation device collects the user's voice data in real time. The voice data includes the user's speech content. Based on the collected voice data, the device analyzes and determines the data content to be translated, and infers the target language that the user wants to translate into.

[0060] In this embodiment, the network server generates translation data based on the data to be translated and the target language using the translation model.

[0061] In this embodiment, the intelligent translation device plays the translated voice result through the voice output sub-device, and displays the translated text content through the display sub-device.

[0062] In this embodiment, the multifunctional intelligent translation method can be applied to an intelligent translation headset device, including a 4G communication module, a Bluetooth module, a WIFI module, a boost module, a battery, a MIC microphone 1, a MIC microphone 2, an RGB indicator light, a speaker, buttons, an ESIM card, a TP module, a display module, a storage module, and a Type-c USB interface; the TWS headset includes: a Bluetooth chip (with internal charging function), a microphone, a speaker, a touch button, and a battery.

[0063] In this embodiment, the application method of the intelligent translation device can be that the TWS headset is placed in the intelligent translation box, and the intelligent translation headset device is turned on by pressing a button. At this time, A carries the intelligent translation headset device, A speaks native language A, and B speaks native language B. During the communication, A presses the button on the intelligent translation headset device to start the translation, and no Bluetooth pairing operation is required in the intermediate process. When A speaks in language A, MIC microphone 1 and MIC microphone 2 collect the voice data, send it to the network server through the 4G communication module, and play it through the speaker on the intelligent translation headset device. At the same time, the translated content will be displayed on the OLED screen, similar to when B speaks in language B, A will display the translated content through the speaker and OLED screen.

[0064] In this embodiment, the application mode of the smart translation device can also be that when the TWS headset is taken out of the smart translation box, the TWS headset and the Bluetooth module in the smart translation box are automatically paired, and the translation mode is started by short pressing the touch button of the TWS headset. At this time, A wears a TWS headset, A speaks native language A, and B speaks native language B. During the communication, A short presses the touch button on the TWS headset to start the translation. When A speaks in language A, the MIC microphone on the TWS headset collects the voice data, transmits it to the smart translation box through the Bluetooth module, and then sends it to the network server through the 4G communication module. It is played through the speaker on the smart translation headset device and the speaker inside the TWS headset. At the same time, the translated content will be displayed on the OLED screen, similar to when B speaks in language B, A will display the translated content through the speaker and the OLED screen.

[0065] In this embodiment, the TWS headset is placed in the smart translation box, and the smart translation headset device is turned on by pressing a button. At this time, the smart translation box contains ESIM, OLED screen, TP, and 4G communication modules. The smart translation box can be used as a smart phone, and can be used to make calls, listen to music, and download APPs.

[0066] The beneficial effects of the above technical solution are: by analyzing the real-time collected voice data to determine the data to be translated and the target language, the translated voice results and translated text results are generated according to the data to be translated and the target language, and the intelligent translation device plays the translated voice results and displays the translated text results. It can provide an efficient and convenient multi-functional translation experience, improve understanding efficiency and translation accuracy, enhance the flexibility and accuracy of intelligent translation devices, meet user needs, enhance user understanding and convenience of use, and improve translation accuracy and user interaction experience.

[0067] Embodiment 2:

[0068] The embodiment of the present invention provides a multifunctional intelligent translation method, wherein an intelligent translation device collects voice data in real time, including:

[0069] The first voice collection sub-device in the intelligent translation device collects user voice data emitted by a user wearing the intelligent translation device;

[0070] The second voice collection sub-device in the intelligent translation device collects the voice data to be translated in addition to the user voice data emitted by the user wearing the intelligent translation device;

[0071] The voice data collected in real time by the intelligent translation device includes: user voice data collected by the first voice collection sub-device and voice data to be translated collected by the second voice collection sub-device.

[0072] In this embodiment, the first voice collection sub-device (such as MIC microphone 1) equipped with the intelligent translation device is used to collect voice data of the user wearing the device in real time. This voice data is usually a voice command or conversation content issued by the user.

[0073] In this embodiment, the intelligent translation device also includes a second voice collection sub-device (such as MIC microphone 2), which is used to collect other voice data to be translated in addition to the voice of the user wearing the device, which is the voice of other people in the environment. Through this sub-device, the translation device can receive the voice content of other speakers.

[0074] In this embodiment, the voice data collection function of the intelligent translation device integrates the input of the first and second voice collection sub-devices. The device collects the wearer's voice data and the voice data to be translated in the surrounding environment in real time to form a complete voice data set.

[0075] Beneficial effects of the above technical solution: The intelligent translation device collects voice data in real time, which can provide data basis for determining the data to be translated and the target language.

[0076] Embodiment 3:

[0077] The embodiment of the present invention provides a multifunctional intelligent translation method, which determines the data to be translated and the target language based on the speech data, including:

[0078] Collecting speech sample data in multiple languages, wherein the speech sample data includes speech sample sub-data in each language, and the speech sample sub-data includes multiple speech samples;

[0079] Preprocessing the speech sample data, extracting features from each speech sample in the preprocessed speech sample sub-data of each language, and determining a feature vector of each speech sample in the speech sample sub-data of each language;

[0080] Determine a language feature matrix based on the feature vectors of all speech samples in the speech sample sub-data of each language;

[0081] Inputting the language feature matrix of each language into the corresponding language model, training the corresponding language model based on the speech sample sub-data of each language, and the language model outputting the corresponding language feature vector;

[0082] Preprocessing the user voice data and the voice data to be translated in the voice data to be translated;

[0083] Performing feature extraction on the preprocessed user voice data to determine a user feature vector based on the user voice data, and at the same time, performing feature extraction on the preprocessed voice data to be translated to determine a feature vector to be translated based on the voice data to be translated;

[0084] Determine the similarity value between the user feature vector and the language feature vector of each language, and determine the language corresponding to the language feature vector with the highest similarity value as the target language;

[0085] Determine a similar sequence in the second language based on the feature vector to be translated and the language feature vectors of all languages;

[0086] Determine a set of languages ​​to be translated based on similar sequences in the second language;

[0087] The speech data to be translated is determined based on the language set to be translated and the preprocessed speech data to be translated.

[0088] In this embodiment, speech sample data in multiple languages ​​are collected, including sub-data in different languages, each language sub-data contains multiple speech samples, and these speech samples are preprocessed, such as denoising, normalization, etc., to ensure data quality.

[0089] In this embodiment, feature extraction is performed on each speech sample in the preprocessed speech sample data to extract a feature vector of each sample. Based on the feature vectors of all speech samples, a language feature matrix of each language is generated to represent the typical features of the language.

[0090] In this embodiment, the language feature matrix of each language is input into the corresponding language model, and the language model is trained based on these data, and the language model will output the language feature vector of the language.

[0091] In this embodiment, the user speech and the speech data to be translated are preprocessed, including noise elimination, segmentation, etc. Then, features are extracted from the two parts of data respectively to generate a user feature vector and a feature vector to be translated.

[0092] In this embodiment, the similarity between the user feature vector and the feature vectors of each language is calculated, and the language with the highest similarity is found as the target language.

[0093] The beneficial effects of the above technical solution are as follows: by determining the data to be translated and the target language based on the voice data, accurate voice recognition can be achieved, the target language can be intelligently determined, the accuracy and efficiency of translation can be improved, and the multilingual adaptability and intelligence level of the intelligent translation equipment can be improved.

[0094] Embodiment 4:

[0095] The embodiment of the present invention provides a multifunctional intelligent translation method, which determines a similar sequence in a second language based on a feature vector to be translated and language feature vectors of all languages, including:

[0096] Comparing the number of features of the feature vector to be translated and the number of features of the language feature vector of each language;

[0097] If the feature numbers of the feature vector to be translated and the feature numbers of the language feature vector are different, embedding the feature vector to be translated or the language feature vector with fewer features so that the feature numbers of the feature vector to be translated and the language feature vector are the same;

[0098] Calculating the similarity values ​​of the embedded feature vector to be translated and the language feature vector of each language, sorting the similarity values ​​of the embedded feature vector to be translated and the language feature vectors of all languages ​​from large to small, and determining a similar sequence of the second language;

[0099]

[0100]

[0101] Among them, S1 represents the similar sequence of the first language, a represents the feature vector to be translated after the speech data to be translated is embedded, and L 1 , L i , L N1 represents the language feature vector after the speech sample sub-data of the 1st language, the i-th language, and the N1-th language in the speech sample data are embedded, Represent the feature vector a to be translated and the language feature vector L of the first language respectively 1 , the language feature vector L of the i-th language i , the language feature vector L of the N1th language N1 The first similarity value, N1 represents the number of languages, a T represents the transposition of the feature vector to be translated after the voice data to be translated is embedded, (aL i ) represents the difference vector between the embedded feature vector to be translated and the language feature vector of the i-th language, (aL i ) T represents the transposition of the difference vector between the embedded feature vector to be translated and the language feature vector of the i-th language, ∑(aLi ) represents the covariance matrix of the embedded feature vector to be translated and the language feature vector of the i-th language, α1 represents the first similarity weight, α2 represents the second similarity weight, S2 represents the second language similarity sequence, Indicates the first similarity value in the sorted first language similarity sequence. represents the jth first similarity value in the sorted first language similarity sequence, Indicates the first similarity value of the N1th similarity sequence in the sorted first language.

[0102] In this embodiment, since speech data in different languages ​​may have different feature dimensions, the feature quantity of the feature vector to be translated is compared with the feature quantity of the feature vector of each language.

[0103] In this embodiment, if the number of features of the feature vector to be translated is different from the number of features of the language feature vector, the vector with fewer features is embedded, and the embedding process makes the number of features of the feature vector to be translated and the language feature vector consistent in some way (such as zero padding).

[0104] In this embodiment, once the dimensions of the feature vector to be translated and the language feature vector are consistent, the similarity between them is calculated, and the similarity values ​​are sorted from large to small to generate a second language similarity sequence, that is, a similarity ranking of all languages ​​with the data to be translated.

[0105] In this embodiment, Represents the covariance similarity value between the embedded feature vector to be translated and the language feature vector of the i-th language.

[0106] In this embodiment, Represents the cosine similarity between the embedded feature vector to be translated and the language feature vector of the i-th language.

[0107] In this embodiment, the first weight α1 represents the weight of the similarity value between the embedded feature vector to be translated and the language feature vector based on the cosine similarity value, and the second weight α2 represents the weight of the similarity value between the embedded feature vector to be translated and the language feature vector based on the covariance similarity value.

[0108] The beneficial effects of the above technical solution are as follows: based on the feature vector to be translated and the language feature vectors of all languages, a similar sequence of the second language is determined, which can provide a data basis for the set of languages ​​to be translated, improve the accuracy of language matching in the translation process, and provide more efficient and accurate translation results.

[0109] Embodiment 5:

[0110] The embodiment of the present invention provides a multifunctional intelligent translation method, which determines a set of languages ​​to be translated based on a similar sequence of a second language, including:

[0111] Determine the number of languages ​​in the language set to be translated based on similar sequences in the second language;

[0112]

[0113] Among them, Nu represents the number of languages ​​in the language set to be translated, CN1 j represents the first cliff value of the jth first similarity value in the second language sequence, CN2 j represents the second cliff value of the j-th first similarity value in the second language sequence, δ1 represents the preset first cliff threshold, δ2 represents the preset second cliff threshold, represents the j+1th first similarity value in the sorted first language similarity sequence, β1 represents the first adjustment factor, and β2 represents the second adjustment factor;

[0114] The language set to be translated is determined based on the second language similar sequence and the number of languages ​​in the language set to be translated.

[0115] In this embodiment, It represents the difference between the jth largest first similarity value and the j+1th largest first similarity value in the first language similarity sequence.

[0116] The first adjustment factor β1 adjusts the difference between two adjacent first similarity values ​​in the second language similarity sequence in the first cliff value.

[0117] The second adjustment factor β2 adjusts the j-th first similarity value and the average value of two first similarity values ​​adjacent to the j-th first similarity value in the second language similarity sequence in the first cliff value.

[0118] The beneficial effects of the above technical solution are as follows: determining the set of languages ​​to be translated based on similar sequences of the second language can provide a data basis for determining the speech data to be translated, thereby improving the accuracy and efficiency of translation, and improving the multilingual adaptability and intelligence level of intelligent translation equipment.

[0119] Embodiment 6:

[0120] The embodiment of the present invention provides a multifunctional intelligent translation method, in which a network server generates a translation speech result and a translation text result based on the data to be translated and the target language, including:

[0121] If the number of languages ​​in the language set to be translated in the data to be translated is greater than 1, performing data recognition on the preprocessed speech data to be translated in the data to be translated to determine speech sub-data to be translated of each language in the language set to be translated;

[0122] Inputting all languages ​​in the language set to be translated and the speech sub-data to be translated of each language into the translation model, and determining preliminary translation data based on the output result of the translation model, wherein the preliminary translation data includes preliminary translation text data of each language in the language set to be translated;

[0123] The preliminary translation data is corrected to determine the corrected translation data, and a translation speech result and a translation text result are generated based on the corrected translation data.

[0124] In this embodiment, if the data to be translated contains multiple languages ​​(that is, the number of languages ​​in the language set to be translated is greater than 1), data recognition is first performed on the speech data to be translated to identify the speech content of each language in the data to be translated, and the speech data of each language is divided into corresponding speech sub-data to be translated.

[0125] In this embodiment, all recognized languages ​​and their corresponding speech sub-data to be translated are input into the translation model for translation. The translation model generates a preliminary translation result based on the input speech data. The preliminary translation data includes the translation text data corresponding to each language.

[0126] The beneficial effects of the above technical solution are as follows: the network server generates translation speech results and translation text results based on the data to be translated and the target language, which can improve the accuracy of the translation results, ensure that the generated speech and text translation results meet user needs, enhance the flexibility and accuracy of the intelligent translation device, and improve the user experience.

[0127] Embodiment 7:

[0128] The embodiment of the present invention provides a multifunctional intelligent translation method, which performs data correction on preliminary translation data to determine corrected translation data, and generates a translation speech result and a translation text result based on the corrected translation data, including:

[0129] Based on the user feature vector of the user voice data, the preliminary translation text data of each language in the preliminary translation data is corrected to determine the corrected translation text data of each language in the set of languages ​​to be translated, wherein the correction processing includes grammar correction, context correction and word meaning correction;

[0130] Determine the translation text result based on the revised translation text data in all languages;

[0131] Converting the corrected translation text data of each language in the translation text result into translation voice data of the corresponding language;

[0132] The translated speech result is determined based on the translated speech data of all languages.

[0133] In this embodiment, a user feature vector is generated using the user's voice data, and the preliminary translation text of each language in the preliminary translation data is corrected based on the feature vector. These corrections include: grammar correction: correcting grammatical errors to make the sentence conform to the grammatical structure of the user's voice data; context correction: adjusting the translation result according to the context of the user's voice data to ensure that the translation is fluent and natural; word meaning correction: modifying the vocabulary selection according to the context of the user's voice data to avoid mistranslation or inaccurate word meaning.

[0134] In this embodiment, after the preliminary translation text data of each language is revised, revised translation text data of each language is generated, and the revised translation data is more accurate and conforms to language habits.

[0135] In this embodiment, the corrected translation text data of all languages ​​are aggregated to generate a final translation text result.

[0136] In this embodiment, the corrected translation text data of each language in the translation text result is converted into translation voice data of the corresponding language.

[0137] In this embodiment, the translated speech data of all languages ​​are merged to finally generate a complete translated speech result.

[0138] The beneficial effects of the above technical solution are: data correction is performed on the preliminary translation data to determine the corrected translation data, and the translation voice results and translation text results are generated based on the corrected translation data, which can improve the accuracy and fluency of the translation quality, make it more natural and in line with user needs, and improve user experience and translation accuracy.

[0139] Embodiment 8:

[0140] The embodiment of the present invention provides a multifunctional intelligent translation method, wherein an intelligent translation device plays a translation voice result and displays a translation text result, including:

[0141] The voice output sub-device of the intelligent translation device plays the translated voice result, and the display sub-device of the intelligent translation device displays the translated text result.

[0142] In this embodiment, the intelligent translation device is equipped with a voice output sub-device, such as a speaker, for playing the translated voice results. The device generates audio through speech synthesis technology from the translated voice data, and plays it through the output sub-device so that the user can hear the translated content.

[0143] In this embodiment, the intelligent translation device also includes a display sub-device, such as a screen or a display, for displaying the translated text results. The translated text will be presented to the user in a visual form, making it convenient for the user to view and understand the translated content.

[0144] The beneficial effects of the above technical solution are as follows: the intelligent translation device plays the translated voice results and displays the translated text results, which can enhance the convenience of user understanding and use, and improve the translation accuracy and user interaction experience.

[0145] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0146] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0147] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multifunctional intelligent translation method, characterized in that: include: 101: The intelligent translation device collects voice data in real time and determines the data to be translated and the target language based on the voice data; 102: The network server generates a translation voice result and a translation text result based on the data to be translated and the target language; 103: The intelligent translation device plays the translation voice result and displays the translation text result.

2. A multifunctional intelligent translation method according to claim 1, characterized in that: Intelligent translation equipment collects voice data in real time, including: The first voice collection sub-device in the intelligent translation device collects user voice data emitted by a user wearing the intelligent translation device; The second voice collection sub-device in the intelligent translation device collects the voice data to be translated in addition to the user voice data emitted by the user wearing the intelligent translation device; The voice data collected in real time by the intelligent translation device includes: user voice data collected by the first voice collection sub-device and voice data to be translated collected by the second voice collection sub-device.

3. A multifunctional intelligent translation method according to claim 2, characterized in that: Determine the data to be translated and the target language based on the voice data, including: Collecting speech sample data in multiple languages, wherein the speech sample data includes speech sample sub-data in each language, and the speech sample sub-data includes multiple speech samples; Preprocessing the speech sample data, extracting features from each speech sample in the preprocessed speech sample sub-data of each language, and determining a feature vector of each speech sample in the speech sample sub-data of each language; Determine a language feature matrix based on the feature vectors of all speech samples in the speech sample sub-data of each language; Inputting the language feature matrix of each language into the corresponding language model, training the corresponding language model based on the speech sample sub-data of each language, and the language model outputting the corresponding language feature vector; Preprocessing the user voice data and the voice data to be translated in the voice data to be translated; Performing feature extraction on the preprocessed user voice data to determine a user feature vector based on the user voice data, and at the same time, performing feature extraction on the preprocessed voice data to be translated to determine a feature vector to be translated based on the voice data to be translated; Determine the similarity value between the user feature vector and the language feature vector of each language, and determine the language corresponding to the language feature vector with the highest similarity value as the target language; Determine a similar sequence in the second language based on the feature vector to be translated and the language feature vectors of all languages; Determine a set of languages ​​to be translated based on similar sequences in the second language; The speech data to be translated is determined based on the language set to be translated and the preprocessed speech data to be translated.

4. A multifunctional intelligent translation method according to claim 3, characterized in that: Based on the feature vector to be translated and the language feature vectors of all languages, a similar sequence in the second language is determined, including: Comparing the number of features of the feature vector to be translated and the number of features of the language feature vector of each language; If the feature numbers of the feature vector to be translated and the feature numbers of the language feature vector are different, embedding the feature vector to be translated or the language feature vector with fewer features so that the feature numbers of the feature vector to be translated and the language feature vector are the same; Calculating the similarity values ​​of the embedded feature vector to be translated and the language feature vector of each language, sorting the similarity values ​​of the embedded feature vector to be translated and the language feature vectors of all languages ​​from large to small, and determining a similar sequence of the second language; Among them, S1 represents the similar sequence of the first language, a represents the feature vector to be translated after the voice data to be translated is embedded, L1, L i , L N1 represents the language feature vector after the speech sample sub-data of the 1st language, the i-th language, and the N1-th language in the speech sample data are embedded, They represent the feature vector a to be translated, the language feature vector L1 of the first language, and the language feature vector L of the i-th language respectively. i , the language feature vector L of the N1th language N1 The first similarity value, N1 represents the number of languages, a T represents the transposition of the feature vector to be translated after the voice data to be translated is embedded, (aL i ) represents the difference vector between the embedded feature vector to be translated and the language feature vector of the i-th language, (aL i ) T represents the transposition of the difference vector between the embedded feature vector to be translated and the language feature vector of the i-th language, ∑(aL i ) represents the covariance matrix of the embedded feature vector to be translated and the language feature vector of the i-th language, α1 represents the first similarity weight, α2 represents the second similarity weight, S2 represents the second language similarity sequence, Indicates the first similarity value in the sorted first language similarity sequence. represents the jth first similarity value in the sorted first language similarity sequence, Indicates the first similarity value of the N1th similarity sequence in the sorted first language.

5. A multifunctional intelligent translation method according to claim 4, characterized in that: Determine the language set to be translated based on the similar sequence of the second language, including: Determine the number of languages ​​in the language set to be translated based on similar sequences in the second language; CN1 j ≥δ1 or CN2 j ≥δ2); Among them, Nu represents the number of languages ​​in the language set to be translated, CN1 j represents the first cliff value of the jth first similarity value in the second language sequence, CN2 j represents the second cliff value of the j-th first similarity value in the second language sequence, δ1 represents the preset first cliff threshold, δ2 represents the preset second cliff threshold, represents the j+1th first similarity value in the sorted first language similarity sequence, β1 represents the first adjustment factor, and β2 represents the second adjustment factor; The language set to be translated is determined based on the second language similar sequence and the number of languages ​​in the language set to be translated.

6. A multifunctional intelligent translation method according to claim 3, characterized in that: The network server generates translated speech results and translated text results based on the data to be translated and the target language, including: If the number of languages ​​in the language set to be translated in the data to be translated is greater than 1, performing data recognition on the preprocessed speech data to be translated in the data to be translated to determine speech sub-data to be translated of each language in the language set to be translated; Inputting all languages ​​in the language set to be translated and the speech sub-data to be translated of each language into the translation model, and determining preliminary translation data based on the output result of the translation model, wherein the preliminary translation data includes preliminary translation text data of each language in the language set to be translated; The preliminary translation data is corrected to determine the corrected translation data, and a translation speech result and a translation text result are generated based on the corrected translation data.

7. A multifunctional intelligent translation method according to claim 6, characterized in that: The preliminary translation data is corrected to determine the corrected translation data, and a translation speech result and a translation text result are generated based on the corrected translation data, including: Based on the user feature vector of the user voice data, the preliminary translation text data of each language in the preliminary translation data is corrected to determine the corrected translation text data of each language in the set of languages ​​to be translated, wherein the correction processing includes grammar correction, context correction and word meaning correction; Determine the translation text result based on the revised translation text data in all languages; Converting the corrected translation text data of each language in the translation text result into translation voice data of the corresponding language; The translated speech result is determined based on the translated speech data of all languages.

8. A multifunctional intelligent translation method according to claim 1, characterized in that: The smart translation device plays the translated voice results and displays the translated text results, including: The voice output sub-device of the intelligent translation device plays the translated voice result, and the display sub-device of the intelligent translation device displays the translated text result.

Citation Information

Patent Citations

  • Pre-training method and device of intelligent translation model and storage medium

    CN111460838A

  • Translation method and device, equipment and storage medium

    CN116702801A

  • Neural network model-based Chinese language translation method and system

    CN118747500A

  • Simultaneous interpretation method and device, and storage medium

    WO2021077333A1