Speech translation method and apparatus
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-09
- Publication Date
- 2026-08-11
AI Technical Summary
[0003]随着全球化进程加速和跨语言交流需求日益增长,IM平台上的用户频繁面临语言障碍问题,尤其是在多语言群组或跨国对话中,语音消息虽提升了表达效率,却可能因接收方不理解源语言而造成沟通中断
[0013]本公开附加的方面和优点将在下面的描述中部分给出,部分将从下面的描述中变得明显,或通过本公开的实践了解到。
Smart Images

Figure CN122549451A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a speech translation method and apparatus. Background Technology
[0002] In modern instant messaging (IM) systems, voice messaging, as a core form of rich media communication, has become deeply integrated into users' daily communication scenarios. Compared to traditional text input, voice messaging, with its natural, efficient, and expressive characteristics, significantly lowers the communication barrier, making it particularly suitable for mobile environments, multitasking situations, or visually limited scenarios. Users can not only quickly convey complex information through voice, but also enhance the vividness and realism of their expression through tone, rhythm, and emotional nuances, thereby increasing the immersion and intimacy of social interactions.
[0003] With the acceleration of globalization and the increasing demand for cross-language communication, users on IM platforms frequently face language barriers, especially in multilingual groups or cross-border conversations. While voice messages improve expression efficiency, they can lead to communication interruptions due to the recipient's lack of understanding of the source language. Therefore, while retaining the original interactive advantages of voice messages, how to translate voice messages has become a key aspect of improving the inclusivity, usability, and user experience of instant messaging. Therefore, how to achieve voice message translation is of paramount importance. Summary of the Invention
[0004] This disclosure aims to at least partially address one of the technical problems in the related art.
[0005] To address this, this disclosure proposes a voice translation method and apparatus. Users only need to perform a translation operation on the voice message to simultaneously obtain the target language translation text and the translated voice with the same timbre in the conversation interface. This preserves the original speaker's emotions and identity characteristics while lowering the threshold for cross-language understanding. By caching the translation results based on voice message identifiers, the system avoids redundant processing, improves response speed, and saves computing resources. At the same time, presenting the translation text and translated audio as independent speech bubbles maintains the integrity of the original message and provides a clear and intuitive multimodal translation display, significantly enhancing the naturalness and efficiency of multilingual communication.
[0006] This disclosure provides a speech translation method in a first aspect, comprising: sending a first speech translation request to a server in response to a translation operation triggered by source speech in a first voice message bubble in a conversation interface; wherein the first speech translation request includes a voice message identifier and a first target language; receiving a first translated text and a first translated audio resource sent by the server in response to the first speech translation request, so as to play the first translated audio indicated by the first translated audio resource in a second voice message bubble in the conversation interface, and simultaneously displaying the first translated text; the first speech translation request is used by the server to translate and generate the first translated text according to the first target language when the server determines, based on the voice message identifier, that there is no translated text and translated audio corresponding to the source speech in the first target language; the first translated audio resource includes the first translated audio and / or the access address of the first translated audio, wherein the first translated audio is synthesized speech with the same timbre as the source speech, generated and stored based on the voiceprint features of the first translated text and the source speech.
[0007] A second aspect of this disclosure provides a voice translation method, comprising: in response to receiving a first voice translation request sent by a client, obtaining a voice message identifier and a first target language in the first voice translation request; wherein the first voice translation request is generated and sent by the client in response to a translation operation triggered by source voice in a first voice message bubble in a session interface; in response to determining, based on the voice message identifier, that there is no corresponding translated text and translated audio for the source voice in the first target language, translating the speech recognition text of the source voice according to the first target language to generate and store a first translated text; generating and storing a first translated audio with the same timbre as the source voice based on the first translated text and the voiceprint features of the source voice; sending the first translated text and the first translated audio resource to the client, so as to play the first translated audio indicated by the first translated audio resource in a second voice message bubble in the session interface, and simultaneously displaying the first translated text; wherein the first translated audio resource includes the first translated audio and / or the access address of the first translated audio.
[0008] A third aspect of this disclosure provides a voice translation device, comprising: a first sending module, configured to send a first voice translation request to a server in response to a translation operation triggered by source speech in a first voice message bubble in a conversation interface; wherein the first voice translation request includes a voice message identifier and a first target language; a first receiving module, configured to receive a first translated text and a first translated audio resource sent by the server in response to the first voice translation request, so as to play the first translated audio indicated by the first translated audio resource in a second voice message bubble in the conversation interface, and simultaneously display the first translated text; the first voice translation request is used by the server to translate and generate the first translated text according to the first target language when the server determines, based on the voice message identifier, that there is no translated text and translated audio corresponding to the source speech in the first target language; the first translated audio resource includes the first translated audio and / or the access address of the first translated audio, wherein the first translated audio is synthesized speech with the same timbre as the source speech, generated and stored based on the voiceprint features of the first translated text and the source speech.
[0009] A fourth aspect of this disclosure provides a voice translation device, comprising: a first acquisition module, configured to, in response to receiving a first voice translation request sent by a client, acquire a voice message identifier and a first target language in the first voice translation request; wherein the first voice translation request is generated and sent by the client in response to a translation operation triggered by source voice in a first voice message bubble in a conversation interface; a first translation module, configured to, in response to determining, based on the voice message identifier, that there is no corresponding translated text and translated audio for the source voice in the first target language, translate the speech recognition text of the source voice to generate and store a first translated text according to the first target language; a second translation module, configured to, based on the first translated text and the voiceprint features of the source voice, generate and store a first translated audio with the same timbre as the source voice; and a first sending module, configured to send the first translated text and the first translated audio resource to the client, so as to play the first translated audio indicated by the first translated audio resource in a second voice message bubble in the conversation interface, and simultaneously display the first translated text; wherein the first translated audio resource includes the first translated audio and / or the access address of the first translated audio.
[0010] A fifth aspect of this disclosure provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the speech translation method described in the first aspect of this disclosure, or implements the speech translation method described in the second aspect of this disclosure.
[0011] A sixth aspect of this disclosure provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the speech translation method as described in the first aspect of this disclosure, or implements the speech translation method as described in the second aspect of this disclosure.
[0012] A seventh aspect of this disclosure provides a computer program product that, when executed by an instruction processor, implements the speech translation method as described in the first aspect of this disclosure, or implements the speech translation method as described in the second aspect of this disclosure.
[0013] Additional aspects and advantages of this disclosure will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this disclosure. Attached Figure Description
[0014] The above and / or additional aspects and advantages of this disclosure will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, in which:
[0015] Figure 1 This is a schematic flowchart of a speech translation method provided in an embodiment of the present disclosure; Figure 2 This is a schematic flowchart of another speech translation method provided in an embodiment of this disclosure; Figure 3 This is a schematic diagram illustrating the effect of voice message bubbles provided in the embodiments of this disclosure; Figure 4 This is a schematic flowchart of another speech translation method provided in an embodiment of this disclosure; Figure 5 This is a schematic flowchart of another speech translation method provided in an embodiment of this disclosure; Figure 6 A schematic diagram illustrating the principle of a speech translation method provided in an embodiment of this disclosure; Figure 7 This is a schematic diagram of a process for obtaining translation results provided in an embodiment of this disclosure; Figure 8 This is a schematic diagram of the structure of a speech translation device provided in an embodiment of the present disclosure; Figure 9 This is a schematic diagram of another voice translation device provided in an embodiment of the present disclosure; Figure 10 This is a block diagram illustrating an electronic device for speech translation according to an exemplary embodiment. Detailed Implementation
[0016] Embodiments of this disclosure are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this disclosure, and should not be construed as limiting this disclosure.
[0017] It should be noted that the collection, gathering, updating, analysis, processing, use, transmission, and storage of user personal information involved in this disclosed technical solution all comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. Necessary measures are taken to prevent unauthorized access to user personal information data and to safeguard user personal information security, network security, and national security.
[0018] Existing instant messaging voice translation technologies generally suffer from problems such as lengthy processing links, fragmented interactive experiences, and response delays. For example, after a user initiates a translation, the system typically needs to complete multiple sequential steps, including voice upload, speech recognition (ASR), text translation (MT), and result feedback. This results in a complex overall processing flow, making it difficult to meet the immediacy requirements of real-time dialogue. At the same time, the translation result is only text and cannot preserve the emotion, tone, and interactive experience of the voice. For longer voice messages, the translation process takes significantly longer, and users often have to wait a long time to obtain the complete translation result, which seriously affects the fluency of use and communication efficiency.
[0019] To address any of the above-mentioned problems, this disclosure proposes a speech translation method and apparatus.
[0020] The speech translation method and apparatus of this disclosure are described below with reference to the accompanying drawings.
[0021] Figure 1 This is a schematic flowchart of a speech translation method provided in an embodiment of the present disclosure.
[0022] It should be noted that this embodiment of the disclosure uses the example of a voice translation method configured in a voice translation device. This voice translation device can be applied to any electronic device to enable the electronic device to perform voice translation functions. The electronic device can be any device with computing capabilities, such as a personal computer (PC), mobile terminal, server, etc. Mobile terminals can be, for example, in-vehicle devices, mobile phones, tablets, personal digital assistants, wearable devices, and other hardware devices with various operating systems, touchscreens, and / or displays.
[0023] For example, in applications such as instant messaging, the voice translation device can be used as part of a client application and run on an electronic device.
[0024] like Figure 1 As shown, the speech translation method may include the following steps: Step 101: In response to the translation operation triggered by the source voice in the first voice message bubble in the conversation interface, send a first voice translation request to the server.
[0025] The first speech translation request includes a speech message identifier and a first target language. When the server determines, based on the speech message identifier, that there is no corresponding translation text and translation audio for the source speech in the first target language, it translates the speech recognition text of the source speech to generate and stores the first translation text.
[0026] To perform voice translation according to user needs, one possible implementation is as follows: when a user clicks on a voice message (i.e., the source voice in the first voice message bubble) in the instant messaging interface and selects translation, the client generates a first voice translation request and sends it to the server. The first voice translation request includes a voice message identifier and a first target language. The voice message identifier uniquely identifies the voice message, and the first target language is the target language the user wishes to translate into (e.g., Chinese, English, etc.). Upon receiving the request, the server first checks whether it has stored the translation result of the voice message in the first target language. If not, it first obtains the speech recognition text (ASR result) of the source voice, then translates the speech recognition text into a first translated text in the first target language, and stores the first translated text for later reuse.
[0027] To effectively avoid repeated translation of the same speech content in the same target language, as another possible implementation, when the server receives the first speech translation request, it queries the existing translated text and translated audio corresponding to the source speech in the specified first target language based on the speech message identifier carried in the first speech translation request. If the server finds that the existing translated text is already in the first target language, it directly reuses the stored translation results, uses the existing translated text as the first translated text of this request, and returns the existing translated audio as the first translated audio to the client.
[0028] Step 102: Receive the first translated text and the first translated audio resource sent by the server in response to the first voice translation request, so as to play the first translated audio resource indicated by the first translated audio resource in the second voice message bubble in the conversation interface, and simultaneously display the first translated text.
[0029] The first translated audio resource includes the first translated audio and / or the access address of the first translated audio. The first translated audio is a synthesized speech that is generated and stored based on the voiceprint features of the first translated text and the source speech and has the same timbre as the source speech.
[0030] To lower the barrier to cross-language understanding and provide users with a more natural, authentic, and efficient communication experience in multilingual exchanges, one possible approach is for the client to create a second voice message bubble in the conversation interface after receiving the first translated text and the first translated audio resource from the server. This bubble synchronously displays the first translated text in the first target language and supports playing the first translated audio indicated by the first translated audio resource. The first translated audio resource can be a complete first translated audio file or a web address for the first translated audio file, allowing for on-demand loading. To preserve the speaker's personality and emotion, the first translated audio is a personalized synthesized speech with the same timbre as the source speech, generated by combining the first translated text with the original speaker's voiceprint features.
[0031] In addition, it should be noted that when the conversation interface is a group conversation interface, the translation result generated for the same voice message can be shared by multiple members in the group. That is, after any member triggers the translation operation, the server will broadcast the generated first translated text and the first translated audio resource to other members in the group, and present the corresponding second voice message bubble in each member's conversation interface, thereby avoiding duplicate translation requests and redundant processing.
[0032] In summary, users only need to trigger a translation operation on the source speech in the first voice message bubble in the conversation interface. The client can then send a first voice translation request containing a voice message identifier and the first target language to the server. The server accurately locates the corresponding source speech based on the voice message identifier. When it detects that the translated text and audio of the source speech in the first target language have not yet been cached, the server generates and persistently stores the first translated text using existing speech recognition text. Furthermore, the server combines the first translated text with the voiceprint features of the source speech to synthesize a personalized translated audio with a timbre highly consistent with the original speaker. This achieves accurate semantic conversion while effectively preserving the speaker's emotional expression and identity recognition. Subsequently, the server returns the first translated text and the first translated audio resource (which can be the audio data itself or its access address) to the client. The client then displays the translated text synchronously in the conversation interface as a second voice message bubble and supports playing the translated audio. This not only avoids interference with the original voice message but also significantly reduces the threshold for cross-language understanding through multimodal presentation, improving the naturalness of communication and overall interaction efficiency.
[0033] To clearly illustrate how the above embodiments receive the first translated text and the first translated audio resource sent by the server in response to the first voice translation request, so as to play the first translated audio indicated by the first translated audio resource in the second voice message bubble in the conversation interface and simultaneously display the first translated text, this disclosure proposes another voice translation method.
[0034] Figure 2 This is a schematic flowchart illustrating another speech translation method provided in an embodiment of this disclosure.
[0035] like Figure 2 As shown, the speech translation method may include the following steps: Step 201: In response to the translation operation triggered by the source voice in the first voice message bubble in the conversation interface, send a first voice translation request to the server.
[0036] The first speech translation request includes a speech message identifier and a first target language. When the server determines, based on the speech message identifier, that there is no corresponding translation text and translation audio for the source speech in the first target language, it translates the speech recognition text of the source speech to generate and stores the first translation text.
[0037] Step 202: Receive the first translated text and the first translated audio resource sent by the server, and load the first translated audio and the first translated text indicated by the first translated audio resource into the second voice message bubble in the conversation interface.
[0038] The first translated audio resource includes the first translated audio and / or the access address of the first translated audio. The first translated audio is a synthesized speech that is generated and stored based on the voiceprint features of the first translated text and the source speech and has the same timbre as the source speech.
[0039] To improve the voice translation experience, as one possible implementation, upon receiving the first translated text and the first translated audio resource from the server, the client loads the first translated audio indicated by the first translated text and the first translated audio resource into a second voice message bubble in the conversation interface, such as... Figure 3 As shown, the second speech bubble can simultaneously carry the source speech, the first translated text in the first target language corresponding to the source speech, and the first translated audio.
[0040] To reduce user waiting time and improve the immediacy of translation, one possible approach is for the client to enter recording mode after the user triggers the voice recording control in the session interface. During recording, the generated audio segments are streamed to the server in real time. The server performs streaming speech recognition based on these audio segments, pre-generating partial or complete speech recognition text. The complete speech data corresponding to the audio segments is used to extract the voiceprint features of the source speech. Simultaneously, upon completion of recording, the client uploads the complete speech file to the server to generate high-precision final speech recognition text and extract stable and reliable voiceprint features. It should be noted that the system prioritizes the combination of speech recognition text and voiceprint features that is completed first and meets availability requirements: if the speech recognition text generated streaming from the audio segments is complete and meets the confidence level, and the voiceprint features extracted from the corresponding complete speech data also reach the reliability threshold, then this combination can be directly used for subsequent translation; otherwise, the system will wait for the high-precision speech recognition text and stable voiceprint features generated after the complete speech file is processed before executing the translation process.
[0041] Step 203: In response to the playback operation triggered by the first audio in the second voice message bubble, play the first translated audio and simultaneously display the first translated text.
[0042] The first translated text is displayed in the second voice message bubble in a manner that includes automatic line wrapping and / or scrolling.
[0043] To achieve interactive playback and highly readable display of translated content, one possible approach is to have the system start playing the translated audio when the user clicks on the translated audio area in the second voice message bubble, while simultaneously ensuring that the corresponding first translated text is displayed synchronously. It should be noted that for longer texts, the system employs adaptation strategies such as automatic line wrapping or scrolling to accommodate different screen sizes and bubble layouts.
[0044] To achieve efficient reuse of translated text, one possible approach is to write the first translated text to the clipboard in response to a copy operation triggered by the first translated text, so that the first translated text in the clipboard can be pasted into a third-party session or third-party program.
[0045] In other words, when the system detects that a user has copied the first translated text, the text is written to the device's system clipboard. As a result, the user can easily paste the translated content into third-party conversations (such as other chat windows) or third-party applications (such as email, document editors, etc.).
[0046] To maintain visual consistency and interactive uniformity in the conversation interface, as a possible implementation, the display style of the source audio in the first voice message bubble (such as bubble shape, color scheme, playback control layout, waveform style, timestamp position, etc.) is consistent with the display style of the first translated audio in the second voice message bubble.
[0047] In summary, by receiving the first translated text and the first translated audio resource returned by the server, and loading the first translated audio indicated by the first translated text and the first translated audio resource into the second voice message bubble in the conversation interface, multimodal presentation of the translated content is achieved. Furthermore, when the user triggers the playback operation of the first translated audio in the second voice message bubble, the system plays the first translated audio and simultaneously displays the first translated text, improving the accuracy of understanding the translated content. At the same time, for translated content of different lengths, the bubble adopts adaptive layout strategies such as automatic line wrapping or scrolling display, effectively solving the readability problem of long texts under limited display space.
[0048] To clearly illustrate how the translation speech switching is performed in the above embodiments, this disclosure proposes another speech translation method.
[0049] Figure 4 This is a schematic flowchart illustrating another speech translation method provided in an embodiment of this disclosure.
[0050] like Figure 4 As shown, based on any of the above embodiments, the speech translation method further includes the following steps: Step 401: In response to the translation language switching operation of the second speech bubble, display the language selection menu.
[0051] To enable users to freely switch between multiple languages for translated voice messages, one possible approach is to display a language selection menu in the conversation interface when the user performs a translation language switching operation on the second voice message bubble, providing the user with an entry point to select the target language.
[0052] Step 402: In response to the selection of the second target language in the language selection menu, a second voice translation request including a voice message identifier and the second target language is sent to the server, and a transition animation is displayed in the session interface.
[0053] The second speech translation request is used by the server to translate and store the speech recognition text of the source speech in the second target language when the server determines, based on the speech message identifier, that there is no corresponding translation text and translation audio for the source speech in the second target language.
[0054] To achieve on-demand multilingual translation and efficient resource utilization, after the user selects a second target language from the language selection menu, the client sends a second speech translation request to the server, containing the original speech message identifier and the second target language. A transition animation is simultaneously displayed in the session interface to alleviate the perceived waiting time. Specifically, this second speech translation request is used by the server to query whether a translated text and audio file for the source speech already exist in the second target language. If not, the server uses the existing speech recognition text to translate it into the second target language, generating and persistently storing the corresponding second translated text and audio file. If the source speech in the second target language exists, the translated text is used as the second translated text, and the audio file is used as the second translated audio file.
[0055] Step 403: Receive the second translated text and the second translated audio resource sent by the server in response to the second voice translation request, stop displaying the transition animation in the session interface, play the second translated audio indicated by the second translated audio resource in the third voice message bubble in the session interface, and simultaneously display the second translated text.
[0056] The second translated audio resource includes the second translated audio and / or the access address of the second translated audio. The second translated audio is a synthesized speech that is generated and stored based on the voiceprint features of the second translated text and the source speech and has the same timbre as the source speech.
[0057] To achieve a smooth transition between multilingual translation results, one possible approach is for the client to terminate the transition animation after receiving the second translated text and the second translated audio resource from the server in response to the second speech translation request. Then, a third voice message bubble is generated in the chat interface to play the second translated audio indicated by the second translated audio resource, while simultaneously displaying the corresponding second translated text. The second translated audio resource can be either the second translated audio itself or its network address. The second translated audio is a personalized voice synthesized based on the voiceprint features of the second translated text and the source speech, with a timbre consistent with the original speaker's voice, thus maintaining the continuity of the speaker's identity and the realism of communication during language switching.
[0058] As another possible implementation, in order to improve the flexibility of user operation and the rational use of system resources, when the user triggers the close control in the language selection menu (such as the "Cancel" button or clicking on a blank area), the client only performs the operation of closing the menu and does not send a second voice translation request to the server. This ensures that the system will not initiate unnecessary translation processing if the user has not actually selected the target language or actively abandoned the intention to switch.
[0059] In addition, in response to the user's operation of switching the playback of the second voice message bubble, the current playback content is dynamically switched between the source voice and the first translated audio in the first target language, and the displayed text information is updated synchronously. For example, when the source voice is played, the speech recognition text corresponding to the source voice is displayed, and when the first translated audio is played, the first translated text corresponding to the first translated audio is displayed.
[0060] In summary, based on the user's triggering of the translation language switch operation by the second voice message bubble, a language selection menu is displayed, allowing the user to re-specify the target language as needed. Once the user selects the second target language, the client immediately sends a second voice translation request containing the original voice message identifier and the second target language to the server, and displays a transition animation in the conversation interface to improve the perceived smoothness of the waiting process. After receiving the request, the server checks whether the translated text and translated audio of the source voice in the second target language already exist based on the voice message identifier. If not, it reuses the cached source voice recognition text for translation, generates and persistently stores the second translated text, and combines the second translated text with the voiceprint features of the source voice to synthesize a second translated audio with the same timbre, so that the translated voice in the second target language still retains the original speaker's identity characteristics and emotional expression. The client then receives the second translated text and second translated audio resources returned by the server, stops the transition animation, and synchronously plays the personalized translated audio and displays the corresponding translation in an independent third voice message bubble in the conversation interface, effectively improving the efficiency of understanding multilingual information and the user experience.
[0061] To achieve the above embodiments, this disclosure proposes another speech translation method.
[0062] Figure 5 This is a schematic flowchart illustrating another speech translation method provided in an embodiment of this disclosure.
[0063] It should be noted that the embodiments of this disclosure can be implemented as a voice translation device; in application scenarios such as instant messaging, the device can be used as part of a server application and run on an electronic device.
[0064] like Figure 5 As shown, the speech translation method may include the following steps: Step 501: In response to receiving the first voice translation request sent by the client, obtain the voice message identifier and the first target language in the first voice translation request.
[0065] The first voice translation request is generated and sent by the client in response to the translation operation triggered by the source voice in the first voice message bubble in the conversation interface.
[0066] To achieve accurate voice translation, as one possible approach, after receiving the first voice translation request initiated by the client, the server extracts the voice message identifier and the first target language specified by the user from the first voice translation request. The first voice translation request is automatically generated and sent by the client when performing a translation operation on the source voice in the first voice message bubble in the user's conversation interface.
[0067] Step 502: In response to determining, based on the voice message identifier, that there is no corresponding translated text and translated audio for the source speech in the first target language, the first translated text is generated and stored by translating the speech recognition text of the source speech according to the first target language.
[0068] To avoid redundant translations and save computing resources, as a possible implementation, the server queries the cache or database based on the voice message identifier. If it finds that the source voice has not yet generated a translated text and translated audio in the first target language, it translates the existing speech recognition text (i.e. the original text transcribed from the source voice) into the first translated text in the first target language and persists the first translated text for later reuse.
[0069] Step 503: Based on the voiceprint features of the first translated text and the source speech, generate and store the first translated audio with the same timbre as the source speech.
[0070] To preserve the speaker's identity and enhance the realism of cross-language communication, as a possible approach, the server utilizes the generated first translated text, combined with voiceprint features extracted from the source speech or voiceprint features matched with the speaker obtained through querying, to generate a synthesized speech with a timbre highly consistent with the original speaker through personalized speech synthesis technology. This first translated audio is then stored.
[0071] Step 504: Send the first translated text and the first translated audio resource to the client so that the first translated audio resource indicates the first translated audio in the second voice message bubble in the conversation interface, and simultaneously display the first translated text.
[0072] The first translated audio resource includes the first translated audio and / or the access address of the first translated audio.
[0073] To improve the efficiency of voice translation, as one possible implementation, the server packages the generated first translated text and the first translated audio resource (first translated audio, or the network access address of the first translated audio, such as sending short voice messages directly and sending the network access address for long voice messages) and returns them to the client. Based on this, the client creates a second voice message bubble in the conversation interface to synchronously play the translated audio and display the corresponding translation, thus realizing multimodal presentation.
[0074] In summary, upon receiving the first voice translation request from the client, the server extracts the voice message identifier and the first target language. This request is sent by the user during the translation operation of the source voice within the first voice message bubble in the conversation interface, ensuring a precise binding between the translation intent and the specific voice content. Furthermore, the server queries the cache based on the voice message identifier. If it confirms that the source voice has not yet generated a translated text and audio in the first target language, it reuses the existing voice recognition text for translation, generating and persistently storing the first translated text, avoiding repeated recognition and redundant calculations. The server further combines the voiceprint features of the translated text and the source voice to synthesize and store a first translated audio with consistent timbre, ensuring that the translation result is not only semantically accurate but also preserves the original speaker's identity and emotional expression. Finally, the server returns the first translated text and the first translated audio resource (which can be audio data or its access address) to the client, who then plays the audio and displays the text synchronously in the conversation interface as a second voice message bubble. This enhances the realism and immersion of cross-language communication, allowing users to efficiently acquire auditory and visual information, lowering the comprehension threshold, and improving the user experience.
[0075] As one possible implementation, a speech synthesis model is used to synthesize a first translated audio with the same timbre as the source speech, based on voiceprint features and the first translated text; the first translated audio is then stored.
[0076] As one possible implementation, the system receives multiple audio data streams sent by a client and associated with the source speech; wherein the multiple audio data streams include at least: audio segments generated in real time during recording and a complete speech file generated after recording is completed; in response to successfully generating complete speech recognition text and voiceprint features corresponding to the source speech based on any audio data stream, the system obtains translation configuration information associated with the source speech; translates the speech recognition text according to the first target language in the translation configuration information to obtain the translated text corresponding to the source speech in the first target language; generates translated audio corresponding to the source speech in the first target language based on the translated text and voiceprint features corresponding to the source speech in the first target language; and stores the translated text and translated audio corresponding to the source speech in the first target language together.
[0077] As one possible implementation, in response to determining, based on the voice message identifier, that there exists a translated text and a translated audio corresponding to the source speech in the first target language, the translated text corresponding to the source speech in the first target language is taken as the first translated text; and the translated audio corresponding to the source speech in the first target language is taken as the second translated text.
[0078] As one possible implementation, in response to receiving a second voice translation request from a client, the system obtains the voice message identifier and the second target language from the second voice translation request; wherein the second voice translation request is generated and sent by the client in response to a translation language switching operation triggered by a second voice bubble; in response to determining, based on the voice message identifier, that there is no corresponding translation text and translation audio for the source speech in the second target language, the system translates the speech recognition text of the source speech according to the second target language, generates and stores the second translation text; based on the second translation text and the voiceprint features of the source speech, the system generates and stores the second translation audio with the same timbre as the source speech; the system sends the second translation text and the second translation audio resource to the client, so that the second translation audio indicated by the second translation audio resource is played in the third voice message bubble in the session interface, and the second translation text is displayed synchronously; wherein the second translation audio resource includes the second translation audio and / or the access address of the second translation audio.
[0079] To achieve personalized speech synthesis and cross-session voiceprint reuse, a mapping relationship can be established between voiceprint features and the user identifier of the voice message sender. This allows for quick lookup of the corresponding voiceprint features based on the user identifier of the message sender, thereby reusing the user's voice identity in translation, transcription, or other speech generation scenarios.
[0080] Based on any embodiment of this disclosure, such as Figure 6 As shown, the speech translation method of this disclosure embodiment may include the following steps: Step A: Sending voice messages (dual-link processing mechanism) 1. Audio data segmentation processing 1.1 During the recording process, the collected audio data is segmented according to time segments; 1.2 Each audio segment contains complete audio frame data and timestamp information. 1.3 The fragment size is dynamically adjusted based on network conditions and processing capacity. 2. Gradual upload mechanism The audio segments are uploaded to the translation service immediately after generation; Asynchronous upload is used to avoid blocking the recording process; Maintain the upload queue to ensure that fragments are transmitted in order; 3. Network Anomaly Handling Monitor network latency and upload failures; Establish a retransmission mechanism to handle fragments that fail to upload; Record the execution status of the optimization process; A2. Basic Link (Complete File Upload Mechanism) Technical highlights: Complete processing workflow after recording is completed 1. Audio file generation A complete audio file is generated after recording is complete; Upload the audio file to the file server to obtain the download address; The message sending process is triggered after the file upload is complete; 2. Message push distribution The messaging service receives a voice message and pushes it to two targets simultaneously. Pushed to the message recipient user; The request is sent to a translation service for processing. The translation service downloads audio data via the voice file address in the message; 3. Safety net mechanism The basic link serves as a backup plan for the priority link; When the optimized link experiences delays or failures, the basic link ensures that the translation task is executed normally; The first translator to complete the translation will be considered the first to complete the translation. Step B: Server Processing (Multi-Model Collaborative Processing) B1. Speech-to-Text Conversion and Feature Extraction Key technical points: ASR model invocation and voiceprint feature acquisition 1. Speech-to-text processing The translation service uses an Automatic Speech-to-Text (ASR) model. Convert the received audio data into text content. Output the text content and extracted speech feature information; 2. Voiceprint Feature Extraction Obtain the corresponding voiceprint features based on the user information of the message sender; Voiceprint features are used in the subsequent speech synthesis stage; Establish a mapping relationship between user voiceprint features and user ID; B2. Speech-to-text translation Key technical points: MT model calling and translation processing 1. Translation language settings The translation service retrieves the translation language settings based on the chat where the message is located. Determine the mapping relationship between the source language and the target language; Supports translation needs for multiple target languages; 2. Text translation processing The translation model (MT) is invoked to translate the text recognized from speech into the target language; Processing and formatting translation results; Ensure the accuracy and fluency of the translated content; B3. Translation Audio Synthesis Key technical points: TTS model invocation and audio synthesis 1. Speech synthesis processing The translation service invokes a text-to-speech (TTS) model. The translated text content is then combined with the extracted speech features; Generate translated audio that preserves the original speaker's vocal characteristics; 2. Audio file management Upload the synthesized translated audio to a file service; Download link for the generated translated audio; Establish the association between audio files and messages; B4. Translation Results Cache Key technical points: Caching mechanism based on message ID and target language 1. Caching key-value design Use message ID and target language as the composite key-value pair for caching; Ensure that the translation results of the same message in different languages are cached independently; Supports multiple users sharing the translation results of the same message; 2. Cache content management The cached content includes: The translated text result; Download link for the translated audio; Avoid translating the same message repeatedly; Improve the response efficiency of translation services; Step C: As Figure 7 As shown, the translation results are obtained (translation request processing). C1. Translation Triggering Method Key technical points: Multiple user interaction triggering mechanisms 1. Manual triggering method The user triggers a translation request by clicking on the message menu; Supports interaction methods such as long-pressing messages or right-clicking menus; Provide a clear entry point for translation operations; 2. Automatic triggering method Triggered based on user-defined automatic translation rules; Supports automatic translation for specific languages or specific contacts; To achieve an intelligent translation triggering mechanism; 3. Language switching trigger A new translation request is triggered when the user switches the target language; Supports switching to other target languages based on existing translations; Enables flexible switching between multiple languages; C2. Cache result delivery Key technical points: Fast response mechanism based on caching 1. Cache query processing The server queries the cache based on the message ID and target language; If the cache is hit, a new translation processing flow is triggered; 2. Result Distribution Mechanism The translated results retrieved will be pushed to the requesting user; Includes download links for the translated text and audio; Ensure the completeness and accuracy of the results; Step D: Results Display (Multimodal Display Interface) D1. Translation Text Display Key technical points: Display and interaction of text content 1. Text display area The translated text content is displayed below the translation speech area; Use clean, easy-to-read fonts and appropriate font sizes; Supports automatic line wrapping and scrolling display of long texts; 2. Text interaction function Supports text selection for translation. Provides a copy function, allowing users to copy translated text to the clipboard; Supports pasting into other applications or conversations; D2. Translation Audio Display Technical highlights: Audio playback controls and interactive features 1. Audio playback interface Use the same speech bubble style as the original speech; Displays a play button, progress bar, duration information, and playback speed control; Maintain consistency in interface style; 2. Playback control function Supports play and pause operations; Provides a progress bar dragging function, supporting fast forward and rewind; Displays the current playback time and total duration; Supports playback speed adjustment (such as 1x, 1.5x, 2x, etc.); 3. Audio download processing Automatically download translated audio files to your local device; Implement audio caching management to avoid duplicate downloads; A retry mechanism for handling audio download failures; D3. Target Language Switching Key technical points: Multilingual interface and processing logic 1. Language switch button The translation area provides a language indicator button (e.g., EN represents English). Click the button to bring up the language selection menu; Displays a list of currently supported target languages; 2. Multilingual cyclic translation Supports switching between multiple target languages; Each time a new translation request or cache query is triggered; Update the translated text and audio content; 3. Language Status Management Records the user's language switching history and preferences; Provides a quick way to switch between commonly used languages; Supports language switching via a latch and a redo operation; In order to achieve Figures 1 to 4 In an embodiment, this disclosure also proposes a speech translation device.
[0081] Figure 8 This is a schematic diagram of the structure of a speech translation device provided in an embodiment of this disclosure.
[0082] like Figure 8 As shown, the voice translation device 800 includes a first transmitting module 810 and a first receiving module 820.
[0083] The first sending module 810 is used to send a first voice translation request to the server in response to a translation operation triggered by the source voice in the first voice message bubble in the conversation interface. The first voice translation request includes a voice message identifier and a first target language. The first receiving module 820 is used to receive the first translated text and the first translated audio resource sent by the server in response to the first voice translation request, so as to play the first translated audio indicated by the first translated audio resource in the second voice message bubble in the conversation interface and simultaneously display the first translated text. The first voice translation request is used by the server to translate the speech recognition text of the source voice according to the first target language when it determines that there is no translated text and translated audio corresponding to the source voice in the first target language, and to generate and store the first translated text. The first translated audio resource includes the first translated audio and / or the access address of the first translated audio. The first translated audio is a synthesized voice with the same timbre as the source voice, generated and stored based on the voiceprint features of the first translated text and the source voice.
[0084] As one possible implementation, the first receiving module 820 is used to receive the first translated text and the first translated audio resource sent by the server, and load the first translated audio and the first translated text indicated by the first translated audio resource into the second voice message bubble in the conversation interface; in response to the playback operation triggered by the first audio in the second voice message bubble, the first translated audio is played and the first translated text is displayed synchronously; wherein, the display method of the first translated text in the second voice message bubble includes automatic line wrapping and / or scrolling display.
[0085] As one possible implementation, the voice translation device 800 also includes a copying module.
[0086] The copy module is used to write the first translated text to the clipboard in response to a copy operation triggered by the first translated text, so that the first translated text in the clipboard can be pasted into a third-party session or a third-party program.
[0087] As one possible implementation, the first speech translation request is also used by the server to determine, based on the speech message identifier, that there exists a translated text and a translated audio corresponding to the source speech in the first target language, and to use the translated text corresponding to the source speech in the first target language as the first translated text, and the translated audio corresponding to the source speech in the first target language as the first translated audio.
[0088] As one possible implementation, the voice translation device 800 also includes a first switching module.
[0089] The switching module is used to display a language selection menu in response to the language switching operation of the second speech bubble; in response to the selection operation of the second target language in the language selection menu, it sends a second speech translation request including a voice message identifier and the second target language to the server and displays a transition animation in the session interface; it receives the second translated text and the second translated audio resource sent by the server in response to the second speech translation request, stops displaying the transition animation in the session interface, plays the second translated audio indicated by the second translated audio resource in the third speech bubble in the session interface, and simultaneously displays the second translated text; wherein, the second speech translation request is used by the server to translate and store the speech recognition text of the source speech according to the second target language when the server determines, based on the voice message identifier, that there is no translated text and translated audio corresponding to the source speech in the second target language; the second translated audio resource includes the second translated audio and / or the access address of the second translated audio, the second translated audio is generated and stored based on the voiceprint features of the second translated text and the source speech, and is synthesized speech with the same timbre as the source speech.
[0090] As one possible implementation, the first switching module is also used to close the language selection menu in response to a trigger operation on the close control in the language selection menu, without sending a second voice translation request to the server.
[0091] As one possible implementation, the voice translation device 800 also includes a second switching module.
[0092] The second switching module is used to switch the currently playing audio between the source audio and the first translated audio in response to the playback switching operation of the second voice message bubble, and to display the text content corresponding to the currently playing audio simultaneously; wherein, when the currently playing audio is the source audio, the speech recognition text of the source audio is displayed, and when the currently playing audio is the first translated audio, the first translated text is displayed.
[0093] As one possible implementation, the voice translation device 800 also includes a recording module.
[0094] The recording module is used to respond to voice recording operations triggered by the voice recording control in the conversation interface. During the recording process, it sends the generated audio segments to the server in real time. The audio segments are used for streaming speech recognition and text recognition. The complete voice data corresponding to the audio segments is used to extract the voiceprint features of the source speech. When the recording is completed, the complete voice file is sent to the server. The complete voice file is used to generate the speech recognition text and voiceprint features corresponding to the source speech.
[0095] As one possible implementation, the display style of the source audio in the first voice message bubble is the same as the display style of the first translated audio in the second voice message bubble.
[0096] In the voice translation device of this embodiment, the user only needs to trigger the translation operation on the source voice in the first voice message bubble in the conversation interface. The client can then send a first voice translation request containing a voice message identifier and a first target language to the server. The server accurately locates the corresponding source voice based on the voice message identifier. When it detects that the translated text and translated audio of the source voice in the first target language have not yet been cached, the server generates and persistently stores the first translated text using existing speech recognition text. Furthermore, the server combines the first translated text with the voiceprint features of the source voice to synthesize a personalized translated audio with a timbre highly consistent with the original speaker. This achieves accurate semantic conversion while effectively preserving the speaker's emotional expression and identity recognition. Subsequently, the server returns the first translated text and the first translated audio resource (which may be the audio data itself or its access address) to the client. The client then displays the translated text synchronously in the conversation interface as a second voice message bubble and supports playing the translated audio. This not only avoids interference with the original voice message but also significantly reduces the threshold for cross-language understanding through multimodal presentation, improving the naturalness of communication and overall interaction efficiency.
[0097] In order to achieve Figures 5 to 7 In an embodiment, this disclosure also proposes a speech translation device.
[0098] Figure 9 This is a schematic diagram of another voice translation device provided in an embodiment of the present disclosure.
[0099] like Figure 9 As shown, the voice translation device 900 includes: a first acquisition module 910, a first translation module 920, a second translation module 930, and a first transmission module 940.
[0100] The system includes the following components: a first acquisition module 910, which, in response to receiving a first voice translation request from a client, acquires the voice message identifier and the first target language from the first voice translation request; wherein the first voice translation request is generated and sent by the client in response to a translation operation triggered by the source voice in the first voice message bubble in the session interface; a first translation module 920, which, in response to determining, based on the voice message identifier, that there is no corresponding translated text and translated audio for the source voice in the first target language, translates the voice recognition text of the source voice to generate and stores the first translated text according to the first target language; a second translation module 930, which, based on the first translated text and the voiceprint features of the source voice, generates and stores the first translated audio with the same timbre as the source voice; and a first sending module 940, which sends the first translated text and the first translated audio resource to the client to play the first translated audio indicated by the first translated audio resource in the second voice message bubble in the session interface, and simultaneously displays the first translated text; wherein the first translated audio resource includes the first translated audio and / or the access address of the first translated audio.
[0101] As one possible implementation, the second translation module 930 is used to synthesize speech based on voiceprint features and the first translated text using a speech synthesis model, synthesizing a first translated audio with the same timbre as the source speech; and storing the first translated audio.
[0102] As one possible implementation, the voice translation device 900 also includes a generation module.
[0103] The generation module is used to receive multiple audio data sent by the client and associated with the source speech; wherein the multiple audio data includes at least: audio segments generated in real time during recording and a complete speech file generated after recording is completed; in response to successfully generating the complete speech recognition text and voiceprint features corresponding to the source speech based on any audio data, it obtains translation configuration information associated with the source speech; according to the first target language in the translation configuration information, it translates the speech recognition text to obtain the translated text corresponding to the source speech in the first target language; according to the translated text and voiceprint features corresponding to the source speech in the first target language, it generates the translated audio corresponding to the source speech in the first target language; and it stores the translated text and translated audio corresponding to the source speech in the first target language together.
[0104] As one possible implementation, the voice translation device 900 also includes a determination module.
[0105] The determining module is used to respond to determining, based on the voice message identifier, the existence of a translated text and a translated audio corresponding to the source speech in the first target language, and to take the translated text corresponding to the source speech in the first target language as the first translated text; and to take the translated audio corresponding to the source speech in the first target language as the second translated text.
[0106] As one possible implementation, the voice translation device 900 also includes a third translation module.
[0107] The third translation module is used to respond to a second voice translation request sent by the client, and to obtain the voice message identifier and the second target language in the second voice translation request; wherein the second voice translation request is generated and sent by the client in response to a translation language switching operation triggered by the second voice bubble; in response to determining, based on the voice message identifier, that there is no corresponding translation text and translation audio for the source voice in the second target language, the module translates the speech recognition text of the source voice according to the second target language, generates and stores the second translation text; based on the voiceprint features of the second translation text and the source voice, the module generates and stores the second translation audio with the same timbre as the source voice; and sends the second translation text and the second translation audio resource to the client so that the second translation audio indicated by the second translation audio resource is played in the third voice message bubble in the session interface, and the second translation text is displayed synchronously; wherein the second translation audio resource includes the second translation audio and / or the access address of the second translation audio.
[0108] The voice translation device of this embodiment, upon receiving a first voice translation request from a client, extracts a voice message identifier and a first target language from it. This request is sent by the user during a translation operation of the source voice within the first voice message bubble in the conversation interface, ensuring accurate binding between the translation intent and the specific voice content. Furthermore, the server queries the cache based on the voice message identifier. If it confirms that the source voice has not yet generated a translated text and translated audio in the first target language, it reuses the existing voice recognition text for translation, generating and persistently storing the first translated text, avoiding repeated recognition and redundant calculations. The server further combines the voiceprint features of the translated text and the source voice to synthesize and store a first translated audio with consistent timbre, ensuring that the translation result is not only semantically accurate but also retains the original speaker's identity characteristics and emotional expression. Finally, the server returns the first translated text and the first translated audio resource (which may be audio data or its access address) to the client, who then synchronously plays the audio and displays the text in the conversation interface using a second voice message bubble. This enhances the realism and immersion of cross-language communication, allowing users to efficiently acquire auditory and visual information, lowering the comprehension threshold, and improving the user experience.
[0109] It should be noted that the foregoing explanation of the speech translation method embodiment also applies to the speech translation device of this embodiment, and will not be repeated here.
[0110] To achieve the above embodiments, this application also proposes an electronic device, such as... Figure 10 As shown, Figure 10This is a block diagram illustrating an electronic device for speech translation according to an exemplary embodiment.
[0111] like Figure 10 As shown, the aforementioned electronic device 1000 includes: The memory 1010 and the processor 1020 are connected by a bus 1030, which connects the different components (including the memory 1010 and the processor 1020). The memory 1010 stores a computer program, and when the processor 1020 executes the program, it implements the speech translation method of this disclosure embodiment.
[0112] Bus 1030 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. For example, these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.
[0113] Electronic device 1000 typically includes a variety of electronic device readable media. These media can be any available media that can be accessed by electronic device 1000, including volatile and non-volatile media, removable and non-removable media.
[0114] The memory 1010 may also include computer system readable media in the form of volatile memory, such as random access memory (RAM) 1040 and / or cache 1050. The electronic device 1000 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 1060 may be used to read and write non-removable, non-volatile magnetic media (…). Figure 10 Not shown; usually referred to as a "hard drive"). Although Figure 10 Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 1030 via one or more data media interfaces. Memory 1010 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of this disclosure.
[0115] A program / utility 1080 having a set (at least one) of program modules 1070 may be stored in, for example, memory 1010. Such program modules 1070 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 1070 typically perform the functions and / or methods described in the embodiments of this disclosure.
[0116] Electronic device 1000 can also communicate with one or more external devices 1090 (e.g., keyboard, pointing device, display, etc.), and with one or more devices that enable a user to interact with the electronic device 1000, and / or with any device that enables the electronic device 1000 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed through input / output (I / O) interface 1092. Furthermore, electronic device 1000 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) through network adapter 1093. Figure 10 As shown, network adapter 1093 communicates with other modules of electronic device 1000 via bus 1030. It should be understood that, although... Figure 10 As not shown, other hardware and / or software modules may be used in conjunction with the electronic device 1000, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0117] The processor 1020 performs various functional applications and data processing by running programs stored in the memory 1010.
[0118] It should be noted that the implementation process and technical principles of the electronic device in this embodiment are explained in the foregoing description of the speech translation method of this disclosure embodiment, and will not be repeated here.
[0119] To implement the above embodiments, this disclosure also proposes a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the speech translation method described in the above embodiments.
[0120] To implement the above embodiments, this disclosure also provides a computer program product that, when the instruction processor in the computer program product is executed, performs the speech translation method described in the above embodiments.
[0121] In the description of this specification, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0122] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0123] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.
Claims
1. A speech translation method, characterized in that, include: In response to a translation operation triggered by the source speech in the first voice message bubble in the conversation interface, a first voice translation request is sent to the server; wherein, the first voice translation request includes a voice message identifier and a first target language; The system receives the first translated text and the first translated audio resource sent by the server in response to the first voice translation request, and plays the first translated audio resource indicated by the first translated audio resource in the second voice message bubble in the conversation interface, and simultaneously displays the first translated text. The first voice translation request is used by the server to translate the speech recognition text of the source speech according to the first target language when the server determines that there is no translated text and translated audio corresponding to the source speech in the first target language based on the voice message identifier. The server then generates and stores the first translated text. The first translated audio resource includes the first translated audio and / or the access address of the first translated audio. The first translated audio is a synthesized speech that has the same timbre as the source speech, generated and stored based on the voiceprint features of the first translated text and the source speech.
2. The method according to claim 1, characterized in that, Receiving the first translated text and the first translated audio resource sent by the server in response to the first voice translation request, and playing the first translated audio resource indicated by the first translated audio resource in the second voice message bubble in the conversation interface, and simultaneously displaying the first translated text, includes: Receive the first translated text and the first translated audio resource sent by the server, and load the first translated audio and the first translated text indicated by the first translated audio resource into the second voice message bubble in the conversation interface; In response to a playback operation triggered by the first translated audio in the second voice message bubble, the first translated audio is played, and the first translated text is displayed simultaneously; wherein the first translated text is displayed in the second voice message bubble in a manner including automatic line wrapping and / or scrolling.
3. The method of claim 2, wherein, The method further includes: In response to a copy operation triggered by the first translated text, the first translated text is written to the clipboard so that the first translated text in the clipboard can be pasted into a third-party session or a third-party program.
4. The method of claim 1, wherein, The first voice translation request is further configured so that when the server determines, based on the voice message identifier, that there exists a translated text and a translated audio corresponding to the source voice in the first target language, the translated text corresponding to the source voice in the first target language is used as the first translated text, and the translated audio corresponding to the source voice in the first target language is used as the first translated audio.
5. The method of claim 1, wherein, The method further includes: In response to the language switching operation of the second speech bubble, a language selection menu is displayed; In response to the selection of a second target language in the language selection menu, a second voice translation request including the voice message identifier and the second target language is sent to the server, and a transition animation is displayed in the conversation interface; The system receives the second translated text and the second translated audio resource sent by the server in response to the second voice translation request, stops displaying the transition animation in the session interface, plays the second translated audio indicated by the second translated audio resource in the third voice message bubble in the session interface, and simultaneously displays the second translated text. The second voice translation request is used by the server to translate the voice recognition text of the source voice according to the second target language when the server determines that there is no translation text and translation audio corresponding to the source voice in the second target language. The translation text is then stored. The second translated audio resource includes the second translated audio and / or the access address of the second translated audio. The second translated audio is a synthesized speech that is generated and stored based on the voiceprint features of the second translated text and the source speech, and has the same timbre as the source speech.
6. The method of claim 5, wherein, The method further includes: In response to a trigger operation on the close control in the language selection menu, the language selection menu is closed, and the second voice translation request is not sent to the server.
7. The method of claim 1, wherein, The method further includes: In response to the playback switching operation of the second voice message bubble, the currently playing voice is switched between the source voice and the first translated audio, and the text content corresponding to the currently playing voice is displayed simultaneously; wherein, when the currently playing voice is the source voice, the speech recognition text of the source voice is displayed, and when the currently playing voice is the first translated audio, the first translated text is displayed.
8. The method of claim 1, wherein, The method further includes: In response to a voice recording operation triggered by the voice recording control in the session interface, the generated audio segments are sent to the server in real time during the recording process; the audio segments are used for streaming speech recognition and text recognition, and the voiceprint features of the source speech are extracted based on the complete voice data corresponding to the audio segments. Upon completion of recording, a complete audio file is sent to the server; wherein, the complete audio file is used to generate speech recognition text and voiceprint features corresponding to the source audio.
9. The method according to any one of claims 1-8, characterized in that, The display style of the source audio in the first voice message bubble is the same as the display style of the first translated audio in the second voice message bubble.
10. A speech translation method, characterized in that, include: In response to receiving a first voice translation request from a client, the system obtains the voice message identifier and the first target language from the first voice translation request; wherein the first voice translation request is generated and sent by the client in response to a translation operation triggered by the source voice in the first voice message bubble in the conversation interface; In response to determining, based on the voice message identifier, that there is no translated text and translated audio corresponding to the source speech in the first target language, the speech recognition text of the source speech is translated and a first translated text is generated and stored according to the first target language; Based on the voiceprint features of the first translated text and the source speech, a first translated audio with the same timbre as the source speech is generated and stored; The first translated text and the first translated audio resource are sent to the client so that the first translated audio resource indicated by the first translated audio resource is played in the second voice message bubble in the session interface, and the first translated text is displayed synchronously; wherein, the first translated audio resource includes the first translated audio and / or the access address of the first translated audio.
11. The method of claim 10, wherein, The step of generating and storing a first translated audio with the same timbre as the source speech based on the voiceprint features of the first translated text and the source speech includes: Using a speech synthesis model, speech synthesis is performed based on the voiceprint features and the first translated text to synthesize a first translated audio with the same timbre as the source speech. Store the first translated audio.
12. The method of claim 10, wherein, The translated text and translated audio corresponding to the source speech in the first target language are generated using the following steps: Receive multiple audio data sent by the client and associated with the source voice; wherein, the multiple audio data includes at least: audio segments generated in real time during recording and a complete voice file generated after recording is completed; In response to successfully generating complete speech recognition text and voiceprint features corresponding to the source speech based on any audio data, the translation configuration information associated with the source speech is obtained; According to the first target language in the translation configuration information, the speech recognition text is translated to obtain the translated text corresponding to the source speech in the first target language; Based on the translated text corresponding to the source speech in the first target language and the voiceprint features, generate the translated audio corresponding to the source speech in the first target language; The translated text and translated audio corresponding to the source speech in the first target language are stored together.
13. The method of claim 10, wherein, The method further includes: In response to determining, based on the voice message identifier, that there exists a translated text and a translated audio corresponding to the source speech in the first target language, the translated text corresponding to the source speech in the first target language is taken as the first translated text; The translated audio corresponding to the source speech in the first target language is used as the second translated text.
14. The method of claim 10, wherein, The method further includes: Upon receiving a second voice translation request from the client, the system obtains the voice message identifier and the second target language from the second voice translation request; wherein the second voice translation request is generated and sent by the client in response to a translation language switching operation triggered by the second voice bubble; In response to determining, based on the voice message identifier, that there is no translated text and translated audio corresponding to the source speech in the second target language, the speech recognition text of the source speech is translated according to the second target language, and a second translated text is generated and stored. Based on the voiceprint features of the second translated text and the source speech, a second translated audio with the same timbre as the source speech is generated and stored; The second translated text and the second translated audio resource are sent to the client so that the second translated audio indicated by the second translated audio resource is played in the third voice message bubble in the session interface, and the second translated text is displayed synchronously; wherein, the second translated audio resource includes the second translated audio and / or the access address of the second translated audio.
15. A speech translation device, characterized by include: The first sending module is configured to send a first voice translation request to the server in response to a translation operation triggered by the source voice in the first voice message bubble in the conversation interface; wherein the first voice translation request includes a voice message identifier and a first target language; The first receiving module is configured to receive the first translated text and the first translated audio resource sent by the server in response to the first voice translation request, so as to play the first translated audio indicated by the first translated audio resource in the second voice message bubble in the conversation interface, and simultaneously display the first translated text. The first voice translation request is used by the server to translate the speech recognition text of the source speech according to the first target language when the server determines that there is no translated text and translated audio corresponding to the source speech in the first target language based on the voice message identifier. The server then generates and stores the first translated text. The first translated audio resource includes the first translated audio and / or the access address of the first translated audio. The first translated audio is a synthesized speech that has the same timbre as the source speech, generated and stored based on the voiceprint features of the first translated text and the source speech.
16. A voice translation device, characterized in that, include: The first acquisition module is configured to, in response to receiving a first voice translation request sent by the client, acquire the voice message identifier and the first target language in the first voice translation request; wherein, the first voice translation request is generated and sent by the client in response to a translation operation triggered by the source voice in the first voice message bubble in the conversation interface; The first translation module is configured to, in response to determining, based on the voice message identifier, that there is no corresponding translation text and translation audio for the source speech in the first target language, translate the speech recognition text of the source speech to generate and store the first translation text according to the first target language; The second translation module is used to generate and store a first translated audio with the same timbre as the source speech based on the voiceprint features of the first translated text and the source speech; A first sending module is configured to send the first translated text and the first translated audio resource to the client, so that the first translated audio indicated by the first translated audio resource is played in the second voice message bubble in the session interface, and the first translated text is displayed synchronously; wherein, the first translated audio resource includes the first translated audio and / or the access address of the first translated audio.
17. An electronic device, comprising: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the speech translation method as described in any one of claims 1-9, or implements the speech translation method as described in any one of claims 10-14.
18. A computer readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed by the processor, it implements the speech translation method as described in any one of claims 1-9, or implements the speech translation method as described in any one of claims 10-14.