Speech translation method and device, readable storage medium and program product
By using an end-to-end speech translation method, the translation direction is dynamically determined, and translated speech and text with the same acoustic features are generated. This solves the problem of low information transmission efficiency in cross-language translation and achieves efficient and natural multimodal output.
Patent Information
- Application Number
- CN202511396597.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2026-01-09
AI Technical Summary
In existing technologies, cross-language translation tools can only perform one-way translation and cannot achieve two-way translation, resulting in low information transmission efficiency and large response delays.
By determining the source language of the speech to be processed in real time, the system generates translated speech and text with the same acoustic features. It adopts an end-to-end speech translation method, dynamically determines the translation direction, directly generates translated speech and displays text simultaneously, simplifies the operation process, and reduces the switching delay between modules.
It enables efficient information transmission in cross-language interaction scenarios, reduces response latency, improves the naturalness and recognizability of translated speech, and meets diverse needs for barrier-free communication.
Smart Images

Figure CN121303151A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a speech translation method, device, readable storage medium, and program product. Background Technology
[0002] In daily life and work, cross-language translation is frequently required. For example, this can occur when different users are communicating via voice; when a speaker at a meeting uses a language different from that of other participants; or when a teacher teaches in a language different from that of the students in an educational setting. Real-time voice translation is necessary in all these scenarios.
[0003] In related technologies, a source language and a target language need to be set before translation. Subsequent input speech is then translated one-way from the source language to the target language, making bidirectional translation impossible. This results in low information transmission efficiency. Summary of the Invention
[0004] This disclosure provides a voice translation method, device, readable storage medium, and program product to address the problem of low information transmission efficiency in information interaction scenarios involving different languages.
[0005] In a first aspect, embodiments of this disclosure provide a speech translation method, comprising: determining the source language corresponding to the speech to be processed; generating translated speech and translated text in the translation language; wherein the translated speech and the speech to be processed have the same acoustic features; the translation language is determined by the user or the usage environment; and displaying the translated speech and the translated text.
[0006] Secondly, embodiments of this disclosure provide a speech translation device, comprising: a determining unit for determining the source language corresponding to the speech to be processed; a generating unit for generating translated speech and translated text in the translation language; wherein the translated speech and the speech to be processed have the same acoustic features; the translation language is determined by a user or the usage environment; and a display unit for displaying the translated speech and the translated text.
[0007] Thirdly, embodiments of this disclosure provide an electronic device, including: a processor and a memory;
[0008] The memory stores computer-executed instructions;
[0009] The processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the speech translation method as described in the first aspect and various possible designs of the first aspect.
[0010] Fourthly, embodiments of this disclosure provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the speech translation method described in the first aspect and various possible designs of the first aspect.
[0011] Fifthly, embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the speech translation method as described in the first aspect and various possible designs of the first aspect.
[0012] The speech translation method, device, readable storage medium, and program product provided in this embodiment first determine the source language, then generate translated speech and translated text in the target language, and finally play and display them simultaneously. This achieves dynamic determination of the translation direction, using the dynamically determined translation direction to perform speech and text translation of the speech to be processed, improving information transmission efficiency in cross-language interaction scenarios. Furthermore, directly generating translated speech from the speech to be processed achieves end-to-end translation, eliminating the switching delays and manual intervention between multiple independent modules such as language recognition, speech recognition, and machine translation, simplifying the operation process, reducing the overhead of data transmission and processing between multiple independent modules, and lowering response latency. In addition, by simultaneously presenting translated speech and corresponding text, a multimodal output method using both visual and auditory information channels produces a synergistic effect, meeting diverse needs for barrier-free communication. Moreover, the translated speech retains the acoustic characteristics of the speech to be processed, improving the naturalness and recognizability of the translated speech. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 Flowchart of the speech translation method provided in the embodiments of this disclosure Figure 1 ;
[0015] Figure 2 Flowchart of the speech translation method provided in the embodiments of this disclosure Figure 2 ;
[0016] Figure 3A This is a diagram illustrating the input of voice to be processed in the dialogue interface.
[0017] Figure 3BThis is a diagram illustrating a multi-person dialogue scenario in a chat interface.
[0018] Figure 4 Flowchart 3 of the speech translation method provided in this embodiment of the disclosure;
[0019] Figure 5A This is a diagram illustrating the input of streaming voice and the display of translated voice and text in the call interface;
[0020] Figure 5B This is an illustration of multiple people inputting voice messages and displaying translated voice messages and text in a call interface.
[0021] Figure 6 This is an illustration of a translation record prompt message;
[0022] Figure 7 A structural block diagram of the speech translation device provided in the embodiments of this disclosure;
[0023] Figure 8 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0025] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0026] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0027] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0028] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0029] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0030] Some voice translation tools can only perform one-way voice translation, that is, they can only translate between one source language and one target language. In other words, they cannot automatically switch from the source language to the target language.
[0031] Some voice translation tools require first converting the source language speech to text, then translating the text into a target language, and finally synthesizing the speech from the target language. This method of voice translation has a relatively large response delay. Therefore, these methods result in low information transmission efficiency when using multilingual voice communication.
[0032] The solution provided in this disclosure dynamically determines the translation language based on the speech to be processed in real time, and translates the speech to be processed into the translation speech through a speech-to-speech translation method. This achieves automatic dynamic translation direction switching and reduces speech translation latency through an end-to-end speech translation method, thereby improving the information transmission efficiency during multilingual speech information interaction.
[0033] Please refer to Figure 1 , Figure 1 Flowchart of the speech translation method provided in this disclosure Figure 1 ,like Figure 1 As shown, the method includes the following steps:
[0034] S101: Determine the source language of the speech to be processed.
[0035] The speech translation method disclosed herein can be implemented by an application client running on a terminal device or an application server running on a server.
[0036] The aforementioned voice to be processed can be various types of input voice, such as voice from currently playing audio or video content, or voice input by the user.
[0037] Upon receiving the voice message to be processed, the aforementioned executing entity can determine the language used in the voice message and use that language as the source language. That is, the source language here is not a fixed language, but is determined based on the voice message to be processed.
[0038] The language corresponding to the speech to be processed can be determined based on phoneme inventory, pronunciation mode, rhythm, etc.
[0039] S102: Generate the translated speech and translated text in the target language; the translated speech has the same acoustic features as the speech to be processed; the target language is determined by the user or the usage environment.
[0040] The translation language can be determined based on the source language in various ways, such as using the historical translation language from historical translation records as the translation language.
[0041] In one implementation, the translation language is determined based on one of the following:
[0042] Determined based on the language configured by the user;
[0043] Determined based on the context of the speech to be processed;
[0044] Determined based on the location and environment;
[0045] It is determined based on the language environment used by both parties during the voice call.
[0046] In some application scenarios, the translation language can be determined based on the language correspondence configured by the user.
[0047] A language pair can include any two languages from multiple languages. These multiple languages include, but are not limited to, English, Chinese, German, and Japanese. In these application scenarios, the specific two languages to be translated can be configured before the execution entity performs the speech translation.
[0048] The language pair consists of two pre-configured languages; the source language is either one of the languages in the language pair. After determining the source language of the speech to be processed, the other language in the above language pair can be used as the translation language.
[0049] By configuring language pairs, the range of translation languages can be clearly defined. The translation language is determined within this range based on the source language of the speech to be processed, avoiding ambiguity and interference in a multilingual environment, reducing the computational complexity and decision-making time of the system, and improving translation efficiency.
[0050] In some applications, the translation language can be determined based on the context of the speech being processed. For example, if a user has previously used a certain language as the translation language, then that language can be used as the translation language for the speech being processed. Alternatively, if there is input speech in a language other than the language being processed in the preceding text, then that language can be used as the translation language.
[0051] In some application scenarios, the translation language can be determined based on the current location environment. This location environment can include geographical location. For example, if resident users in this location environment typically communicate in their first language, then the translation language for the speech to be processed will be their first language. Understandably, in this application scenario, obtaining the user's location environment requires their authorization.
[0052] In some applications, the translation language can be determined based on the language environment used by both parties during a voice call. During a voice call, both parties can interact using different languages. In this scenario, the translation language for one party's voice message can be set to the language used by the other party.
[0053] In these implementations, the translation language is determined in different ways, allowing the method to be applied to different scenarios and meet the needs of users in different situations.
[0054] In this step, translated speech can be generated based on the acoustic features of the original speech. The translated speech retains the acoustic features of the original speech, making it more natural and recognizable. Acoustic features include, but are not limited to, timbre, pitch, and prosody.
[0055] One approach is to use an end-to-end speech translation method to generate translated speech based on the target language.
[0056] In some examples, the waveform of the speech to be processed can be encoded using an encoder to obtain a speech encoding sequence. This speech encoding sequence captures information such as the content and acoustic features of the speech. The speech encoding sequence is then input into a decoder, which generates an abstract linguistic content representation sequence corresponding to the target language and an acoustic feature sequence inheriting the acoustic features of the speech to be processed. Subsequently, based on this abstract linguistic content representation sequence and acoustic feature sequence, a translated speech waveform with the acoustic features of the speech to be processed is synthesized. The translated speech waveform (digital signal) can be converted from digital to analog and amplified using a speaker, which then drives the speaker to reconstruct its vibrations as sound waves, thus playing the translated speech. This process achieves end-to-end speech translation, enabling rapid translation of speech from one language to another.
[0057] It is understandable that the encoder and decoder described above can be pre-trained using multiple training sample pairs. The training sample pairs include speech from two languages, where the waveform of one language's speech serves as the encoder's input, and the acoustic feature sequence of the other language, carrying abstract linguistic content, serves as the decoder's target output. Furthermore, the input and output languages can be interchanged. It is also understandable that multiple language pairs can be formed from multiple languages, and for each language pair, multiple training sample pairs corresponding to that language pair can be obtained to train the encoder and decoder, resulting in an encoder and decoder capable of end-to-end speech translation.
[0058] In one example, the translated text can be obtained through speech recognition. It's understandable that speech recognition can be used to convert the input speech into text at the beginning of the process of translating the input speech into translated speech.
[0059] In one example, the decoder described above can be a multi-task decoder, which may include a text decoder and a speech decoder. The text decoder can generate translated text based on the speech encoding sequence output by the encoder; the speech decoder can generate translated speech based on the acoustic feature sequence output by the encoder.
[0060] S103: Display the translated audio and translated text.
[0061] An audio player can be used to play the translated audio, and the translated text can be displayed on the application client. This presents the translation result both visually and aurally.
[0062] In this embodiment, the source language is first determined, then the translated speech and text in the target language are generated, and finally played and displayed simultaneously. This achieves dynamic determination of the translation direction, which is then used to translate the speech and text of the speech to be processed, improving the efficiency of information transmission in cross-language interaction scenarios. Furthermore, the direct generation of translated speech from the speech to be processed achieves end-to-end translation, eliminating the switching delays and manual intervention between multiple independent modules such as language identification, speech recognition, and machine translation. This simplifies the operation process, reduces the overhead of data transmission and processing between multiple independent modules, and lowers response latency. Moreover, by simultaneously presenting the translated speech and corresponding text, a multimodal output method using both visual and auditory information channels produces a synergistic effect, meeting diverse needs for barrier-free communication. Additionally, the translated speech retains the acoustic characteristics of the speech to be processed, improving the naturalness and recognizability of the translated speech.
[0063] In some embodiments, step S102 includes:
[0064] In response to the instruction to use the original speech for translation, a translated speech is generated based on the acoustic features of the speech to be processed and the target language.
[0065] In one example, a user can issue the command to use native audio through the aforementioned application client. For instance, the user can select native audio as the audio type for translation from multiple audio types within the application client.
[0066] Understandably, other voice types can also be set in the application client. Other voice types may include, for example, machine voice.
[0067] Using native speech commands indicates that translated speech is generated based on the acoustic features of the input speech. Acoustic features include, but are not limited to, one or more of timbre, pitch, and prosody. Since the acoustic features of the input speech are preserved, translated speech using native speech is naturally recognizable.
[0068] In these implementations, translated speech is generated based on the acoustic features of the speech to be processed, according to instructions to use the original speech, thereby enabling precise user control over the translated speech. Furthermore, the extraction of acoustic features from the speech to be processed through explicit instruction authorization enhances privacy protection and security control in acoustic feature processing.
[0069] In some implementations, the method further includes:
[0070] Acoustic features are extracted from at least a portion of the speech to be processed.
[0071] Acoustic features can be extracted using a portion of the speech to be processed, which can improve the speed of acquiring acoustic features and thus reduce the response delay of translated speech.
[0072] Acoustic features can also be extracted using the entire speech to be processed, which can improve the accuracy of the extracted acoustic features and help improve the consistency between the generated translated speech and the speech to be processed.
[0073] Specifically, the acoustic features of the speech to be processed can be obtained through preprocessing (including digitization, pre-emphasis, framing, windowing, etc.), time-frequency analysis, and acoustic feature calculation (including Mel frequency cepstral coefficients, fundamental frequency, energy and chromaticity features).
[0074] In some embodiments, step S102 includes:
[0075] First, the speech to be processed is sent to the speech model, which then extracts the acoustic features of the speech frame by frame.
[0076] Secondly, the speech model is used to generate the translated speech based on the acoustic features of the speech to be processed extracted frame by frame and the translation language.
[0077] In these implementations, the aforementioned speech model can be any model capable of achieving end-to-end speech translation functionality.
[0078] In one example, the above speech model could be a model that includes an encoder and a decoder.
[0079] In one example, the aforementioned speech model could be a model obtained by training a language model with natural language processing capabilities on speech translation.
[0080] In these embodiments, the above-mentioned speech model can extract acoustic features of the speech to be processed frame by frame, and generate translated speech of each frame of speech to be processed based on the acoustic features of each frame.
[0081] This approach utilizes a speech model to extract acoustic features frame-by-frame from the source speech and generate translated speech. Compared to traditional cascaded systems (speech recognition, machine translation, and speech synthesis), this method uses a single model to convert source and target speech, reducing error propagation and improving the stability and accuracy of the generated translated speech. Combining multiple steps into a single model simplifies the system architecture and reduces the complexity of deploying and maintaining the speech translation system. By using the same model to process acoustic features and generate translated speech, the acoustic features of the source speech (such as timbre, intonation, and speech rate) are better preserved, making the translated speech closer to the original speaker's voice. Furthermore, the frame-by-frame extraction of acoustic features by the model enables ultra-refined analysis and processing of the speech signal, making the translated speech closer to the source speech, better preserving its acoustic characteristics, and accurately conveying its emotional information.
[0082] Please refer to Figure 2 The figure shows a flowchart of the speech translation method provided in an embodiment of this disclosure. Figure 2 ,like Figure 2 As shown, the method includes the following steps:
[0083] S201: Receive voice input in the dialog interface.
[0084] S202: In response to the voice input end command, the voice input in the dialog interface will be treated as voice to be processed.
[0085] In this embodiment, the user can input voice in the dialogue interface. This dialogue interface can be the one displayed in the aforementioned application client, where the user can converse with the intelligent agent.
[0086] The aforementioned unprocessed voice input by the user in the dialogue interface can be the unprocessed voice input.
[0087] The aforementioned voice input termination command can be, for example, a command sent by the user triggering a send control, a voice command indicating the end of the user's voice input, a voice input termination command generated based on the voice input duration reaching a duration threshold, or a voice input termination command generated based on the input voice length reaching a length threshold.
[0088] S203: Determine the source language of the speech to be processed.
[0089] S204: Generate the translated speech and translated text in the target language; the translated speech has the same acoustic features as the speech to be processed; the target language is determined by the user or the usage environment.
[0090] S205: Display the translated audio and translated text in the dialog interface.
[0091] Please refer to Figure 3A , Figure 3A This is a diagram illustrating the input of voice to be processed in a dialog interface. The application client can display a dialog interface 31, which includes voice input controls, such as a voice input control 32 displayed as "Press and hold to speak". Figure 3A As shown in the left figure, the user can perform a trigger operation on the voice input control 32 to enter the voice input state, as shown in the left figure. Figure 3A As shown in the diagram. In voice input mode, the user can input a voice message as the voice to be processed. After the user finishes voice input, the voice message entered by the user in the dialogue interface is taken as the voice to be processed. The application client can send the voice to be processed to the agent, which determines the source language of the voice message and the translation language. Based on the acoustic features of the voice message, a translation voice message corresponding to the translation language is generated. The agent can send the translation voice message and the translation text to the application client, which can then display the translation text and play the translation voice message in the dialogue interface. The text corresponding to the voice message can also be displayed in the dialogue interface, such as... Figure 3A The phrase "Hello, I have booked accommodation for 3 days" shown in the right image can be the text corresponding to the speech to be processed. "Hello, I have made accommodation for three days." is the translated text. Furthermore, the translated speech can be played simultaneously while displaying the translated text.
[0092] In this embodiment, the user-input speech is used as the speech to be processed in the human-computer dialogue interface. The source language of the speech to be processed is dynamically determined, and a translated speech with the same acoustic features as the input speech is generated in the corresponding translation language. The translated speech and translated text are displayed in the dialogue interface. Since end-to-end processing from speech input to multilingual translation is achieved in a single dialogue interface, the user does not need to perform any language switching operations. The entire process of language recognition, translation generation, and result presentation can be completed automatically, eliminating technical barriers and operational obstacles to cross-language communication. In addition, the multimodal (speech and text) information fusion output in the dialogue interface can meet the receiving preferences of different users and improve the redundancy and reliability of information transmission. The comprehensibility and emotional fidelity of the speech translation results are enhanced: since the translated speech reproduces the acoustic characteristics of the speech to be processed, it not only conveys the semantic content but also retains the speaker's emotions, tone, and other paralinguistic information, helping users to more accurately understand the semantic intent of the speech to be processed.
[0093] In some implementations of this embodiment, the speech to be processed is speech in a multi-turn dialogue; the participants in the multi-turn dialogue include at least two speakers, wherein the at least two speakers use two languages to conduct the dialogue;
[0094] Prior to step S203 above, the method further includes the following steps:
[0095] First, determine the first speaker of the voice input in the current round. The first speaker can be any one of at least two speakers.
[0096] Secondly, in response to the fulfillment of the triggering condition, the voice input by the first speaker in the current round is taken as the voice to be processed.
[0097] In each round, one participant can input the speech to be processed. The aforementioned execution entity can use the intelligent agent to determine the source language of the speech to be processed in that round and generate the corresponding translated speech in the target language.
[0098] The aforementioned triggering conditions may include, for example, receiving a voice input end command. This voice input end command may include a user-triggered voice input end command, a voice input end command determined by the input duration, or a voice input end command determined by detecting a speaker switch.
[0099] In these implementations, the speaker is identified in the dialogue interface by identifying the speaker in the current turn, and then combined with the trigger condition judgment. This ensures that the speaker's voice is processed only when the user's true intention is clear. This can reduce false triggers in noisy environments and improve the intelligence and accuracy of voice translation.
[0100] In some implementations, the voice input by the first speaker is used to respond to the historical translated voice input corresponding to the second speaker; the language used in the historical translated voice input is the same as the language used in the voice input to be processed.
[0101] For example, if the second speaker in a previous round inputs English speech to be processed, a Chinese translation speech 1 can be generated. The first speaker hears the Chinese translation speech 1. In the current round, if the first speaker inputs Chinese speech to be processed 1, an English translation speech 2 can be generated, and the second speaker can hear the English translation speech 2. Speech to be processed 1 can be a response to translation speech 1. Furthermore, the second speaker can also start a new round of dialogue based on the translation speech 2 they hear.
[0102] Figure 3B This is a diagram illustrating a multi-person dialogue scenario in the dialogue interface. "I" and "Speaker 1" can input the voice to be processed in multiple rounds in the dialogue interface 31.
[0103] After the user inputs the voice message "Hello, how do I get to the nearest subway station?" in dialogue interface 31, the text of the voice message input by the user will be displayed. The source language is Chinese. Dialogue interface 31 will then play the translated voice message in the target language (English): "Excuse me, how can I get to the nearest subway station?". Simultaneously, the translated text "Excuse me, how can I get to the nearest subway station?" will be displayed.
[0104] Speaker 1 inputs the English voice message "Sure. Go straight for two blocks, then turn left. The station entrance is right there." Dialogue interface 31 displays the text of this voice message. Dialogue interface 31 then plays the translated Chinese voice message "Of course. Go straight for two blocks, then turn left. The subway entrance is right there." and displays the translated text.
[0105] In these implementations, the translation language of the first speaker's input is the same as that of the second speaker's input, and the translation language of the second speaker's input is the same as that of the first speaker's input. This greatly facilitates the interaction efficiency between the first and second speakers and improves the coherence of multilingual communication.
[0106] The above solution intelligently associates the current speech with historical translated speech and automatically inherits its translation language, giving the system the ability to perceive the context of dialogue, thereby greatly improving the coherence, naturalness, accuracy and efficiency of multilingual communication.
[0107] Please refer to Figure 4 , Figure 4 The flowchart of the speech translation method provided in this embodiment is shown in Figure 3. Figure 4 As shown, the method includes the following steps:
[0108] S401: Receives streaming voice input in the call interface.
[0109] In this embodiment, the call interface can be a call interface between the user and the intelligent agent. In this call interface, streaming voice can be input to the intelligent agent. That is, the user can continuously input voice to the intelligent agent in the call interface, and the intelligent agent can translate the streaming voice in real time.
[0110] That is, in this embodiment, the object of processing is not the "complete speech" that waits for the entire speech to finish before processing begins, but the continuous streaming speech input from the microphone.
[0111] In some implementations, the streaming speech to be processed is obtained based on the following method:
[0112] The audio is obtained by intercepting the streaming input audio based on the first audio length; or...
[0113] The streaming input speech is extracted based on the detected semantic boundaries.
[0114] To enable low-latency real-time processing, continuous streaming speech can be segmented into processable segments, each of which can be a stream of speech to be processed.
[0115] The streaming speech to be processed in this step can be the most recently input streaming speech.
[0116] In some application scenarios, a first speech length can be preset, which can be a time window, such as 300 milliseconds or 500 milliseconds. The aforementioned execution entity continuously receives streaming input speech and stores it in a first-in-first-out buffer. When the speech data accumulated in the buffer reaches the preset first speech length, the execution entity can extract the input speech within that time period as a streaming speech to be processed. After extraction, the buffer can be cleared, and the accumulation of input speech for the next time window can begin, and so on.
[0117] In some examples, the aforementioned first voice length is determined based on real-time network bandwidth and / or available system resources.
[0118] The initial speech length can be determined by real-time network bandwidth; a shorter initial speech length indicates lower network bandwidth, while a longer initial speech length indicates higher network bandwidth. By determining the initial speech length based on real-time network bandwidth and extracting the speech to be processed from the streaming input, the transmission latency of the speech in the network can be reduced, which is beneficial for improving the response speed of the translated speech.
[0119] The first speech length can also be determined based on available system resources. These available system resources may include, for example, the available processor and memory resources of the electronic device hosting the application client, the electronic device hosting the application server, and the electronic device running the relevant speech processing algorithms. If the available processor and memory are limited, the first speech length will be shorter; otherwise, the first speech length can be longer. By determining the first speech length based on available system resources, the translation efficiency of the speech extracted from the first speech length can be improved.
[0120] In this application scenario, the translated speech has a fixed latency and good real-time performance. Moreover, the algorithm has low complexity, does not rely on complex semantic analysis, and consumes few computational resources.
[0121] In some application scenarios, the audio is truncated based on detected semantic boundaries to obtain the streaming speech to be processed. In these scenarios, using a dynamic, adaptive, and semantically driven truncation method to truncate the streaming input speech can ensure the semantic integrity of each processing unit and the quality of translation processing.
[0122] Specifically, speech endpoint detection can be performed first. Streaming input speech can be analyzed in real time to detect low-energy silence segments. When the duration of a detected silence exceeds a preset threshold (e.g., 200 milliseconds), a potential semantic boundary (e.g., a pause between sentences) can be preliminarily identified. The speech segment preceding this potential semantic boundary point is then input into a lightweight automatic speech recognition and natural language understanding module (or punctuation prediction model) for analysis. This model analyzes the text content of the segment to determine if it constitutes a complete semantic unit. If the model confirms that the point is a reasonable semantic boundary, the continuous speech preceding this point can be extracted as a single streaming speech segment to be processed.
[0123] This extraction method allows the streaming speech to be processed into phrases or sentences with complete meaning, which reduces the burden on the subsequent translation process and improves the accuracy and coherence of the output results.
[0124] S402: Determine the source language of the speech to be processed.
[0125] S403 receives the translated speech and translated text obtained from the agent's convection speech translation.
[0126] The aforementioned execution entity can continuously send streaming speech to the intelligent agent, which can continuously translate each speech to obtain the translated speech and translated text of each streaming speech.
[0127] The intelligent agent can continuously send the translated speech and translated text corresponding to each stream of speech to be processed to the application client.
[0128] S404: Displays translated voice and translated text on the call interface.
[0129] Please refer to Figure 5A , Figure 5A This diagram illustrates the input of streaming voice and the display of translated voice and text in a call interface. Users can open call interface 51 in the application client and conduct voice calls with the intelligent agent from within call interface 51, such as... Figure 5A As shown in the left figure, the application client receives the streaming voice input in the call interface 51 and sends the currently input voice segment as streaming voice to be processed to the intelligent agent.
[0130] The agent detects the source language of the input streaming speech in real time and generates translated speech and text in the target language. The agent can also send the real-time translated speech and text to the application client.
[0131] The translated voice and translated text are displayed by the application client in the call interface 51. For example... Figure 5A As shown in the right figure, the source language of the detected streaming input speech is English. The translation language is Chinese.
[0132] When a user inputs the streaming speech "Good morning. Today I'd like to introduce some plant varieties to you. Plants are usually classified by family, for example, Rosaceae, Orchidaceae, Poaceae, Fabaceae, and Asteraceae.", the application client can truncate the streaming speech based on semantic boundary detection or the first speech length, and send the truncated input speech as the speech to be processed to the intelligent agent for translation. The intelligent agent sends the translated speech and translated text of each streaming speech to the application client in real time for display. That is, the application client also streams the translated speech and translated text.
[0133] As shown Figure 5A in the right figure, the current streaming voice to be processed is "and Asteraceae", the voice played in the call interface 51 is "and Asteraceae, etc." in Chinese, and "and Asteraceae, etc." is highlighted in the call interface 51.
[0134] In some embodiments, the above step S402 includes:
[0135] In response to the language of the voice to be processed not changing compared to the language of the processed voice, generating a first translated voice and a first translated text of the voice to be processed based on the first translated language corresponding to the processed voice; the language of the processed voice block is the first source language. And the above step S404 includes:
[0136] Displaying the first translated voice and the first translated text in the call interface.
[0137] The above execution entity can perform context detection. When a new streaming voice to be processed is input, it can be used to perform language analysis on it. Comparing the identified language characteristics of the voice to be processed with the first source language characteristics of the processed voice stored in the context state. If the similarity is higher than a preset threshold, it is determined that the language has not changed.
[0138] The source language of the historical processed voice (for example, the nearest processed voice to the current voice to be processed) is the first source language. If it is detected that the source language of the current voice to be processed has not changed compared to the processed voice, then the first translated language corresponding to the processed voice can be used to translate the voice to be processed.
[0139] In the call interface, the above first translated voice can be played through the speaker. The translated text corresponding to the processed voice can also be displayed in the call interface. That is, the above first translated text is displayed after the translated text of the processed voice in the call interface. The translated text of the processed voice and the first translated text displayed in the call interface are coherent in language.
[0140] In these embodiments, when it is detected that the language of the voice to be processed has not changed, the translated voice of the voice to be processed can be generated through the first translated language of the processed voice. There is no need to determine the source language of the voice to be processed again, thereby improving the speed of generating the translated voice, reducing the response delay for the voice to be processed, and reducing the consumption of computing resources.
[0141] In some embodiments, the above step S402 further includes: In response to the language of the voice to be processed changing compared to the language of the processed voice, determining the second source language of the voice to be processed.
[0142] Generate the second translated speech and the second translated text in the second translated language;
[0143] Step S404 above includes:
[0144] The call interface displays the second translated voice and the second translated text.
[0145] In these implementations, when a change in the language of the language to be processed is detected between the language of the processed speech and the language of the processed speech, a second source language of the speech to be processed can be determined in real time, and translated speech corresponding to the second translation language can be generated.
[0146] The aforementioned execution entity can perform context detection. When the similarity between the language features of a newly input streaming speech and the first source language features stored in the context state is lower than a preset threshold, it determines that the language has changed. This confirms that the language of the current speech being processed is a second source language (e.g., changing from Chinese to Japanese).
[0147] In these implementations, by sensing the dynamic changes in the language of the streaming input speech, the system dynamically identifies the changed second source language and generates a second translated speech and text in the corresponding second translation language. These are then displayed in real-time on the call interface. This eliminates the need for manual user intervention or resetting, automatically sensing and responding to language switching behavior. This enables continuous and accurate translation in multilingual scenarios, significantly improving the fluency of cross-language interaction. Simultaneously, the second translated speech maintains the same acoustic characteristics as the current streaming speech, ensuring the naturalness of the speech output and the speaker's identity after language switching. The real-time updating and display of the translated speech and text corresponding to the new language on the call interface allows all parties in the conversation to understand the semantic changes instantly and accurately, effectively supporting dynamic multilingual communication scenarios.
[0148] In some implementations, the streaming voice to be processed includes voices from at least two users; step S403 includes:
[0149] For each user's speech to be processed, receive translated speech and translated text generated by the agent that have the same acoustic features as the user's speech to be processed, wherein at least two users use different translation languages for their respective translated speech; and the above step S404 includes:
[0150] The call interface displays the translated audio and translated text corresponding to the voice messages to be processed for at least two users.
[0151] Each user's translated text includes a corresponding identifier. By displaying the association between user identifiers and translated text, clear differentiation and orderly management of multi-user dialogues are achieved. In scenarios where multiple people communicate simultaneously, users can accurately identify the source of each sentence, maintaining the logic and traceability of the dialogue. This solves the problem of source confusion in multi-user scenarios.
[0152] In one example, the speech samples from at least two users can be input sequentially. Each speech sample carries its own timestamp. The agent can determine the source language of each speech sample according to the input order and perform translation using the corresponding target language.
[0153] For each user, the aforementioned agent can determine the source language and target language of the user's speech to be processed. An encoder generates a speech encoding sequence, which includes linguistic content representations of acoustic features (such as timbre, prosody, intonation, and speech rate). A decoder generates an abstract second language content representation sequence and an acoustic feature sequence corresponding to the target language based on the speech encoding sequence. A vocoder then generates the translated speech waveform corresponding to the abstract second language content representation sequence based on the acoustic feature sequence. Furthermore, the agent simultaneously generates the translated text. This translated text can be obtained through speech recognition of the translated speech, or it can be obtained by mapping the speech encoding sequence to text using a text decoder in a multi-task decoder.
[0154] In these implementations, the aforementioned intelligent agent sends the translated audio and translated text corresponding to the voice recordings of at least two users, along with their respective timestamps, to the application client. The application client can then play the translated audio corresponding to each user's voice recording based on the relative relationships between the translated audio recordings reflected by the timestamps. Similarly, the translated text can be displayed sequentially in the call interface according to the relative relationships between the translated text recordings reflected by the timestamps, showing the translated text corresponding to each user's voice recording.
[0155] Figure 5B This is an illustration of multiple people inputting voice messages and displaying translated voice and text in a call interface. (Example) Figure 5BAs shown, in the call interface 51 of the application client, the speaker 2 inputs the streaming English input voice "Good morning. Today I’d like to introduce some plant varieties to you. Plants are usually classified by family for example, Rosaceae, Orchidaceae, Poaceae, Fabaceae, and Asteraceae." The English text 53 of the input voice can be streamed and displayed in the call interface 51. The application client can intercept the streaming input voice into multiple segments of streaming voice to be processed according to the first voice length or semantic detection, and send the streaming voice to be processed to the agent. The agent real-time identifies the source language (English) of the streaming voice to be processed and generates the translated voice and translated text in the translated language (Chinese). The agent real-time sends the translated voice and translated text to the application client, and the application client can stream and display the translated voice and translated text in the call interface. As Figure 5B shown, the translated voice "and Asteraceae etc." corresponding to the current streaming voice to be processed "and Asteraceae" can be played by voice in the call interface 51. The translated text "Good morning, today I would like to introduce some plant varieties to you. Usually, we classify plants by family. For example, Rosaceae, Orchidaceae, Poaceae, Fabaceae, and Asteraceae etc." 54 can be streamed and displayed in the call interface 51. Among them, "and Asteraceae etc." is highlighted.
[0156] The speaker 3 can input the streaming Chinese input voice "Are there any other classification methods for plants? And how do these classification methods differ from each other?" The Chinese text 55 of the input voice can be streamed and displayed in the call interface 51. And the translated voice "differ from each other" corresponding to the current streaming voice to be processed "how do these classification methods differ from each other" is played by voice. The translated text "Are there any other ways to classify plants?And how do these classification methods differ from each other?" 56 can also be streamed and displayed in the call interface 51. Among them, "differ from each other" is highlighted.
[0157] In one example, if the voices to be processed of at least two users are input synchronously, the voices to be processed of the above at least two users can be input successively. Each voice to be processed carries its own timestamp.
[0158] In these examples, the aforementioned agent can first separate the speech from at least two users, creating multiple speech streams. The agent can then perform translation operations on each of these multiple speech streams in parallel, including determining the source language and generating the corresponding translation speech and text for the target language.
[0159] When presented in the call interface, spatial audio separation can be performed, for example, using different channels (left channel, right channel, and center channel) to play the translation voice corresponding to different users.
[0160] In these implementations, when the streaming speech includes at least two users' speech, the agent can dynamically determine the source language and translation language for each user and generate corresponding translated speech for each user. This enables barrier-free real-time cross-language communication by processing the speech input of multiple users in parallel and generating translated speech that preserves the original acoustic features. Each user can communicate naturally in their native language, and the agent automatically completes real-time translation between multiple languages. This transforms the traditional serial translation mode into a parallel processing mode, significantly improving the timeliness of cross-language communication. It solves the problems of high latency and low efficiency caused by intermediate translation steps in multilingual communication. Furthermore, by maintaining the consistency of the acoustic features of the translated speech with the original speech, effective transmission of emotional information is achieved.
[0161] In this embodiment, by inputting streaming speech into the call interface, the intelligent agent performs real-time translation of the streaming speech, achieving low-latency, high-accuracy real-time translation of long speech. This provides users with an immersive and efficient cross-language information transmission mode, overcoming the limitations of language barriers on real-time information transmission and reception during cross-language information interaction.
[0162] In some embodiments, the method further includes:
[0163] After translating the streaming input speech, a translation record is generated; among which,
[0164] The translation record includes a summary described by the target language, a streamed input speech consisting of multiple streamed speech segments to be processed, a first text, and a second text. The first text is the text corresponding to the streamed input speech, and the second text consists of the texts corresponding to the multiple translated speech segments respectively.
[0165] In these implementations, the user can end the translation of the streaming voice input within the call interface, for example, by closing the call interface during the call or ending the call with the agent within the call interface. After ending the translation of the streaming voice input, the agent can generate a translation record and display a prompt message in the application client indicating that a translation record of the streaming voice input is available for viewing.
[0166] Users can view the above translation history to review the streaming input speech and its corresponding translated text.
[0167] Figure 6 This is an illustration of a translation log prompt. After the translation of streaming speech ends in the call interface, a translation log is generated. A prompt message for the translation log is then displayed on the homepage (61) of the application client. Figure 6 As shown, the prompt information includes phrases like "Translation Record," the title of the translation record "Plant Variety Introduction," the generation time of the translation record, the streaming audio information of the translation record, controls for playing the streaming audio, and may also include a summary of the translation record, such as... Figure 6 The text describes how plant varieties are categorized by "family," with detailed descriptions of the characteristics of each family.
[0168] In these embodiments, traceability of the translation process is achieved by constructing a complete translation record containing multimodal elements. Furthermore, by presenting streaming input speech, first text, and second text, the original content can be compared with the translation result, providing data support for the agent's translation optimization. Intelligent summarization and rapid browsing of the translated content are achieved by generating a summary described in the translation language.
[0169] In some implementations, the method further includes:
[0170] In response to a playback command for the translation recording, play the input audio.
[0171] Based on the playback progress of the input voice, the first and second text segments corresponding to the playback progress are highlighted synchronously.
[0172] In these implementations, the timestamp of each stream of speech to be processed can be used to associate and store the stream of speech to be processed, the text of the stream of speech to be processed, and the translated text of the stream of speech to be processed. When playing the speech in the translation record, the text (as a first text segment) and the translated text (as a second text segment) of the currently playing speech segment (corresponding to an input speech to be processed) can be located according to the timestamp of the playback progress indicator, and the first text segment and the second text segment are highlighted first.
[0173] In these implementations, the problem of poor audio-text synchronization can be improved by aligning the timing of the audio and text streams. Furthermore, by automatically scrolling and positioning the first and second texts during audio playback, it is ensured that the currently highlighted text remains in the visible area, which helps users obtain synchronized information both auditorily and visually, thus improving information retrieval efficiency.
[0174] Corresponding to the speech translation method in the above embodiments, Figure 7This is a structural block diagram of a speech translation device provided in an embodiment of this disclosure. For ease of explanation, only the parts relevant to the embodiments of this disclosure are shown. (Refer to...) Figure 7 The voice translation device 70 includes: a determining unit 701, a generating unit 702, and a display unit 703.
[0175] in,
[0176] The determining unit 701 is used to determine the source language of the speech to be processed.
[0177] The generation unit 702 is used to generate the translated speech and translated text in the target language; the translated speech has the same acoustic features as the speech to be processed; the target language is determined by the user or the usage environment.
[0178] Display unit 703 is used to display the translated audio and translated text.
[0179] The media content-based information processing apparatus provided in this disclosure first determines the source language, then generates translated speech and translated text in the target language, and finally plays and displays them simultaneously. This achieves dynamic determination of the translation direction, using this dynamically determined direction to perform speech and text translation of the speech to be processed, thus improving information transmission efficiency in cross-language interaction scenarios. Furthermore, directly generating translated speech from the speech to be processed achieves end-to-end translation, eliminating the switching delays and manual intervention between multiple independent modules such as language recognition, speech recognition, and machine translation, simplifying the operation process, reducing the overhead of data transmission and processing between multiple independent modules, and lowering response latency. In addition, by simultaneously presenting translated speech and corresponding text, a multimodal output method using both visual and auditory information channels produces a synergistic effect, meeting diverse needs for barrier-free communication. Moreover, the translated speech retains the acoustic characteristics of the speech to be processed, improving the naturalness and recognizability of the translated speech.
[0180] In one embodiment of this disclosure, the generation unit 702 is further configured to:
[0181] In response to the instruction to use the original speech for translation, a translated speech is generated based on the acoustic features of the speech to be processed and the target language.
[0182] In summary, and further, by using the original speech as an instruction, translated speech is generated based on the acoustic features of the speech to be processed, thus enabling precise user control over the translated speech. Furthermore, the extraction of acoustic features from the speech to be processed through explicit instruction authorization enhances privacy protection and security control in acoustic feature processing.
[0183] In one embodiment of this disclosure, the apparatus 70 further includes a feature extraction unit (not shown in the figures), the feature extraction unit being used for:
[0184] Acoustic features are extracted from at least a portion of the speech to be processed.
[0185] In summary, further, using a portion of the speech to be processed to extract acoustic features can improve the speed of acoustic feature acquisition and thus reduce the response delay of the translated speech. Using the entire speech to be processed to extract acoustic features can improve the accuracy of the extracted acoustic features, which is beneficial to improving the consistency between the generated translated speech and the original speech.
[0186] In one embodiment of this disclosure, the generation unit 702 is further configured to:
[0187] The speech to be processed is sent to the speech model, which extracts the acoustic features of the speech frame by frame.
[0188] The speech model is used to generate translated speech based on the acoustic features of the speech to be processed extracted frame by frame and the target language.
[0189] In summary, this approach further utilizes a speech model to extract auxiliary acoustic features from the speech to be processed and generate translated speech. Compared to traditional cascaded systems (speech recognition first, then machine translation, and finally speech synthesis) for speech translation, this scheme uses a single model to complete the conversion from source speech to target speech, reducing error propagation and improving the stability and accuracy of the generated translated speech. Combining multiple steps into a single model simplifies the system architecture and reduces the complexity of deployment and maintenance. By using the same model to process acoustic features and generate translated speech, the acoustic characteristics of the source speech (such as timbre and intonation) are better preserved, making the translated speech closer to the original speaker's voice. Furthermore, by extracting acoustic features frame by frame, the model achieves ultra-refined analysis and processing of the speech signal, making the translated speech closer to the original speech and better preserving its acoustic characteristics (such as timbre, intonation, and speech rate), accurately conveying the emotional information of the original speech.
[0190] In one embodiment of this disclosure, the device 70 further includes a first receiving unit (not shown in the figures), the first receiving unit being used for:
[0191] Receive voice input in the chat interface;
[0192] In response to the voice input end command, the voice input in the dialog interface will be treated as voice to be processed.
[0193] Display unit 703 is further used for:
[0194] The translated audio and translated text are displayed in the dialog interface.
[0195] In summary, further, within the human-computer dialogue interface, the user-input speech is used as the speech to be processed, and the source language of the speech is dynamically determined. A translated speech with the same acoustic features as the input speech is generated in the corresponding translation language. The translated speech and translated text are then displayed in the dialogue interface. Because end-to-end processing from speech input to multilingual translation is achieved within a single dialogue interface, users do not need to perform any language switching operations. The entire process of language recognition, translation generation, and result presentation is automatically completed, eliminating technical barriers and operational obstacles to cross-language communication. Furthermore, the multimodal (speech and text) information fusion output in the dialogue interface can meet the receiving preferences of different users and improve the redundancy and reliability of information transmission. The comprehensibility and emotional fidelity of the speech translation results are enhanced: because the translated speech reproduces the acoustic characteristics of the speech to be processed, it not only conveys the semantic content but also retains the speaker's emotions, tone, and other paralinguistic information, helping users to more accurately understand the semantic intent of the speech to be processed.
[0196] In one embodiment of this disclosure, the device 70 further includes a second receiving unit (not shown in the figures), the second receiving unit being used for:
[0197] Receive streaming voice messages input in the call interface;
[0198] The generation unit 702 is further used for:
[0199] Receive the translated speech and translated text obtained from the agent's concurrency-based speech translation;
[0200] Display unit 703 is further used for:
[0201] The call interface displays the translated voice and translated text.
[0202] In summary, going a step further, by inputting streaming speech into the call interface and having the intelligent agent translate the streaming speech in real time, it is possible to achieve low-latency, high-accuracy real-time translation of long speech. This provides users with an immersive and efficient cross-language information transmission mode, overcoming the limitations of language barriers on real-time information transmission and reception during cross-language information interaction.
[0203] In one embodiment of this disclosure, the speech to be processed is speech in a multi-turn dialogue; the participants in the multi-turn dialogue include at least two speakers, wherein the at least two speakers use two languages to conduct the dialogue; the device 70 further includes a speech determination unit (not shown in the figure), which is configured to perform the following operations before the determination unit determines the source language corresponding to the speech to be processed:
[0204] Determine the first speaker of the voice input in the current round. The first speaker can be any one of at least two speakers.
[0205] In response to the fulfillment of the trigger condition, the voice input by the first speaker in the current round is taken as the voice to be processed.
[0206] In summary, going a step further, identifying the speaker in the current round of the dialogue interface and then combining this with trigger condition judgment can ensure that the speaker's voice is only processed when the user's true intention is clear. This can reduce false triggers in noisy environments and improve the intelligence and accuracy of voice translation.
[0207] In one embodiment of this disclosure, the voice input by the first speaker is used to reply to the historical translated voice corresponding to the second speaker; the language used in the historical translated voice is the same as the language used in the voice to be processed.
[0208] In summary, by intelligently associating the current speech with historical translated speech and automatically inheriting its translation language, the system is endowed with the ability to perceive the context of dialogue, thereby greatly improving the coherence, naturalness, accuracy and efficiency of multilingual communication.
[0209] In one embodiment of this disclosure, the streaming voice to be processed includes voices from at least two users;
[0210] The generation unit 702 is further used for:
[0211] For each user's speech to be processed, a translated speech and a translated text generated by the agent that have the same acoustic features as the user's speech to be processed are received, wherein at least two users use different translation languages for their respective translated speech;
[0212] Display unit 703 is further used for:
[0213] The call interface displays the translated audio and translated text corresponding to the voice messages to be processed for at least two users; the translated text for each user includes the identifier corresponding to that user.
[0214] In summary, when the streaming speech to be processed includes speech from at least two users, the agent can dynamically determine the source language and translation language for each user and generate corresponding translated speech for each user. This achieves barrier-free real-time cross-language communication by processing the speech input of multiple users in parallel and generating translated speech that preserves the original acoustic features. Each user can communicate naturally in their native language, and the agent automatically completes real-time translation between multiple languages. This transforms the traditional serial translation mode into a parallel processing mode, significantly improving the timeliness of cross-language communication. It solves the technical problems of high latency and low efficiency caused by intermediate translation steps in multilingual communication. Furthermore, by maintaining the consistency of the acoustic features of the translated speech with the original speech, effective transmission of emotional information is achieved.
[0215] In one embodiment of this disclosure, the determining unit 701 is further configured to:
[0216] Since the language of the speech to be processed has not changed compared to the language of the processed speech, a first translated speech and a first translated text are generated based on the first translation language corresponding to the processed speech block; the language of the processed speech block is the first source language.
[0217] Display unit 703 is further used for:
[0218] The first translated voice and the first translated text are displayed in the call interface.
[0219] In summary, further, when no language change is detected in the speech to be processed, the translated speech can be generated from the first translation language of the processed speech. This eliminates the need to determine the source language of the speech again, thereby improving the speed of generating translated speech, reducing response latency to the speech being processed, and reducing computational resource consumption.
[0220] In one embodiment of this disclosure, the determining unit 701 is further configured to:
[0221] In response to a change in the language of the speech to be processed compared to the language of the processed speech, a second source language of the speech to be processed is determined;
[0222] Generate the second translated speech and the second translated text in the second translated language;
[0223] Display unit 703 is further used for:
[0224] The call interface displays the second translated voice and the second translated text.
[0225] In summary, by sensing the dynamic changes in the language of the streaming input speech, the system dynamically identifies the changed second source language and generates a second translated speech and text in the corresponding second translation language. These are then displayed in real-time on the call interface. This eliminates the need for manual user intervention or resetting, automatically sensing and responding to language switching behavior. This enables continuous and accurate translation in multilingual scenarios, significantly improving the fluency and practicality of cross-language interaction. Furthermore, the second translated speech maintains the same acoustic characteristics as the current streaming speech, ensuring the naturalness of the speech output and the speaker's identity after language switching. The real-time updating and display of the translated speech and text corresponding to the new language on the call interface allows all parties in the conversation to understand semantic changes instantly and accurately, effectively supporting dynamic multilingual communication scenarios.
[0226] In one embodiment of this disclosure, the apparatus 70 further includes a segmentation unit (not shown), which is configured to acquire streaming speech to be processed based on the following steps:
[0227] The audio is obtained by intercepting the streaming input audio based on the first audio length; or...
[0228] The streaming input speech is extracted based on the detected semantic boundaries.
[0229] In one embodiment of this disclosure, the first voice length is determined based on real-time network bandwidth and / or available system resources.
[0230] In one embodiment of this disclosure, the apparatus 70 further includes a recording generation unit (not shown in the figures), the recording generation unit being used for:
[0231] After translating the streaming input speech, a translation record is generated; among which,
[0232] The translation record includes a summary described by the target language, a streamed input speech consisting of multiple streamed speech segments to be processed, a first text, and a second text. The first text is the text corresponding to the streamed input speech, and the second text consists of the texts corresponding to the multiple translated speech segments respectively.
[0233] In summary, further, by constructing a complete translation record containing multimodal elements, traceability of the translation process is achieved. Furthermore, by presenting streaming input speech, the first text, and the second text, the original content can be compared with the translation result, providing data support for the agent's translation optimization. By generating summaries described in the translation language, intelligent summarization and rapid browsing of the translated content are achieved.
[0234] In one embodiment of this disclosure, the apparatus 70 further includes a recording and display unit (not shown in the figures), the recording and display unit being used for:
[0235] In response to a playback command for the translation recording, play the input audio.
[0236] Based on the playback progress of the input voice, the first and second text segments corresponding to the playback progress are highlighted synchronously.
[0237] In summary, further, by using the timestamp of each stream of audio to be processed, the stream of audio to be processed, the text of the stream of audio to be processed, and the translated text of the stream of audio to be processed are associated and stored. When playing audio in the translation record, the text of the currently playing stream of audio to be processed (as a first text segment) and the translated text (as a second text segment) can be located according to the timestamp of the playback progress indicator, and the first and second text segments are highlighted first. In these embodiments, by aligning the timing of the audio stream and the text stream, the problem of poor audio-text synchronization can be improved. In addition, by automatically scrolling and positioning the first and second texts during audio playback, it is ensured that the currently highlighted text is always in the visible area, which helps users obtain synchronized information from both auditory and visual perspectives, thus improving information acquisition efficiency.
[0238] In one embodiment of this disclosure, the translation language is determined based on one of the following:
[0239] Determined based on the language configured by the user;
[0240] Determined based on the context of the speech to be processed;
[0241] Determined based on the location and environment of the speech to be processed;
[0242] It is determined based on the language environment used by both parties during the voice call.
[0243] To implement the above embodiments, this disclosure also provides an electronic device.
[0244] refer to Figure 8 The diagram illustrates a structural schematic of an electronic device 800 suitable for implementing embodiments of the present disclosure. The electronic device 800 can be a terminal device or a server. The terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, personal digital assistants (PDAs), tablet computers, portable media players (PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 8 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0245] like Figure 8As shown, the electronic device 800 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. The RAM 803 also stores various programs and data required for the operation of the electronic device 800. The processing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0246] Typically, the following devices can be connected to I / O interface 805: input devices 806 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 807 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 808 including, for example, magnetic tapes, hard disks, etc.; and communication devices 809. Communication device 809 allows electronic device 800 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8 An electronic device 800 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0247] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 809, or installed from a storage device 808, or installed from a ROM 802. When the computer program is executed by a processing device 801, it performs the functions defined in the methods of embodiments of this disclosure.
[0248] It should be noted that the computer-readable storage medium described in this disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable storage medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0249] The aforementioned computer-readable storage medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0250] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the method shown in the above embodiments.
[0251] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0252] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0253] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".
[0254] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0255] The electronic device, computer-readable storage medium, and computer program product provided in this disclosure first determine the source language, then generate translated speech and translated text in the target language, and finally play and display them simultaneously. This achieves dynamic determination of the translation direction, using the dynamically determined translation direction to perform speech and text translation of the speech to be processed, improving information transmission efficiency in cross-language interaction scenarios. Furthermore, directly generating translated speech from the speech to be processed achieves end-to-end translation, eliminating the switching delays and manual intervention between multiple independent modules such as language recognition, speech recognition, and machine translation, simplifying the operation process, reducing the overhead of data transmission and processing between multiple independent modules, and lowering response latency. In addition, by simultaneously presenting translated speech and corresponding text, a multimodal output method using both visual and auditory information channels produces a synergistic effect, meeting diverse needs for barrier-free communication. Moreover, the translated speech retains the acoustic characteristics of the speech to be processed, improving the naturalness and recognizability of the translated speech.
[0256] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0257] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0258] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A speech translation method, characterized in that, include: Determine the source language of the speech to be processed; Generate the translated audio and translated text for the target language; The translated speech has the same acoustic features as the speech to be processed; The translation language is determined by the user or the usage environment; The translated audio and the translated text are displayed.
2. The method according to claim 1, characterized in that, The generation of translated speech and translated text in the translated language includes: In response to an instruction to use the original speech for translation, a translated speech is generated based on the acoustic features of the speech to be processed and the target language.
3. The method according to claim 2, characterized in that, The method further includes: The acoustic features are extracted from at least a portion of the speech to be processed.
4. The method according to claim 2, characterized in that, The generation of translated speech and translated text in the translated language includes: The speech to be processed is sent to a speech model, which extracts the acoustic features of the speech frame by frame. The speech model is used to generate a translated speech based on the acoustic features of the speech to be processed extracted frame by frame and the target language.
5. The method according to any one of claims 1-4, characterized in that, The method further includes: Receive voice input in the chat interface; In response to the voice input end command, the voice input in the dialogue interface is taken as the voice to be processed; The presentation of the translated speech and the translated text includes: The translated audio and translated text are displayed in the dialogue interface.
6. The method according to any one of claims 1-4, characterized in that, The method further includes: Receive streaming voice messages input in the call interface; The generation of translated speech and translated text in the translated language includes: Receive the translated speech and the translated text obtained by the agent's concurrency-based speech translation; The presentation of the translated speech and the translated text includes: The translated voice and the translated text are displayed on the call interface.
7. The method according to claim 5, characterized in that, The speech to be processed is speech from a multi-turn dialogue; the participants in the multi-turn dialogue include at least two speakers, wherein the at least two speakers are conversing in two languages; Before determining the source language corresponding to the speech to be processed, the method further includes: Determine the first speaker of the voice input in the current round, wherein the first speaker is any one of the at least two speakers; In response to the fulfillment of the triggering condition, the speech input by the first speaker in the current round is taken as the speech to be processed.
8. The method according to claim 7, characterized in that, The voice input by the first speaker is used to respond to the historical translated voice corresponding to the second speaker; the language used in the historical translated voice is the same as the language used in the voice to be processed.
9. The method according to claim 6, characterized in that, The streaming voice to be processed includes voice messages from at least two users. The receiving of the translated speech and the translated text obtained by the agent's streaming speech translation includes: For each user's speech to be processed, a translated speech and a translated text generated by the agent that have the same acoustic features as the user's speech to be processed are received, wherein the translated speech for the at least two users are translated in different languages; The display of the translated speech and the translated text includes: The call interface displays the translated voice and translated text corresponding to the voices to be processed for at least two users, respectively; wherein, the translated text for each user includes the identifier corresponding to that user.
10. The method according to claim 6, characterized in that, The process of determining the source language of the speech to be processed includes: In response to the fact that the language of the speech to be processed has not changed compared to the language of the processed speech, a first translated speech and a first translated text of the speech to be processed are generated based on the first translation language corresponding to the processed speech block; the language of the processed speech block is the first source language; The display of the translated speech and the translated text includes: The first translated voice and the first translated text are displayed in the call interface.
11. The method according to claim 10, characterized in that, The step of determining the source language of the speech to be processed also includes: In response to a change in the language of the speech to be processed compared to the language of the processed speech, a second source language of the speech to be processed is determined; Generate the second translated speech and the second translated text in the second translated language; The display of the translated speech and the translated text includes: The second translated voice and the second translated text are displayed in the call interface.
12. The method according to claim 6, characterized in that, The streaming speech to be processed is obtained based on the following method: The audio is obtained by intercepting the streaming input audio based on the first audio length; or... The streaming input speech is intercepted based on the detected semantic boundaries.
13. The method according to claim 12, characterized in that, The first voice length is determined based on real-time network bandwidth and / or available system resources.
14. The method according to claim 12, characterized in that, The method further includes: After translating the streaming input speech, a translation record is generated; wherein, The translation record includes a summary described by the translation language, a streaming input speech consisting of multiple streaming speech segments to be processed, a first text, and a second text. The first text is the text corresponding to the streaming input speech, and the second text consists of the text corresponding to each of the multiple translated speech segments.
15. The method according to claim 14, characterized in that, The method further includes: In response to a playback command for the translation recording, the input voice is played; Based on the playback progress of the input voice, the first text segment and the second text segment corresponding to the playback progress are highlighted synchronously.
16. The method according to claim 1, characterized in that, The translation language is determined based on one of the following: Determined based on the language configured by the user; Determined based on the context of the speech to be processed; Determined based on the location and environment of the speech to be processed; It is determined based on the language environment used by both parties during the voice call.
17. A voice translation device, characterized in that, include: The determining unit is used to determine the source language of the speech to be processed. The generation unit is used to generate the translated speech and translated text for the target language; The translated speech has the same acoustic features as the speech to be processed; The translation language is determined by the user or the usage environment; The display unit is used to display the translated speech and the translated text.
18. An electronic device, characterized in that, include: Processor and memory; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1 to 14.
19. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the method as described in any one of claims 1 to 14.
20. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 14.
Citation Information
Patent Citations
Translation voice output method and device, storage medium and electronic equipment
CN110083846A
Voice information processing method and device and electronic equipment
CN113571044A
Speech translation method, system and device and storage medium
CN118228742A
Speech translation method, device, and storage medium
US20240028841A1