A method and system for translating VoIP calls

By mixing and synthesizing the caller's original voice and the transliterated voice during VoIP calls, the problem of long waiting time for the transliteration is solved, achieving real-time translation of the caller's continuous speech and improving the user experience.

CN122372680APending Publication Date: 2026-07-10QUEEN BEE NETWORK TECH (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
QUEEN BEE NETWORK TECH (SHENZHEN) CO LTD
Filing Date
2026-05-22
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

In existing VoIP call translation technology, the called party has to wait a long time for the translated voice when the caller speaks continuously, resulting in poor real-time performance. Furthermore, if the silence timeout threshold is too high, the caller cannot speak multiple sentences continuously, leading to a poor experience as well.

Method used

The system generates a transliteration of the Nth sentence immediately after the caller finishes speaking, mixes it with the original audio of the (N+1)th sentence to create a mixed frame, and sends it to the called party's terminal for playback. Different sound features are used to control parameters to make the transliteration more prominent.

Benefits of technology

This allows the called party to hear the translated audio faster and the called party to speak multiple sentences continuously, improving the real-time performance and user experience of the translation function.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122372680A_ABST
    Figure CN122372680A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for translating VoIP calls. The method includes: when a server receives a first original audio media stream uploaded by a first terminal, it detects whether a first original audio frame and a first translated audio frame exist simultaneously within a first time overlap window; if not, it sends either the first original audio frame or the first translated audio frame existing within the first time overlap window to a second terminal for playback; if so, it mixes and synthesizes the corresponding first original audio frame and first translated audio frame within the first time overlap window to generate a first mixed frame, and sends the first mixed frame to the second terminal for playback. This invention enables the called party to hear the translated audio faster and allows the calling party to speak multiple sentences continuously.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of real-time communication translation technology, and in particular to a method and system for translating VoIP calls. Background Technology

[0002] To provide services to multilingual customers, related technologies have been introduced into specific real-time communication services to achieve real-time call translation functionality. In this specific real-time communication service, the caller can initiate a voice call through a real-time communication application (APP), and the called party does not need to install the same real-time communication application as the caller. They can simply answer the call on their mobile phone to hear the translation of the caller's speech. Furthermore, during the call, the caller can also hear the translation of the called party's speech through the real-time communication application. The aforementioned real-time call translation function adopts a "question-and-answer" format. For the called party's terminal, it strictly follows the first round of question-and-answer: "Play the caller's first round of original audio -> Translation waiting (server translating the caller's original audio) -> Play the translated version of the caller's first round of original audio -> Wait for the translated version to finish playing -> Collect the called party's original audio." The second round of question-and-answer is: "Play the caller's second round of original audio -> Translation waiting (server translating the caller's original audio) -> Play the translated version of the caller's second round of original audio -> Wait for the translated version to finish playing -> Collect the called party's original audio," and so on, realizing a multi-round "question-and-answer" real-time call translation function. While the called party's terminal plays the translated version of the caller's original audio, the calling party's terminal also simultaneously plays the translated version of the caller's original audio, allowing the calling party's terminal to hear the translated result of their own speech in real time.

[0003] Specifically, in scenarios where the caller needs to speak continuously, the shortcomings of the aforementioned real-time call translation function in the "question and answer" format are illustrated by two cases: a high and a low post-silence timeout threshold. In the first case, assuming the caller needs to speak three sentences continuously in the first round of speech before stopping, when the post-silence timeout threshold is low (e.g., 300 milliseconds), the pause time after each sentence (i.e., the post-silence duration) exceeds the post-silence timeout threshold, thus triggering each sentence to be treated as a speech segment and immediately performing speech recognition and translation on the corresponding speech segment. For the called party's terminal, the real-time call translation function in the form of "question and answer" is as follows: "Play the first original voice of the calling party -> translation waiting (the server is translating the original voice of the calling party) -> play the transliteration of the first original voice of the calling party -> wait for the transliteration to finish playing -> play the second original voice of the calling party -> translation waiting (the server is translating the original voice of the calling party) -> play the transliteration of the second original voice of the calling party -> wait for the transliteration to finish playing -> play the third original voice of the calling party -> translation waiting (the server is translating the original voice of the calling party) -> play the transliteration of the third original voice of the calling party -> wait for the transliteration to finish playing -> collect the original voice of the called party." In another scenario, assuming the caller needs to speak three sentences consecutively in the first round, when the back-silence timeout threshold is high (e.g., 1200 milliseconds), the pause duration of the first two sentences (i.e., the back-silence duration) does not exceed the back-silence timeout threshold. Only the pause duration of the third sentence exceeds the back-silence timeout threshold. Therefore, only after the caller finishes speaking the third sentence will the three sentences be treated as a single speech segment, triggering a speech recognition and translation of the entire speech segment corresponding to these three sentences. For the called party's terminal, the process of the real-time call translation function in the "question and answer" format is as follows: "Play the original audio of the three sentences spoken by the caller in the first round -> translation waiting (the server is translating the original audio of the caller's three sentences) -> play the transliterated audio of the original audio of the three sentences spoken by the caller in the first round -> wait for the transliteration to finish playing -> collect the original audio of the called party."

[0004] It is evident that the real-time call translation function provided by the relevant technology has the following conflicts in continuous speaking scenarios: If the after-silence timeout threshold is too low, the called party can hear the translated voice more quickly, with a shorter continuous waiting time, resulting in better real-time performance and a better experience. However, the caller cannot speak multiple sentences continuously, leading to a poor experience. On the other hand, if the after-silence timeout threshold is too high, the caller can speak multiple sentences continuously, resulting in a better experience. However, the called party needs to wait longer to hear the translated voice, resulting in poorer real-time performance and a poor experience. Summary of the Invention

[0005] The purpose of this invention is to at least solve one of the technical problems existing in the prior art, and to provide a method and system for translating VoIP calls, which can enable the called party to hear the translated voice faster and enable the caller to speak multiple sentences continuously.

[0006] To achieve the above-mentioned objectives, the present invention provides the following technical solution: Firstly, a method for translating VoIP calls is provided, the method comprising: When the server receives the first original audio media stream uploaded by the first terminal, it detects whether the first original audio frame and the first transliterated audio frame exist simultaneously within the first time overlap window. If not, the first original audio frame or the first transliterated audio frame existing in the first time overlap window will be sent to the second terminal for playback; If so, the first original audio frame and the first transliterated audio frame corresponding to the first time overlap window will be mixed and synthesized to generate a first mixed frame, and the first mixed frame will be sent to the second terminal for playback; Wherein, the first original audio media stream is composed of a plurality of first original audio frames arranged in chronological order; the first transliterated frame is a constituent unit for generating a corresponding first transliterated segment based on the speech segment extracted from the first original audio media stream, and the first transliterated segment is composed of a plurality of first transliterated frames arranged in chronological order. The audio mixing process includes applying different sound feature control parameters to the first transliterated frame and the first original frame, such that the first sound feature parameter of the first transliterated frame in the first mixed frame and the second sound feature parameter of the first original frame in the first mixed frame satisfy a preset difference condition.

[0007] Secondly, the present invention also provides a VoIP call translation system, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to perform the steps of the VoIP call translation method described above.

[0008] Compared with the prior art, the VoIP call translation method and system provided in this application have at least the following beneficial effects; Compared to existing technologies where the calling party's continuous speaking results in a longer waiting time for the called party to hear the translated audio, leading to poor real-time performance and a less than ideal user experience, this embodiment provides a VoIP call translation method. After the calling party finishes speaking the original audio of the Nth sentence, the method generates the translated audio corresponding to the Nth sentence while the calling party continues speaking the original audio of the N+1th sentence. The translated audio of the Nth sentence is then mixed with the original audio of the N+1th sentence to obtain a mixed media stream (composed of multiple mixed frames arranged chronologically), which is then sent to the called terminal. Because different sound feature control parameters are applied to the first transliterated frame and the first original frame in the mixed media stream, the first sound feature parameter of the first transliterated frame in the first mixed frame and the second sound feature parameter of the first original frame in the first mixed frame meet the preset difference conditions. Therefore, the called party can hear the mixed media stream with the first transliterated segment being more prominent, which allows the called party to hear the transliteration faster and the calling party to speak multiple sentences continuously.

[0009] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0010] The present invention will be further described below with reference to the accompanying drawings and embodiments; Figure 1 This is a structural block diagram of a VoIP call translation system in one embodiment; Figure 2 This is a flowchart illustrating a VoIP call translation method in yet another embodiment; Figure 3 This is a schematic diagram of the process for generating the first transliterated segment in one embodiment; Figure 4 This is a schematic diagram of the process for generating the first transliterated segment in another embodiment; Figure 5 This is a schematic diagram of the process of sending a marker signal to the second terminal after entering the Nth round of dialogue in one embodiment; Figure 6 This is a schematic diagram of the state of each queue in the first round of dialogue in one embodiment. Detailed Implementation

[0011] This section will describe in detail specific embodiments of the present invention. Preferred embodiments of the present invention are shown in the accompanying drawings. The purpose of the drawings is to supplement the textual description with graphics, so that people can intuitively and vividly understand each technical feature and overall technical solution of the present invention, but they should not be construed as limiting the scope of protection of the present invention.

[0012] It should be noted that, firstly, this application involves the following steps: On the App side, the original audio from the caller is collected and recognized. Then, it is translated according to the called party's language, synthesized according to the recipient's language, and the synthesized media stream is converted into telecommunications signaling. This signaling is then relayed to major partner telecommunications operators worldwide. Conversely, after the called party speaks the original audio, the telecommunications operator relays the original audio signaling back to our media conversion server, converts the signaling into a media stream, separates the audio from the stream, extracts and recognizes the original sound, translates it according to the caller's language, synthesizes it, and plays it back. Secondly, the software is cross-server. While other software requires both parties to download the same software for interconnection and translation, our software only requires the caller to download it to directly call mobile phones or landlines. Currently, it can call over 100 countries and regions worldwide.

[0013] Reference Figure 1 This is a diagram illustrating the application environment of a VoIP call translation method provided in this application. Figure 1As shown, the real-time translation system for VoIP calls provided in this embodiment includes a first terminal, a second terminal, and a server. The server communicates with the first terminal via the Internet and with the second terminal via the Internet telephony service provider (ITSP) network and the public switched telephone network (PSTN). The first terminal is a mobile terminal with a real-time communication application (APP) installed, used to initiate VoIP calls, collect the user's original voice, and play the other party's voice / translated voice. The first terminal establishes a service signaling connection with the server via an IP network and transmits voice media streams. The second terminal is a mobile terminal with a native telephone application, which answers incoming calls and conducts calls via a mobile phone number, without needing to install the real-time communication application on the first terminal. Calls between the second terminal and the server are connected via the ITSP / PSTN network. The server is used to complete call connection, voice media relay, and real-time translation processing. The server specifically includes a signaling control module, a media gateway module, a real-time translation module, and [other components not specified in the original text]. The signaling control module is used for session establishment and control, including service signaling control on the first terminal side (e.g., initiating calls, hanging up, status notifications, etc. via HTTP / HTTPS) and call control facing the ITSP network (e.g., SIP session establishment / release, routing control, etc.); and performs session management, including assigning session identifiers to each call, maintaining the call state machine, binding both terminals, and media resources. The media gateway module is used for voice media transmission and reception and processing, including RTP media transmission and reception, jitter buffering, encoding / decoding conversion, mixing / synthesis, and outputting downlink voice channels to the ITSP / PSTN side. The real-time translation module performs speech recognition, machine translation, and speech synthesis on voice segments, generating translated audio segments, and sending the translated audio segments back to the media gateway module for downlink playback. Specifically, the server accesses the number network through the ITSP network and interconnects with the PSTN network, enabling the server to connect VoIP calls initiated by the first terminal to the mobile phone number of the second terminal, thereby allowing the second terminal to answer the call in a native telephone manner. The signal tagging management module is used to generate, embed, and assign specific "tag signals" to the audio stream sent to the second terminal, which are used to assist in the task of extracting the target speaker from the call recording.

[0014] It should be noted that SIP (Session Initiation Protocol) in the signaling control module refers to the session initiation protocol, used to initiate a session. A SIP session is a session between two user terminals based on an IP network, i.e., a VoIP session. HTTP is an abbreviation for Hypertext Transfer Protocol, a protocol used to transfer hypertext from a WWW server to a local browser. In the real-time translation module, ASR (Automatic Speech Recognition), MT (Machine Translation), and TTS (Text-to-Speech) are used for automatic speech recognition, machine translation, and text-to-speech, respectively. Specifically, the real-time translation module in this embodiment can automatically determine the called party's language corresponding to the first translated segment sent to the second terminal, and can also automatically determine the calling party's language corresponding to the second translated segment sent to the first terminal. In one example, when registering for the real-time communication app on the first terminal, the caller must select a corresponding language. For example, if the caller selects English, the real-time translation module determines that the language of the second transliterated segment sent to the first terminal is English, based on the caller's registration information provided by the app. In another example, the server determines the language of the called party by the location of the phone number dialed by the real-time communication app. For example, if the caller selects Chinese, the real-time translation module determines that the language of the first transliterated segment sent to the second terminal is Chinese, based on the location information determined from the called party's phone number.

[0015] Reference Figure 1 The process of the first terminal calling the second terminal and conducting a real-time translation call is as follows: Step S1, Call Initiation and Connection. Specifically, this includes: the first terminal enters or selects the second terminal's mobile phone number in the real-time communication APP and sends a call initiation request to the server; the server's signaling control module creates a call session, assigns a session identifier, and performs routing selection, initiating call connection to the second terminal's mobile phone number through the ITSP network; after the second terminal's native phone application rings and answers, the server completes session establishment.

[0016] Step S2, original audio media uplink and relay. Specifically, this includes: after the call is established, the first terminal collects the original audio of the calling user and forms an original audio media stream, which is then sent to the server via the IP network; the server's media gateway module receives the original audio media stream and performs buffering / encoding / decoding processing.

[0017] Step S3, speech segment detection and real-time translation. Specifically, the server extracts speech segments from the original audio media stream based on speech endpoint detection, and inputs the speech segments into the real-time translation module to perform recognition, translation, and speech synthesis to generate corresponding transliterated segments.

[0018] Step S4, downstream playback of the translated audio. Specifically, the server media gateway module sends the translated audio segment to the voice channel of the second terminal, allowing the second terminal to hear the translated audio in the native phone application; when necessary, the server can also perform mixing / gain control on the downstream media to meet the requirements of interstitial playback and intelligibility.

[0019] Step S5, reverse communication. Specifically, this includes: the voice spoken by the second terminal is transmitted back to the server via the PSTN / ITSP network. The server also performs voice segment detection and real-time translation, and sends the translated audio to the first terminal APP for playback, thereby realizing two-way real-time translation communication.

[0020] like Figure 2 As shown, in one embodiment, a VoIP call translation method is provided, which is applied to the VoIP call translation system described above. The method includes: In step S202, when the server receives the first original audio media stream uploaded by the first terminal, it detects whether the first original audio frame and the first translated audio frame exist simultaneously within the first time overlap window.

[0021] In this embodiment, the first terminal is the calling terminal, and the second terminal is the called terminal. It should be noted that the first original audio media stream consists of multiple first original audio frames arranged in chronological order. The first translated audio frame is a constituent unit that generates a corresponding first translated audio segment based on a speech segment extracted from the first original audio media stream, and the first translated audio segment consists of multiple first translated audio frames arranged in chronological order. For ease of description, in this embodiment, the first original audio queue and the first translated audio queue are collectively referred to as the first time overlap window. In some embodiments, the second original audio queue and the second translated audio queue are collectively referred to as the second time overlap window. The reason is that, in this embodiment, if the first original audio frame and the first translated audio frame exist simultaneously within the first time overlap window, that is, if the first original audio frame and the first translated audio frame exist simultaneously in the first original audio queue and the first translated audio queue, according to the method of the server using four queues to manage the audio data of the real-time communication process between the first terminal and the second terminal with translation function in this embodiment, whether the first original audio frame and the first translated audio frame exist simultaneously within the first time overlap window will be mixed and synthesized to generate a first mixed frame and then sent to the second terminal. Therefore, from the effect, the playback time of the first original audio frame and the first translated audio frame that exist simultaneously within the first time overlap window on the second terminal is overlapping.

[0022] In this embodiment, the server uses four queues to manage the audio data during real-time communication between the first terminal and the second terminal with translation functionality. Specifically, the four queues include a first original audio queue, a first translated audio queue, a second original audio queue, and a second translated audio queue. When the server receives a first original audio frame uploaded by the first terminal, it writes each first original audio frame into the first original audio queue frame by frame according to the chronological order of the received first original audio frames, thus caching the first original audio media stream. When the server generates a corresponding first translated audio segment based on the audio segment extracted from the first original audio media stream, the server writes each first translated audio frame into the first translated audio queue frame by frame according to the playback progress of each first translated audio frame constituting the first translated audio segment, thus caching the first translated audio segment. Similarly, when the server receives a second original audio frame uploaded by the second terminal, it writes each second original audio frame into the second original audio queue frame by frame according to the chronological order of the received second original audio frames, thus caching the second original audio media stream. When the server generates the corresponding second transliteration segment based on the audio segment extracted from the second original audio media stream, the server writes each second transliteration frame into the second transliteration queue frame by frame according to the playback progress of each second transliteration frame constituting the second transliteration segment, thereby realizing the caching of the second transliteration segment.

[0023] On the other hand, the server monitors the status of four queues in real time: the first original sound queue, the first transliteration queue, the second original sound queue, and the second transliteration queue. Once it detects that any queue is not empty, it will take the element at the head of the corresponding queue according to the first-in-first-out rule, perform preset processing, and then forward it to the designated terminal.

[0024] Step S204: If not, the first original audio frame or the first transliterated audio frame existing in the first time overlap window will be sent to the second terminal for playback.

[0025] If so, in step S206, the first original audio frame and the first transliterated audio frame corresponding to the first time overlap window will be mixed and synthesized to generate a first mixed frame, and the first mixed frame will be sent to the second terminal for playback.

[0026] For step S204, in a scenario where the first terminal has just spoken and the content of the speech has not yet been translated, the server detects that only the first original audio queue is not empty. Once the server detects that any queue is not empty and only the first original audio queue is not empty, the preset processing is to take the first original audio frame from the head of the first original audio queue according to the first-in-first-out rule and send it to the second terminal. At this time, the head of the first original audio queue is dequeued, so the number of elements in the first original audio queue is reduced by one.

[0027] In step S206, in another scenario, after a period of time, the content spoken by the first terminal has been translated. The server writes each first translated frame into the first translated queue frame by frame according to the playback progress of each first translated frame in the first translated segment. Therefore, the server detects that only the first original audio queue and the first translated queue are not empty. If the server detects that any queue is not empty and only the first original audio queue and the first translated queue are not empty, the preset processing is to take one first original audio frame from the head of the first original audio queue and one first translated frame from the head of the first translated queue according to the first-in-first-out rule. Then, the first original audio frame and the first translated frame taken from the head of the two queues are mixed and synthesized to generate a first mixed frame, which is then sent to the second terminal for playback. At this time, the head of both the first original audio queue and the first translated queue have been dequeued, so the number of elements in both the first original audio queue and the first translated queue is reduced by one.

[0028] Specifically, the audio mixing and synthesis includes: applying different sound feature control parameters to the first transliterated frame and the first original frame respectively, so that the first sound feature parameter of the first transliterated frame in the first mixed frame and the second sound feature parameter of the first original frame in the first mixed frame satisfy a preset difference condition.

[0029] In step S204, in another scenario, after some time has passed, the number of elements in the first original sound queue is zero, meaning the first original sound queue is empty, while the number of elements in the first transliterated sound queue is not zero, meaning the first transliterated sound queue is not empty. At this point, the server detects that only the first transliterated sound queue is not empty. The corresponding preset processing is to take the first transliterated sound frame from the head of the first transliterated sound queue according to the first-in-first-out rule and send it to the second terminal. At this time, the head of the first transliterated sound queue has been dequeued, so the number of elements in the first transliterated sound queue is reduced by one.

[0030] In summary, the server monitors and manages the first original audio queue and the first transliteration queue to send three forms of audio data to the second terminal: a single first original audio media stream, a first transliteration segment, and a first mixed frame generated by mixing and synthesizing the first original audio frame and the first transliteration frame.

[0031] Compared to existing technologies where the calling party's continuous speaking results in a longer waiting time for the called party to hear the translated audio, leading to poor real-time performance and a less than ideal user experience, this embodiment provides a VoIP call translation method. After the calling party finishes speaking the original audio of the Nth sentence, the method generates the translated audio corresponding to the Nth sentence while the calling party continues speaking the original audio of the N+1th sentence. The translated audio of the Nth sentence is then mixed with the original audio of the N+1th sentence to obtain a mixed media stream (composed of multiple mixed frames arranged chronologically), which is then sent to the called terminal. Because different sound feature control parameters are applied to the first transliterated frame and the first original frame in the mixed media stream, the first sound feature parameter of the first transliterated frame in the first mixed frame and the second sound feature parameter of the first original frame in the first mixed frame meet the preset difference conditions. Therefore, the called party can hear the mixed media stream with the first transliterated segment being more prominent, which allows the called party to hear the transliteration faster and the calling party to speak multiple sentences continuously.

[0032] In one example, different sound feature control parameters are applied to the first transliterated frame and the first original frame, respectively, so that the first sound feature parameter of the first transliterated frame in the first mixed frame and the second sound feature parameter of the first original frame in the first mixed frame satisfy a preset difference condition, specifically including: The server performs amplitude scaling on the first transliterated frame and the first original frame, respectively. The amplitude scaling includes multiplying the sample value of the first transliterated frame by a second gain coefficient G2 and multiplying the sample value of the first original frame by a first gain coefficient G1, satisfying Formula 1: The first mixed frame is then obtained by weighted summation of the two time-aligned scaled frames.

[0033] Amplitude scaling refers to multiplicative scaling of PCM sample values, a common gain control method in audio engineering. G1 and G2 correspond to the gain coefficients of the first original sound frame and the first transliterated sound frame, respectively. For example, if G1 = 0.6, then G2 can be greater than or equal to 1.07 to meet the 5dB difference. In this embodiment, the first original sound frame taken from the head of the first original sound queue and the first transliterated sound frame taken from the head of the first transliterated sound queue are defined as time-aligned because they have the same output timestamp (or are mapped to the same output frame number) after generating the first mixed frame. Weighted summation refers to adding samples one by one to generate a mixed frame, and can be limited / compressed before output. For example, the aligned frames x[n] and z[n] are scaled to obtain x'[n] and z'[n], respectively, and then y[n] = x'[n] + z'[n] is calculated as the first mixed frame and sent.

[0034] In this embodiment, when the first original audio frame and the first translated audio frame are output in overlapping order, the energy of the first translated audio frame is relatively more prominent (not less than a preset difference, such as 5dB), which reduces the energy masking of the first original audio frame on the first translated audio frame and improves the intelligibility of the first translated audio frame. The gain difference ensures that the single-channel playability is maintained after mixing.

[0035] like Figure 3 As shown, in one embodiment, generating the corresponding first transliterated segment based on the speech segment extracted from the first original audio media stream specifically includes: In step S302, the server performs voice endpoint detection based on the back-silence timeout threshold, and when the back-silence timeout event is triggered, it extracts a voice segment from the first original audio media stream.

[0036] Specifically, speech endpoint detection is used to determine the "speech start / end" process in streaming audio, and can be based on VAD, energy threshold, or speech probability. The post-silence timeout threshold is the threshold for the duration of silence after the speech endpoint. A "post-silence timeout event" is triggered when the silence duration is greater than or equal to the threshold, confirming the "end of this speech segment." For example, the threshold is 300ms. The server starts timing the silence after detecting the end of the caller's speech; when the cumulative number of consecutive silence frames reaches 300ms, the post-silence timeout event is triggered, and the first set of original audio frames from the speech start frame to the speech endpoint frame is extracted as the "speech segment."

[0037] In step S304, the server performs recognition and translation on the speech segment and performs speech synthesis to obtain the first transliterated segment.

[0038] Specifically, recognition and translation involve feeding the speech segment into ASR to obtain text, and then feeding it into MT to obtain the translation. Speech synthesis involves feeding the translation into TTS to generate audio, i.e., the first transliterated segment. For example, the speech segment content is the Chinese "I am now at the airport". ASR outputs text -> MT outputs the English "I'm at the airport now." -> TTS generates the English transliterated segment and obtains the first transliterated frame sequence {T_1…T_n}.

[0039] This embodiment provides a controllable and reproducible "first transliteration segment generation trigger point" to avoid having to wait until the caller finishes multiple consecutive sentences before translating; it also avoids premature truncation leading to unstable translation quality. Specifically, this embodiment uses a "post-silence threshold" to divide the continuous media stream into speech segments, enabling the translation module to work immediately after each segment ends and output the first transliteration frame sequence, improving translation real-time performance. It should be noted that performing speech endpoint detection based on the post-silence timeout threshold, and extracting the speech segment from the first original audio media stream when the post-silence timeout event is triggered, is existing technology. Technical details can be found in Chinese Patent Publication No. CN114155839A, and will not be repeated here.

[0040] like Figure 4 As shown, in one embodiment, the step of performing recognition and translation on the speech segment and then synthesizing the speech to obtain a first transliterated segment specifically includes: Step S402: The server obtains the first acoustic attribute tag corresponding to the speech segment.

[0041] In one example, the first acoustic attribute label refers to the acoustic category label extracted from the speech segment that reflects "male / female voice type" (it should be noted that the acoustic category label is an acoustic attribute and is not necessarily equivalent to a real identity). For example, the server inputs the speech segment into a pre-trained gender recognition model and obtains a gender recognition result, such as "male". It should be noted that the pre-trained gender recognition model is prior art and will not be elaborated here. For example, the method provided in Chinese Patent Publication No. CN116631436A can be used to train the gender recognition model.

[0042] In step S404, the server selects a target timbre parameter from a preset synthesized timbre parameter library that is inconsistent with the first acoustic attribute label. Each timbre parameter in the synthesized timbre parameter library corresponds to one and only one acoustic attribute label; the acoustic attribute label indicates the attribute corresponding to the timbre parameter, and the attribute corresponding to the timbre parameter includes: male and female.

[0043] The synthesized voice parameter library contains two pre-set sets of TTS voice parameters, each bound to an acoustic attribute label, such as male or female. Given that the first acoustic attribute label corresponding to a speech segment is "male," inconsistent selection means that if the first acoustic attribute label corresponding to a speech segment is "male," then the target voice parameter with the acoustic attribute label "female" is selected; and vice versa. For example, if the first acoustic attribute label corresponding to a speech segment is determined to be male, then the target voice parameter with the acoustic attribute label "female" is selected. Specifically, the statement that each voice parameter in the synthesized voice parameter library corresponds to exactly one acoustic attribute label means that each voice parameter corresponds to only one gender. For example, the voice parameter F_voice corresponds to only one acoustic attribute label, namely "female"; while the voice parameter M_voice corresponds to only one acoustic attribute label, namely "male."

[0044] In step S406, the server performs recognition and translation on the speech segment based on the target timbre parameters and performs speech synthesis to obtain the first transliterated segment.

[0045] Transliteration based on target timbre parameters refers to generating a first transliteration segment with corresponding timbre attributes using target timbre parameters as conditions during the TTS stage. In one example, the translation is an English sentence, and TTS synthesizes it using a female timbre, resulting in the first transliteration frame sequence {T_1…T_n} with a female timbre.

[0046] In this embodiment, when the first original audio frame and the first translated audio frame are mixed, the strong difference in "timbre category" between the first translated audio frame and the first original audio frame makes them easier for listeners to distinguish, thus improving the intelligibility of the translated audio, attention, and diffusing ability. The difference in timbre / acoustic attributes provides auditory diffusing cues, reducing information confusion caused by the overlap of the two voices.

[0047] Reference Figure 5 In one embodiment, the method further includes performing the process of sending a marker signal: Step 502: When the server detects the switch from the first state to the second state for the Nth time, it enters the Nth round of dialogue; where N is a positive integer variable; the first state refers to the state where the first original sound queue, the first transliteration queue, the second original sound queue, and the second transliteration queue are all empty, and the second state refers to the state where only the first original sound queue is not empty.

[0048] For example, after a call is established between the first terminal and the second terminal, when the calling party speaks for the first time, the first terminal captures the caller's initial audio and converts it into a first original audio media stream, which is then uploaded to the server. It can be seen that before the server receives the first original audio media stream for the first time, the first original audio queue, the first transliteration queue, the second original audio queue, and the second transliteration queue on the server are all empty. When the server receives the first original audio media stream for the first time, it temporarily stores the corresponding first original audio frames frame by frame in the first original audio queue, ensuring that only the first original audio queue is not empty. Since this is the first time the server detects a switch from the first state to the second state, the first round of dialogue begins. It should be noted that when the N+1th round of dialogue is detected, the Nth round of dialogue ends. For example, when the second round of dialogue begins, the first round of dialogue ends.

[0049] Step 504: After entering the Nth round of dialogue, the server first sends the first marker signal to the second terminal, and then sends the first original audio frame in the first original audio queue to the second terminal.

[0050] Taking the first round of dialogue as an example, after entering the first round of dialogue, the server first sends the first marker signal to the second terminal, and then sends the first original audio frame from the first original audio queue to the second terminal. In other words, after entering the Nth round of dialogue, the server always sends the first marker signal to the second terminal first, and then starts sending the first original audio frame from the first original audio queue to the second terminal. This allows the server to determine the starting boundary of a segment of the first original audio media stream in the call recording made on the second terminal through the first marker signal.

[0051] Step 506: In the Nth round of dialogue, after the first marker signal is sent to the second terminal, when the server detects a switch from the second state to the first state or the third state, it sends the second marker signal to the second terminal; wherein, the third state refers to a state in which neither the first original sound queue nor the second original sound queue is empty.

[0052] In one example, such as Figure 6 As shown, in the first round of dialogue, the caller needs to speak two sentences consecutively. After the first marker signal is sent to the second terminal, the server detects the switch from the second state to the first state. This indicates that the caller has just finished speaking the first sentence, and the server has detected the pause after speaking the first sentence. Specifically, the server performs voice endpoint detection based on the back-silence timeout threshold. When the back-silence timeout event is triggered, the server extracts the first audio segment from the first original audio media stream. It should be noted that when the first audio segment is extracted from the first original audio media stream... Since the pause duration has reached the post-silence timeout threshold, the server has forwarded all the first original audio frames corresponding to the first audio segment to the second terminal. At this time, since the caller has not yet started speaking the next sentence, the first original audio queue is empty. Therefore, before the audio segment is translated into the corresponding first transliterated segment, the server detects that it is in the first state, that is, the server detects a switch from the second state to the first state. At this time, the server sends a second marker signal to the second terminal. This second marker signal is used to determine the end boundary of a segment of the first original audio media stream. By using the start boundary of a segment of the first original audio media stream determined by the first marker signal provided in step S504, a segment of the first original audio media can be extracted from the call recording.

[0053] In one example, such as Figure 6As shown, in the first round of dialogue, the calling party needs to speak two sentences consecutively. After the first marker signal is sent to the second terminal, the server detects a switch from the second state to the third state, indicating that the called party interrupted while the calling party was speaking the first sentence. This results in both the first and second original audio queues being non-empty. In this case, it means that only the moment before the called party interrupted, the voice heard in the call recording on the second terminal is the voice of the calling party. Therefore, when the server detects a switch from the second state to the third state, it sends the second marker signal to the second terminal. This second marker signal is used to determine the end boundary of a segment of the first original audio media stream. By using the start boundary of the segment of the first original audio media stream determined by the first marker signal provided in step S504, a segment of the first original audio media can be extracted from the call recording.

[0054] Step 508: In the Nth round of dialogue, when the server detects a switch from the first state or the seventh state to the fourth state, it sends a third marker signal to the second terminal; wherein, the fourth state refers to the state in which only the first transliteration queue is not empty; the seventh state refers to the state in which only the first original sound queue and the first transliteration queue are not empty.

[0055] Step 510: When the server detects a switch from the fourth state to the first state during the Nth round of dialogue, it sends the fourth flag signal to the second terminal.

[0056] In one example, such as Figure 6As shown, in the first round of dialogue, the caller needs to speak two sentences consecutively. After the caller finishes speaking the first sentence, there is a pause. The server performs voice endpoint detection based on the after-silence timeout threshold. When the after-silence timeout event is triggered, the server extracts the first audio segment from the first original audio media stream. It should be noted that when the first audio segment is extracted from the first original audio media stream, since the pause duration has reached the after-silence timeout threshold, the server has already forwarded the entire first original audio frame corresponding to the first audio segment to the second terminal. At this point, since the caller has not yet started speaking the next sentence, the first original audio queue is empty. After translation, the first audio segment extracted from the first original audio media stream is translated into the corresponding first transliterated segment and written into the first transliteration queue. At this point, the caller has not yet started speaking the second sentence, so only the first transliteration queue is not empty. That is, the server detects the switch from the first state to the fourth state and sends a third marker signal to the second terminal. This third marker signal is used to determine the starting boundary of a first transliterated segment. However, if the caller starts speaking the second sentence while the first transliterated segment corresponding to the caller's first sentence is being sent to the second terminal for playback, and if the caller's second sentence ends during the process of sending the first transliterated segment corresponding to the caller's first sentence to the second terminal for playback, resulting in a switch from the seventh state to the fourth state, then the server will send another third marker signal to the second terminal at that moment. Two adjacent third marker signals can then be extracted from the call recording on the second terminal. Next, if the first translated segment corresponding to the caller's first sentence is sent to the second terminal and played, at this time the last frame in the first translated segment queue has just been dequeued and forwarded to the second terminal, while the first translated segment corresponding to the caller's second sentence has not yet been generated, the server, upon detecting a switch from the fourth state to the first state, sends the fourth marker signal to the second terminal. Then, the first translated segment corresponding to the caller's second sentence is generated, and the server, upon detecting a switch from the first state to the fourth state, sends the third marker signal to the second terminal. Similarly, if the first translated segment corresponding to the caller's second sentence is sent to the second terminal and played, at this time the last frame in the first translated segment queue has just been dequeued and forwarded to the second terminal, the server, upon detecting a switch from the fourth state to the first state, sends the fourth marker signal to the second terminal. It is evident that in the first round of dialogue, more than one pair of third and fourth marker signals are sent to the second terminal. The signals are sent in the following order: third marker signal, third marker signal, fourth marker signal, third marker signal, and fourth marker signal. The last pair of third and fourth marker signals in the time sequence can be used to determine the start and end positions of the first transliterated segment corresponding to the caller's second sentence.

[0057] Step 512: In the Nth round of dialogue, when the server detects a switch from the first state to the fifth state, it begins monitoring whether the first original audio queue remains empty. If so, when the server detects a switch from the sixth state to the first state while the first original audio queue remains empty, it sends the fifth marker signal to the second terminal. If not, when the server detects that the first original audio queue is no longer empty, it generates a task end signal. The task end signal is used to instruct the server not to send any marker signals to the second terminal in the Nth round of dialogue. The fifth state refers to the state where only the second original audio queue is not empty; the sixth state refers to the state where only the second transliteration queue is not empty.

[0058] In one example, such as Figure 6 As shown, in the first round of dialogue, the calling party has already spoken two sentences, and the corresponding first-translated audio segments have been forwarded to the second terminal for playback. The called party has not yet spoken, therefore, the first original audio queue, the first translated audio queue, the second original audio queue, and the second translated audio queue are all empty at this point. Next, the called party speaks through the second terminal, and their speech is uploaded to the server via the second original audio media stream. The server writes the second original audio frames constituting the second original audio media stream into the second original audio queue, at which point only the second original audio queue is not empty. When the server detects that only the second original audio queue is not empty, it immediately forwards the second original audio frames in the second original audio queue to the first terminal according to the first-in-first-out rule. At this point, in the first round of dialogue, when the server detects a switch from the first state to the fifth state, it begins monitoring whether the first original audio queue remains empty. If the first original audio queue remains empty, it means that the calling party did not interrupt while the called party was speaking, and therefore, the call recording on the second terminal only contains the called party's speech.

[0059] Specifically, after receiving the second original audio media stream, the server writes the second original audio frames constituting the second original audio media stream into the second original audio queue. Based on a back-silence timeout threshold, it performs voice endpoint detection. When a back-silence timeout event is triggered, it extracts a speech segment from the second original audio media stream, translates the speech segment, outputs the corresponding second transliterated segment, and writes multiple second transliterated frames constituting the second transliterated segment into the second transliterated queue. When the server detects that only the second transliterated queue is not empty, it immediately forwards the second transliterated frames in the second transliterated queue to the first terminal according to a first-in, first-out rule. It should be noted that at the time of the back-silence timeout event, all second original audio frames in the second original audio queue have already been forwarded to the first terminal.

[0060] In this example, if the called party only speaks one sentence in the first round of dialogue, and the calling party does not interrupt during the called party's speech, then only the second transliteration queue is not empty during the process of the second transliteration frame corresponding to the called party's speech being dequeued from the second transliteration queue and forwarded to the first terminal. Therefore, after the last second transliteration frame is dequeued from the second transliteration queue and forwarded to the first terminal, when the server detects that the first original audio queue is continuously empty (i.e., the first original audio queue is continuously empty during the process of the server forwarding the second original audio frame and the second transliteration frame to the first terminal), and detects the switch from the sixth state to the first state, it sends the fifth flag signal to the second terminal.

[0061] In another example, the called party only spoke one sentence in the first round of the conversation, but the calling party interrupted during the called party's speech. For example, while the server was forwarding the second original audio frame or the second translated audio frame to the first terminal, the first original audio queue was no longer empty. Therefore, upon detecting that the first original audio queue was no longer empty, a task end signal was generated. When the server detected the generation of the task end signal, it would not send any marker signals to the second terminal while in the first round of the conversation. Therefore, in this example, the fifth marker signal was not sent to the second terminal in the first round of the conversation.

[0062] In another example, for instance, the called party speaks two sentences in the second round of dialogue. Only after the called party has finished speaking the first sentence and all the corresponding transliterations have been forwarded to the first terminal for playback does the called party begin speaking the second sentence. While the called party is speaking the second sentence, the calling party interrupts. It can be seen that in this scenario, after the last transliteration frame corresponding to the first sentence is dequeued from the second transliteration queue and forwarded to the first terminal, the server sends a fifth marker signal to the second terminal. Furthermore, a task completion signal is generated when the calling party interrupts during the called party's second sentence. In this scenario, although a fifth marker signal is sent to the second terminal, a task end signal is also sent. This causes the server to determine that the current round of dialogue has not successfully sent all valid marker signals to the second terminal. This is because, in this scenario, the second terminal receives the fifth marker signal, and the called party may just be speaking the second sentence. This would cause the fifth marker signal on the second terminal to overlap with part of the second sentence's sound, which may make it impossible to effectively extract the fifth marker later. Therefore, to avoid this problem, this embodiment determines that the current round of dialogue has not successfully sent all valid marker signals to the second terminal, thereby triggering the process of sending marker signals to continue in the next round of dialogue until all valid marker signals are successfully sent to the second terminal in a certain round of dialogue.

[0063] Understandably, if the caller begins speaking the second sentence after uttering the first sentence but before the corresponding first transliterated segment is written into the first transliteration queue, it's equivalent to the server detecting a second transition from the first state to the second state, thus entering the second round of dialogue. It should be noted that due to the uncertainty of the caller and callee's speaking times during real-time communication, it cannot be guaranteed that all five marker signals (first, second, third, fourth, and fifth) will be sent to the second terminal in a single round. Therefore, if the fifth marker signal cannot be sent to the second terminal in a single round, the process of sending marker signals will continue. Specifically, in this embodiment, upon detecting the entry into the Nth round of dialogue, in the Nth round, when the conditions are met, the first marker signal is sent to the second terminal first, followed by the second marker signal, then the third and fourth marker signals, and finally the fifth marker signal. That is, the server can determine whether the task of sending the first to fifth marker signals in this round of dialogue has been successfully completed by detecting the order in which the marker signals are sent. Specifically, if the fifth marker signal is successfully sent to the second terminal in the Nth round of the dialogue without generating a task end signal, it means that all valid marker signals (i.e., valid first, second, third, fourth, and fifth marker signals) have been successfully sent to the second terminal in this round. Therefore, it is unnecessary to send the first to fifth marker signals to the second terminal in the N+1th round. This means that after the N+1th round, the server no longer needs to continuously monitor the status of the four queues, and the corresponding computer resources can be released. Conversely, if the server detects that the fifth marker signal has not been sent to the second terminal or a task end signal has been generated in the Nth round, it determines that all valid marker signals have not been successfully sent to the second terminal in this round. This will trigger the process of sending marker signals to continue in the N+1th round until all valid marker signals are successfully sent to the second terminal in some round.

[0064] It should be noted that if the requirement is to successfully send all valid marker signals to the second terminal in two or more rounds of dialogue, the corresponding computer resources should only be released after the task of successfully sending all valid marker signals to the second terminal in two or more rounds of dialogue is completed. For example, in one example, the task of successfully sending all valid marker signals to the second terminal in two rounds of dialogue can be required. For instance, after a preset time period following the first successful sending of all valid marker signals to the second terminal, such as 3 minutes, the process of sending marker signals can be executed again until the second successful sending of all valid marker signals to the second terminal is achieved. By executing the task of sending all valid marker signals to the second terminal in two separate executions at preset time intervals, it is possible to ensure that all valid marker signals appear in the call recording even if the second terminal does not start recording the call from the beginning.

[0065] Since the called party in this application answers the call through a native phone app on a second terminal, and the audio received contains a mixture of the first original audio frame and the first translated audio frame, the called party, in order to preserve evidence of the call and address potential disputes, will utilize the call recording function provided by the native phone app on the second terminal, or the recording or screen recording function built into the second terminal. For the called party on the first terminal, since they are using a specific real-time communication app, and the data such as the first original audio media stream, the first translated audio segment, the second original audio media stream, and the second translated audio segment from both the first and second terminals can be individually obtained by the server, the called party can directly request the corresponding data from the server through the real-time communication app when they need each part of the audio content. To subsequently separate the first original audio frame and the first translated audio frame from the mixed media stream, a target speaker extraction task or directional speech separation technology can be used. The target speaker extraction task is a mature existing technology and will not be elaborated here; for specific details, please refer to Chinese Patent Publication No. CN117809675A. The improvement of this embodiment is that, during the call between the first terminal and the second terminal, a process of sending marker signals is executed to send a first marker signal, a second marker signal, a third marker signal, a fourth marker signal, and a fifth marker signal to the second terminal, which can automatically and efficiently extract reference audio for the target speaker extraction task from the subsequent call recording of the second terminal.

[0066] It should be noted that determining the boundary of the second original audio media stream corresponding to the called party in the call recording on the second terminal is a technical challenge. This is because determining the start position of the second original audio media stream requires prior knowledge of when the called party will speak. Obviously, the server only knows the called party has spoken after receiving the second original audio media stream. However, if a marker signal is sent to the second terminal as the start position immediately after receiving the second original audio media stream, the marker signal may overlap with the called party's speech, preventing subsequent early recognition. Similarly, determining the end position of the second original audio media stream requires waiting for the called party to finish speaking. However, even if the marker signal for the end position is sent to the second terminal during a brief pause, it cannot be guaranteed that the called party will have just finished speaking when the second terminal receives the marker signal. To address this technical difficulty, this embodiment uses the concept of entering the Nth round of dialogue, a method for determining the boundary of the second original audio media stream corresponding to the called party from the perspective of one round of dialogue. In a single dialogue round, the calling party will inevitably speak first, followed by the called party. By defining a single dialogue round and monitoring the four queues, the timing of the calling party's speech ending (i.e., the moment corresponding to the last fourth marker signal sent in a single dialogue round) can be accurately determined. Therefore, after the fourth marker signal, the called party will inevitably begin speaking, thus determining the starting position of the second original audio media stream corresponding to the called party. Furthermore, the system monitors whether the first original audio queue is continuously empty. If so, when the first original audio queue is continuously empty, and a switch from the sixth state to the first state is detected, the fifth marker signal is sent to the second terminal, thereby determining the boundary of the end position of the second original audio media stream. In this embodiment, the boundary between the starting and ending positions of the second original audio media stream includes the called party's voice and any silence.

[0067] In one example, the first, second, third, fourth, and fifth marker signals are all audio segments carrying marker information. For example, the length of each audio segment is 60ms. In order to distinguish different marker signals, different marker signals use different DTMF (Dual-Tone Multi-Frequency) key values, which is convenient for subsequent identification in call recordings.

[0068] In one embodiment, a method is also provided for automatically separating a first original audio media stream, a first transliterated segment, and a second original audio media stream from a call recording. This method is operable on a computer device, including a second terminal or a PC. The method further includes: Step S602: Obtain the call recording from the second terminal.

[0069] Step S604: Extract the marker signal sequence from the call recording; wherein, the elements of the marker signal sequence include the number of each marker signal and the playback progress of each marker signal in the call recording, and the elements in the marker signal sequence are sorted according to the order of the playback progress of the corresponding marker signal in the call recording.

[0070] In one example, the extracted marker signal sequence is [(1,t_1),(2,t_2),(3,t_3),(3,t_4),(4,t_5),(3,t_6),(4,t_7),(5,t_8)]. Taking the element (1,t_1) of the marker signal sequence as an example, 1 represents the marker signal number, and t_1 represents the playback progress, i.e., the timestamp in the call recording. For example, the DTMF key value of the marker signal numbered 1 is "A"; the DTMF key value of the marker signal numbered 2 is "B"; the DTMF key value of the marker signal numbered 3 is "C"; the DTMF key value of the marker signal numbered 4 is "D"; and the DTMF key value of the marker signal numbered 5 is "#".

[0071] Specifically, detecting the aforementioned marker signals in a call recording essentially involves detecting the corresponding DTMF key values ​​for each marker signal. How to detect DTMF key values ​​from audio is a matter of existing technology. For example, the ITU-T Q.23 standard specifies the low-frequency and high-frequency groups of DTMF, as well as the mapping relationship between key pairs and frequencies. As long as the corresponding two-frequency combination is detected in the audio, the key value can be determined. Specifically, during extraction, a sliding window analysis is performed on the call recording to detect whether the preset frequency points of the low-frequency and high-frequency groups exist simultaneously and meet a duration threshold (e.g., 60ms), thereby decoding the corresponding DTMF key value and its timestamp.

[0072] Step S606: Extract the first playback progress corresponding to the target first marker signal and the second playback progress corresponding to the second signal marker adjacent to the target first marker signal from the marker signal sequence; wherein, the target first marker signal is any one of the first marker signals in the marker signal sequence.

[0073] Step S608: Extract the audio data between the first playback progress corresponding to the target first marker signal and the second playback progress corresponding to the second signal marker adjacent to the target first marker signal from the call recording, and use it as the first reference audio.

[0074] In this example, the extracted marker signal sequence is [(1,t_1),(2,t_2),(3,t_3),(3,t_4),(4,t_5),(3,t_6),(4,t_7),(5,t_8)]. The first playback progress corresponding to the target first marker signal is t_1, and the second playback progress corresponding to the second marker adjacent to the target first marker signal is t_2. Therefore, the playback range of the first reference audio in the call recording is (t_1+△t, t_2). The audio data with a playback range of (t_1+△t, t_2) in the call recording is used as the first reference audio, where △t is the length of the audio segment corresponding to each marker signal, and the unit of △t is the same as t_1, which is milliseconds.

[0075] Step S610: Input the feature representation of the call recording and the feature representation of the first reference audio into the pre-trained target speaker speech separation model to obtain the first target speaker audio belonging to the first target speaker in the call recording; wherein, the first reference audio is used to characterize the speech features of the first target speaker.

[0076] This embodiment can automatically extract the caller's first original audio media stream from call recordings, improving the efficiency of separating the target speaker from call recordings.

[0077] In one embodiment, a method is also provided for automatically separating a first original audio media stream, a first transliterated segment, and a second original audio media stream from a call recording. This method is operable on a computer device, including a second terminal or a PC. The method further includes: Step S702: Obtain the call recording from the second terminal.

[0078] Step S704: Extract the marker signal sequence from the call recording; wherein, the elements of the marker signal sequence include the number of each marker signal and the playback progress of each marker signal in the call recording, and the elements in the marker signal sequence are sorted according to the order of the playback progress of the corresponding marker signal in the call recording.

[0079] In one example, the extracted marker signal sequence is [(1,t_1),(2,t_2),(3,t_3),(3,t_4),(4,t_5),(3,t_6),(4,t_7),(5,t_8)]. Taking the element (1,t_1) of the marker signal sequence as an example, 1 represents the marker signal number, and t_1 represents the playback progress, i.e., the timestamp in the call recording. For example, the DTMF key value of the marker signal numbered 1 is "A"; the DTMF key value of the marker signal numbered 2 is "B"; the DTMF key value of the marker signal numbered 3 is "C"; the DTMF key value of the marker signal numbered 4 is "D"; and the DTMF key value of the marker signal numbered 5 is "#".

[0080] Specifically, detecting the aforementioned marker signals in a call recording essentially involves detecting the corresponding DTMF key values ​​for each marker signal. How to detect DTMF key values ​​from audio is a matter of existing technology. For example, the ITU-T Q.23 standard specifies the low-frequency and high-frequency groups of DTMF, as well as the mapping relationship between key pairs and frequencies. As long as the corresponding two-frequency combination is detected in the audio, the key value can be determined. Specifically, during extraction, a sliding window analysis is performed on the call recording to detect whether the preset frequency points of the low-frequency and high-frequency groups exist simultaneously and meet a duration threshold (e.g., 60ms), thereby decoding the corresponding DTMF key value and its timestamp.

[0081] Step S706: Extract the fourth playback progress corresponding to the target fourth marker signal and the third playback progress corresponding to the third signal marker adjacent to the target fourth marker signal from the marker signal sequence; wherein, the target fourth marker signal is any fourth marker signal in the marker signal sequence that is adjacent to the fifth marker signal.

[0082] Step S708: Extract the audio data between the fourth playback progress corresponding to the target fourth marker signal and the third playback progress corresponding to the third signal marker adjacent to the target fourth marker signal from the call recording, and use it as the second reference audio.

[0083] In this example, the extracted marker signal sequence is [(1,t_1),(2,t_2),(3,t_3),(3,t_4),(4,t_5),(3,t_6),(4,t_7),(5,t_8)]. The fourth playback progress corresponding to the target fourth marker signal is t_7, and the third playback progress corresponding to the third signal marker adjacent to the target fourth marker signal is t_6. Therefore, the playback range of the second reference audio in the call recording is (t_6+△t, t_7). The audio data with a playback range of (t_6+△t, t_7) in the call recording is used as the second reference audio, where △t is the length of the audio segment corresponding to each marker signal, and the unit of △t is the same as t_6, which is milliseconds.

[0084] Step S710: Input the feature representation of the call recording and the feature representation of the second reference audio into the pre-trained target speaker speech separation model to obtain the second target speaker audio belonging to the second target speaker in the call recording; wherein, the second reference audio is used to characterize the speech features of the second target speaker.

[0085] This embodiment can automatically extract the caller's first original audio media stream from call recordings, improving the efficiency of separating the target speaker from call recordings.

[0086] In one embodiment, a method is also provided for automatically separating a first original audio media stream, a first transliterated segment, and a second original audio media stream from a call recording. This method is operable on a computer device, including a second terminal or a PC. The method further includes: Step S802: Obtain the call recording from the second terminal.

[0087] Step S804: Extract the marker signal sequence from the call recording; wherein, the elements of the marker signal sequence include the number of each marker signal and the playback progress of each marker signal in the call recording, and the elements in the marker signal sequence are sorted according to the order of the playback progress of the corresponding marker signal in the call recording.

[0088] In one example, the extracted marker signal sequence is [(1,t_1),(2,t_2),(3,t_3),(3,t_4),(4,t_5),(3,t_6),(4,t_7),(5,t_8)]. Taking the element (1,t_1) of the marker signal sequence as an example, 1 represents the marker signal number, and t_1 represents the playback progress, i.e., the timestamp in the call recording. For example, the DTMF key value of the marker signal numbered 1 is "A"; the DTMF key value of the marker signal numbered 2 is "B"; the DTMF key value of the marker signal numbered 3 is "C"; the DTMF key value of the marker signal numbered 4 is "D"; and the DTMF key value of the marker signal numbered 5 is "#".

[0089] Specifically, detecting the aforementioned marker signals in a call recording essentially involves detecting the corresponding DTMF key values ​​for each marker signal. How to detect DTMF key values ​​from audio is a matter of existing technology. For example, the ITU-T Q.23 standard specifies the low-frequency and high-frequency groups of DTMF, as well as the mapping relationship between key pairs and frequencies. As long as the corresponding two-frequency combination is detected in the audio, the key value can be determined. Specifically, during extraction, a sliding window analysis is performed on the call recording to detect whether the preset frequency points of the low-frequency and high-frequency groups exist simultaneously and meet a duration threshold (e.g., 60ms), thereby decoding the corresponding DTMF key value and its timestamp.

[0090] Step S806: Extract the fifth playback progress corresponding to the target fifth marker signal and the fourth playback progress corresponding to the fourth signal marker adjacent to the target fifth marker signal from the marker signal sequence; wherein, the target fifth marker signal is any fifth marker signal in the marker signal sequence.

[0091] Step S808: Extract the audio data between the fifth playback progress corresponding to the target fifth marker signal and the fourth playback progress corresponding to the fourth signal marker adjacent to the target fifth marker signal from the call recording, and use it as the third reference audio.

[0092] In this example, the extracted marker signal sequence is [(1,t_1),(2,t_2),(3,t_3),(3,t_4),(4,t_5),(3,t_6),(4,t_7),(5,t_8)]. The fifth playback progress corresponding to the target fifth marker signal is t_8, and the fourth playback progress corresponding to the fourth signal marker adjacent to the target fifth marker signal is t_7. Therefore, the playback range of the third reference audio in the call recording is (t_7+△t, t_8). The audio data with a playback range of (t_7+△t, t_8) in the call recording is used as the third reference audio, where △t is the length of the audio segment corresponding to each marker signal, and the unit of △t is the same as t_7, which is milliseconds.

[0093] The third reference audio obtained in this embodiment contains a lot of silent parts. Therefore, it is necessary to divide the third reference audio into frames and calculate the frame energy, filter out silent frame segments with energy below a preset threshold and duration not less than a preset duration, and splice the remaining non-silent frame segments to obtain the de-silent reference audio.

[0094] Step S810: Input the feature representation of the call recording and the feature representation of the third reference audio into the pre-trained target speaker speech separation model to obtain the third target speaker audio belonging to the third target speaker in the call recording; wherein, the third reference audio is used to characterize the speech features of the third target speaker.

[0095] This embodiment can automatically extract the caller's first original audio media stream from call recordings, improving the efficiency of separating the target speaker from call recordings.

[0096] In one embodiment, the durations of the first target speaker audio, the second target speaker audio, and the third target speaker audio are all aligned with the call recording on the same timeline, and their total durations are consistent with the total duration of the call recording. The method further includes: The audio of the first target speaker is mixed and synthesized with the audio of the third speaker to generate the first synthesized call audio.

[0097] The second target speaker's audio is mixed and synthesized with the third speaker's audio to generate a second synthesized call audio.

[0098] In this embodiment, being aligned on the same timeline and having the same total duration means that the three target speaker audios and the call recording have the same start and end points (e.g., all T seconds). Blank audio (with a sample value of 0 or low energy noise) is output during the time slice when the target speaker is not speaking, thus ensuring that the calculation can be performed sample by sample. For example, if the call recording is 180 seconds long, then the extracted first, second, and third target speaker audios are also all 180 seconds long.

[0099] Specifically, the first synthesized call audio D1[n] is the synthesized result obtained by adding the audio of the first target speaker and the audio of the third target speaker according to sample alignment. For example, D1[n] = α·A[n] + β·C[n], where A[n] is the audio of the first target speaker and C[n] is the audio of the third target speaker; α and β can be 1 or adjustable weight coefficients. Specifically, the second synthesized call audio D2[n] is the synthesized result obtained by mixing and synthesizing the audio of the second target speaker and the audio of the third target speaker. For example, D2[n] = α'·B[n] + β'·C[n], where B[n] is the audio of the second target speaker; α' and β' can be 1 or adjustable weight coefficients.

[0100] This embodiment transforms "dialogue structure restoration" into "linear mixing operation of coaxial audio tracks"; equal-length alignment ensures that sample-by-sample synthesis is legal and stable. Since the audio of the three target speakers is aligned with the call recording of equal length, the synthesis process can obtain different versions of call audio (e.g., the call version of "first target + third target", the call version of "second target + third target") without the need for complex time reconstruction, which is convenient for evidence preservation, review, quality inspection or training data construction.

[0101] In one embodiment, a VoIP call translation system is also provided, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to perform the steps of the VoIP call translation method described above. The steps of the VoIP call translation method described above may be the steps from the various embodiments above.

[0102] In one embodiment, a computer-readable storage medium is also provided, storing computer-executable instructions for causing a computer to perform the steps of the VoIP call translation method described above. The steps of the VoIP call translation method here can be the steps in the VoIP call translation methods of the various embodiments described above.

[0103] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0104] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0105] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain.

[0106] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRA), direct RAM via Rambus (RDRA), direct memory bus dynamic RAM (DRDRAM), and dynamic RAM via Rambus (RDRAM), etc.

[0107] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

Claims

1. A method for translating VoIP calls, characterized in that, The method includes: When the server receives the first original audio media stream uploaded by the first terminal, it detects whether the first original audio frame and the first transliterated audio frame exist simultaneously within the first time overlap window. If not, the first original audio frame or the first transliterated audio frame existing in the first time overlap window will be sent to the second terminal for playback; If so, the first original audio frame and the first transliterated audio frame corresponding to the first time overlap window will be mixed and synthesized to generate a first mixed frame, and the first mixed frame will be sent to the second terminal for playback; Wherein, the first original audio media stream is composed of a plurality of first original audio frames arranged in chronological order; the first transliterated frame is a constituent unit for generating a corresponding first transliterated segment based on the speech segment extracted from the first original audio media stream, and the first transliterated segment is composed of a plurality of first transliterated frames arranged in chronological order. The audio mixing process includes applying different sound feature control parameters to the first transliterated frame and the first original frame, such that the first sound feature parameter of the first transliterated frame in the first mixed frame and the second sound feature parameter of the first original frame in the first mixed frame satisfy a preset difference condition.

2. The method for translating VoIP calls according to claim 1, characterized in that, The process of generating the corresponding first transliterated segment based on the speech segment extracted from the first original audio media stream specifically includes: The server performs voice endpoint detection based on the back-silence timeout threshold, and extracts the voice segment from the first original audio media stream when the back-silence timeout event is triggered. The server performs recognition and translation on the speech segment and synthesizes the speech to obtain the first transliterated segment.

3. The VoIP call translation method according to claim 2, characterized in that, The step of recognizing and translating the speech segment and then synthesizing it to obtain the first transliterated segment specifically includes: The server obtains the first acoustic attribute label corresponding to the speech segment; The server selects a target timbre parameter from a preset synthesized timbre parameter library that is inconsistent with the first acoustic attribute label; wherein, each timbre parameter in the synthesized timbre parameter library corresponds to one and only one acoustic attribute label; the acoustic attribute label is used to indicate the attribute corresponding to the timbre parameter, and the attribute corresponding to the timbre parameter includes: male and female; The server performs recognition and translation on the speech segment based on the target timbre parameters and then performs speech synthesis to obtain the first transliterated segment.

4. A VoIP call translation method according to claim 1 or 3, characterized in that, The step of applying different sound feature control parameters to the first transliterated frame and the first original frame, respectively, so that the first sound feature parameter of the first transliterated frame in the first mixed frame and the second sound feature parameter of the first original frame in the first mixed frame satisfy a preset difference condition, specifically includes: The server performs amplitude scaling on the first transliterated frame and the first original frame, respectively. The amplitude scaling includes multiplying the sample value of the first transliterated frame by a second gain coefficient G2 and multiplying the sample value of the first original frame by a first gain coefficient G1, satisfying Formula 1: The first mixed frame is then obtained by weighted summation of the two time-aligned scaled frames.

5. The method for translating VoIP calls according to claim 1, characterized in that, The method further includes: When the server detects a switch from the first state to the second state for the Nth time, it enters the Nth round of dialogue; where N is a positive integer variable; the first state refers to the state where the first original sound queue, the first transliteration queue, the second original sound queue, and the second transliteration queue are all empty, and the second state refers to the state where only the first original sound queue is not empty; After entering the Nth round of dialogue, the server first sends the first marker signal to the second terminal, and then sends the first original audio frame in the first original audio queue to the second terminal. In the Nth round of dialogue, after the first marker signal is sent to the second terminal, when the server detects a switch from the second state to the first state or the third state, it sends the second marker signal to the second terminal; wherein, the third state refers to a state in which neither the first original sound queue nor the second original sound queue is empty; In the Nth round of dialogue, when the server detects a switch from the first state to the fourth state, it sends a third flag signal to the second terminal; where the fourth state refers to the state where only the first transliteration queue is not empty. In the Nth round of dialogue, when the server detects a switch from the fourth state to the first state, it sends the fourth flag signal to the second terminal. In the Nth round of dialogue, when the server detects a switch from the first state to the fifth state, it begins monitoring whether the first original audio queue remains empty. If so, when the server detects a switch from the sixth state to the first state while the first original audio queue remains empty, it sends the fifth marker signal to the second terminal. If not, when the server detects that the first original audio queue is no longer empty, it generates a task end signal. The task end signal is used to instruct the server not to send any marker signals to the second terminal in the Nth round of dialogue. The fifth state refers to the state where only the second original audio queue is not empty; the sixth state refers to the state where only the second transliteration queue is not empty.

6. The method for translating VoIP calls according to claim 5, characterized in that, The method further includes: Obtain call recordings from the second terminal; A sequence of marker signals is extracted from the call recording; wherein the elements of the marker signal sequence include the number of each marker signal and the playback progress of each marker signal in the call recording, and the elements in the marker signal sequence are sorted according to the order of the playback progress of the corresponding marker signals in the call recording. Extract the first playback progress corresponding to the target first marker signal and the second playback progress corresponding to the second signal marker adjacent to the target first marker signal from the marker signal sequence; wherein, the target first marker signal is any one of the first marker signals in the marker signal sequence; The audio data between the first playback progress corresponding to the first target marker signal and the second playback progress corresponding to the second signal marker adjacent to the first target marker signal is extracted from the call recording and used as the first reference audio. The feature representation of the call recording and the feature representation of the first reference audio are input into a pre-trained target speaker speech separation model to obtain the first target speaker audio belonging to the first target speaker in the call recording; wherein, the first reference audio is used to characterize the speech features of the first target speaker.

7. The VoIP call translation method according to claim 5, characterized in that, The method further includes: Obtain call recordings from the second terminal; A sequence of marker signals is extracted from the call recording; wherein the elements of the marker signal sequence include the number of each marker signal and the playback progress of each marker signal in the call recording, and the elements in the marker signal sequence are sorted according to the order of the playback progress of the corresponding marker signals in the call recording. Extract the fourth playback progress corresponding to the target fourth marker signal and the third playback progress corresponding to the third signal marker adjacent to the target fourth marker signal from the marker signal sequence; wherein, the target fourth marker signal is any fourth marker signal in the marker signal sequence that is adjacent to the fifth marker signal; The audio data between the fourth playback progress corresponding to the target fourth marker signal and the third playback progress corresponding to the third signal marker adjacent to the target fourth marker signal is extracted from the call recording and used as the second reference audio; The feature representation of the call recording and the feature representation of the second reference audio are input into a pre-trained target speaker speech separation model to obtain the second target speaker audio belonging to the second target speaker in the call recording; wherein, the second reference audio is used to characterize the speech features of the second target speaker.

8. The method for translating VoIP calls according to claim 5, characterized in that, The method further includes: Obtain call recordings from the second terminal; A sequence of marker signals is extracted from the call recording; wherein the elements of the marker signal sequence include the number of each marker signal and the playback progress of each marker signal in the call recording, and the elements in the marker signal sequence are sorted according to the order of the playback progress of the corresponding marker signals in the call recording. Extract the fifth playback progress corresponding to the target fifth marker signal and the fourth playback progress corresponding to the fourth signal marker adjacent to the target fifth marker signal from the marker signal sequence; wherein, the target fifth marker signal is any fifth marker signal in the marker signal sequence; The audio data between the fifth playback progress corresponding to the target fifth marker signal and the fourth playback progress corresponding to the fourth signal marker adjacent to the target fifth marker signal is extracted from the call recording and used as the third reference audio; The feature representations of the call recording and the third reference audio are input into a pre-trained target speaker speech separation model to obtain the third target speaker audio belonging to the third target speaker in the call recording; wherein, the third reference audio is used to characterize the speech features of the third target speaker.

9. The method for translating VoIP calls according to claim 1, characterized in that, The durations of the first target speaker audio, the second target speaker audio, and the third target speaker audio are all aligned with the call recording on the same timeline, and their total durations are consistent with the total duration of the call recording. The method further includes: The audio of the first target speaker is mixed and synthesized with the audio of the third speaker to generate the first synthesized call audio; The second target speaker's audio is mixed and synthesized with the third speaker's audio to generate a second synthesized call audio.

10. A VoIP call translation system, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, when the processor executes the program, it performs the steps of a VoIP call translation method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Voice endpoint detection method and device, equipment and storage medium

    CN114155839A

  • Gender recognition model processing method and device, computer equipment and storage medium

    CN116631436A

  • Directional voice separation method and system for target speaker in conference scene

    CN117809675A