Simultaneous interpretation delay cancellation method, device, equipment, medium and program product
By separating and masking the original speech of the target speaker from the mixed audio, combined with background audio fusion and user selection functions, the latency problem in machine simultaneous interpretation systems is solved, improving user experience and translation quality, and is suitable for multi-person conversation scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- IFLYTEK CO LTD
- Filing Date
- 2025-12-29
- Publication Date
- 2026-04-17
AI Technical Summary
Existing machine simultaneous interpretation systems suffer from latency, resulting in a poor user experience. In particular, listeners can perceive the time difference between the translated speech and the original audio, which disrupts continuity and immersion.
By separating the original speech of the target speaker from the mixed audio, masking its background audio, and fusing the target speech with the background audio, the delay perception anchor is eliminated, achieving delay pseudo-cancellation. At the same time, the speaker information set is displayed to give users the right to actively choose.
It improves the user experience of simultaneous interpretation, enhances the sense of presence and immersion, improves the accuracy and applicability of translation, solves the problem of voice blocking in multi-person conversation scenarios, and improves the accuracy and convenience of users in selecting target speakers.
Smart Images

Figure CN121438857B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of simultaneous interpreting technology, and more particularly to a method, apparatus, device, medium, and program product for delay elimination in simultaneous interpreting. Background Technology
[0002] Simultaneous interpreting is a crucial technology in scenarios such as international conferences, cross-language live broadcasts, and multilingual communication. Ideal simultaneous interpreting requires the translated text to be conveyed to the audience in real time and fluently, with virtually no perceptible delay. In practice, professional human interpreters can typically control the delay to within a few seconds, providing the audience with a coherent auditory experience.
[0003] In recent years, with the development of deep learning technology, machine simultaneous interpretation systems have made significant progress. However, these systems generally face a trade-off between accuracy and latency. To ensure the accuracy and fluency of the translation, machine simultaneous interpretation systems need to recognize, understand, translate, and synthesize the mixed audio input. This series of computational processes inevitably introduces inherent system latency; for example, it may take several seconds to output a fluent and accurate translation, leading to a degraded user experience. Therefore, eliminating or reducing latency in simultaneous interpretation is a pressing technical requirement.
[0004] Currently, at the algorithm level, the system is evolving from traditional cascaded systems to end-to-end models to reduce intermediate steps and lower inference latency. At the hardware level, processing speed is being improved by using high-performance computing chips and optimizing model deployment. However, existing technologies can only continuously reduce latency, not eliminate it completely. When secondary sounds (such as translated speech) arrive more than 35ms later than the original sound, they are perceived by the listener as independent echoes, thus disrupting the continuous auditory impression. In other words, listeners can perceive this delay by comparing the original and translated speech, leading to a poor user experience. Summary of the Invention
[0005] This invention provides a method, apparatus, device, medium, and program product for eliminating delays in simultaneous interpreting, in order to solve the defect in the prior art where simultaneous interpreting delays cannot be completely eliminated, resulting in a poor user experience, and to achieve pseudo-delay elimination, thereby improving the user experience of simultaneous interpreting.
[0006] This invention provides a delay cancellation method for simultaneous interpreting, comprising:
[0007] Acquire mixed audio including the original speech of at least one speaker;
[0008] The original speech of the target speaker is separated from the mixed audio to obtain background audio that masks the original speech of the target speaker; the target speaker is one of the at least one speakers, and the original speech of the target speaker is the speech to be interpreted simultaneously;
[0009] Simultaneous interpretation is performed on the original speech of the target speaker to obtain the target speech of the target speaker.
[0010] The target language speech is fused with the background audio to obtain the output audio;
[0011] The target speaker is determined based on the following method:
[0012] Display speaker information set; the speaker information set includes the text content corresponding to the original speech of the at least one speaker;
[0013] The target speaker is determined based on the speaker selected by the speaker selection instruction; the speaker selection instruction is an instruction triggered based on the speaker information set.
[0014] According to the simultaneous interpreting delay elimination method provided by the present invention, the speaker information set further includes the identification information of the at least one speaker.
[0015] According to the simultaneous interpreting delay cancellation method provided by the present invention, the identification information of the at least one speaker is determined based on the following method:
[0016] The sound source is located in the mixed audio to obtain the spatial orientation features of the sound source of the at least one speaker, and the voiceprint features are extracted from the mixed audio to obtain the voiceprint features of the at least one speaker.
[0017] The spatial orientation features of the sound source of the at least one speaker and the voiceprint features of the at least one speaker are input into the target association model to obtain the speech feature identification vector of the at least one speaker output by the target association model; the target association model is used to fuse the spatial orientation features of the sound source of the same speaker and the voiceprint features to obtain the speech feature identification vector.
[0018] Based on the speech feature identification vector of the at least one speaker, the identification information of the at least one speaker is determined; the speech feature identification vector of the at least one speaker corresponds one-to-one with the identification information of the at least one speaker.
[0019] According to a delay cancellation method for simultaneous interpreting provided by the present invention, the step of separating the original speech of the target speaker from the mixed audio to obtain background audio that masks the original speech of the target speaker includes:
[0020] Generate a mask matrix for reconstructing the original speech of the target speaker;
[0021] Based on the mask matrix and the mixed audio, speech reconstruction is performed to obtain the original speech of the target speaker;
[0022] The original speech of the target speaker in the mixed audio is masked to obtain the background audio.
[0023] According to a delay cancellation method for simultaneous interpreting provided by the present invention, the step of generating a mask matrix for reconstructing the original speech of the target speaker includes:
[0024] The mixed audio and the speech feature vector of the target speaker are input into the speech separation model to obtain the mask matrix output by the speech separation model; the speech feature vector of the target speaker is obtained by feature fusion based on the sound source spatial orientation features and voiceprint features of the target speaker;
[0025] The speech separation model is obtained by training based on sample mixed audio.
[0026] According to a delay cancellation method for simultaneous interpreting provided by the present invention, the step of simultaneously interpreting the original speech of the target speaker to obtain the target speech of the target speaker includes:
[0027] The original speech of the target speaker is encoded to obtain speech coding features;
[0028] The input features, converted from the speech coding features, are input into the speech feature generation model to obtain the speech features output by the speech feature generation model; the speech feature generation model is used to generate speech features of the target language based on the speech coding features of the original language speech; the input features are features adapted to the speech feature generation model.
[0029] The speech features are decoded to obtain the target language speech of the target language.
[0030] The present invention also provides a delay cancellation device for simultaneous interpretation, comprising:
[0031] An audio acquisition module is used to acquire mixed audio including the original speech of at least one speaker;
[0032] The speech separation module is used to separate the original speech of the target speaker from the mixed audio to obtain background audio that masks the original speech of the target speaker; the target speaker is one of the at least one speakers, and the original speech of the target speaker is the speech to be interpreted simultaneously;
[0033] The simultaneous interpretation module is used to simultaneously interpret the original speech of the target speaker to obtain the target speech of the target speaker.
[0034] An audio fusion module is used to fuse the target language speech with the background audio to obtain output audio;
[0035] The target speaker is determined based on the following modules:
[0036] The information display module is used to display a speaker information set; the speaker information set includes the text content corresponding to the original speech of the at least one speaker;
[0037] The speaker determination module is used to determine the target speaker based on the speaker indicated by the speaker selection instruction; the speaker selection instruction is an instruction triggered based on the speaker information set.
[0038] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the delay cancellation method for simultaneous interpretation as described above.
[0039] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the delay elimination method for simultaneous interpretation as described above.
[0040] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the delay elimination method for simultaneous interpretation as described above.
[0041] The present invention provides a method, apparatus, device, medium, and program product for delay elimination in simultaneous interpreting. It separates the original speech of the target speaker from mixed audio, obtaining background audio that masks the original speech, thereby masking the original speech as an anchor point for delay perception. This directly eliminates the possibility of time comparison for the user, thus completely solving the negative impact of delay at the subjective experience level, ultimately achieving pseudo-delay elimination to improve the user experience of simultaneous interpreting. Furthermore, while masking the original speech of the target speaker, it retains all other background ambient sounds, greatly enhancing the sense of presence and immersion, thereby improving the user experience of simultaneous interpreting. Since translation quality is no longer sacrificed for extremely low latency, simultaneous interpreting of the original speech of the target speaker can yield highly accurate target speech, thus improving the accuracy of simultaneous interpreting and ultimately enhancing the user experience of simultaneous interpreting. Simultaneously, it combines the target speech with… Background audio is fused to obtain the output audio, so that the listener only receives the translated target language speech and natural background ambient sound, thereby eliminating the acoustic basis for delay comparison and achieving delay pseudo-cancellation, ultimately improving the user experience of simultaneous interpretation. Furthermore, a speaker information set is displayed, including the text content corresponding to the original speech of at least one speaker. Based on the speaker information set, a speaker selection instruction is triggered to determine the target speaker, thus giving the user the right to actively choose, enhancing listening flexibility and scenario applicability, thereby improving the user experience of simultaneous interpretation. It also effectively solves the problem of speech masking in multi-person conversation scenarios, further improving the scenario applicability of simultaneous interpretation and ultimately enhancing the user experience. Moreover, by displaying the text content corresponding to the original speech of at least one speaker, the accuracy and convenience of the user's selection of the target speaker are greatly improved, thus enhancing the user experience. Attached Figure Description
[0042] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0043] Figure 1 This is one of the flowcharts illustrating the delay elimination method for simultaneous interpretation provided by the present invention.
[0044] Figure 2 This is the second flowchart of the simultaneous interpretation delay elimination method provided by the present invention.
[0045] Figure 3This is the third flowchart of the simultaneous interpretation delay elimination method provided by the present invention.
[0046] Figure 4 This is the fourth flowchart of the simultaneous interpretation delay elimination method provided by the present invention.
[0047] Figure 5 This is the fifth flowchart of the simultaneous interpretation delay elimination method provided by the present invention.
[0048] Figure 6 This is a schematic diagram of the delay elimination device for simultaneous interpretation provided by the present invention.
[0049] Figure 7 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0050] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0051] Currently, efforts are being made to reduce latency in machine simultaneous interpretation by improving algorithms and enhancing hardware performance. For example, traditional machine simultaneous interpretation methods often employ a cascaded approach, linking an automatic speech recognition module (for transcribing source speech into source text), a machine translation module (for translating from source text to target text), and a text-to-speech module (for generating target language speech) to achieve automatic speech translation. This traditional approach ensures the accuracy of sentence meaning, but requires each module to process input data sequentially, falling far short of the real-time requirements of simultaneous interpretation. Therefore, merging the automatic speech recognition model, machine translation model, and text-to-speech module into a single integrated model requires only a single decoder, eliminating the need for intermediate stages like in cascaded systems. This significantly reduces computational requirements and inference latency, but still results in significant latency leading to a degraded user experience. Another approach involves using high-performance chips, deploying the model on near-field devices (such as smart headphones, conference control systems, and translation terminals), employing lightweight models and dedicated inference engines, optimizing audio acquisition and transmission hardware, increasing data transmission rates, and enabling module parallelism. However, this still results in significant latency leading to a degraded user experience and increases hardware costs.
[0052] To address the issues with latency reduction schemes in simultaneous interpreting, this study initially focused on the need for machine simultaneous interpreting strategies to balance translation quality and latency, regardless of whether the approach is cascaded or end-to-end. This involves determining when to translate the input speech to ensure users receive accurate translations quickly. Consequently, translation strategies have evolved from fixed Wait-k strategies (fixed translation strategies) to more flexible adaptive strategies, aiming to output the translation as quickly as possible while maintaining quality. Fixed translation strategies primarily involve "waiting" for a certain number of input words (e.g., k words) in the source language speech / text before translating. While simple in structure and easy to control latency, they lose contextual information and cannot guarantee accuracy, especially in languages with significant subject-verb differences. Recent fixed translation strategies, such as the Wait-k strategy, ensure that the model can translate based on the context of at least k words, where k is the number of words the model reads first, and then translates simultaneously with the rest of the source sentence; meaning the output is always the k words following the input. Adaptive translation strategies typically employ chunk-based speech processing (processing speech into chunks) to incrementally generate translations. This approach segments the source text based on dynamic context information before translation. These strategies use specific models or strategies to segment streaming source text, offering greater flexibility than fixed strategies. Recent adaptive strategies utilize minimum information units (the smallest segments whose translation content doesn't change with the following text) for segmentation, dynamically breaking down the source text into independently translatable segments to ensure low-latency, high-quality translations. Current techniques for achieving this segmentation include: 1. CTC (Connectionist Temporal Classification)-assisted boundary detection mechanisms, which determine semantic boundary points based on CTC results, allowing the system to automatically identify translatable units and balance latency and translation quality; 2. Segmentation strategies based on semantic segmentation points, eliminating the need for additional modules and enabling the LLM (Large Language Model) to translate semantically intact, achieving a balance between latency and quality; 3. Utilizing dynamic caching of dialogue context to reduce repetitive semantic analysis time, improve translation coherence, and reduce redundant computational latency.
[0053] Further research into the aforementioned delay reduction schemes for simultaneous interpreting revealed that while delays can be further reduced, they cannot be completely eliminated.
[0054] To address the shortcomings of the aforementioned latency reduction schemes in simultaneous interpreting, ongoing research has revealed that as long as objective latency persists, listeners can perceive this latency by comparing the original and translated audio, thus the problem of poor user experience remains. Specifically, the root cause of the negative impact of this objective latency on the listener's subjective experience is not merely the absolute duration of the latency, but more importantly, the fact that listeners simultaneously hear the speaker's original language (i.e., the original audio) and the system's output target language (i.e., the translated audio). In this situation, the speaker's original audio becomes a real-time time reference anchor, allowing listeners to clearly compare and perceive the lag in the translated audio, resulting in a disjointed and discordant auditory experience. Furthermore, psychoacoustic research shows that when an audio signal (such as the translated audio) is delayed by more than a certain threshold (e.g., 35 milliseconds) compared to the main audio signal (such as the original audio), it is recognized by the brain as an independent echo, thereby disrupting auditory continuity and immersion. Based on this, the ultimate goal is to reduce the perceived latency from the listener's perspective, without altering the system latency, while simultaneously improving translation quality.
[0055] As mentioned earlier, for listeners, what affects their sensory experience is the asynchrony between hearing the speaker's voice and hearing the translated voice. This delay uses the speaker's voice as a reference anchor. From this perspective, as long as the listener cannot hear the speaker's original voice, the sensory experience can be greatly improved. On the other hand, this method can also greatly reduce the need to reduce system latency, thereby improving translation quality.
[0056] Based on the aforementioned continuous research, a delay cancellation method for simultaneous interpreting was finally proposed. This method separates the original speech of the target speaker from the mixed audio, obtaining background audio that masks the original speech. This masks the original speech, which serves as an anchor point for delay perception, directly eliminating the possibility of time comparison for the user. Therefore, it completely solves the negative impact of delay at the subjective experience level, ultimately achieving pseudo-delay cancellation to improve the user experience of simultaneous interpreting. Furthermore, while masking the original speech, it retains all other background ambient sounds, greatly enhancing the sense of presence and immersion, thus improving the user experience. Since translation quality is no longer sacrificed for extremely low latency, simultaneous interpreting the original speech yields highly accurate target speech, improving the accuracy of simultaneous interpreting and ultimately enhancing the user experience. Simultaneously, by fusing the target speech with the background audio, a further improvement is achieved. The system outputs audio so that listeners only receive the translated target language speech and natural background ambient sound, thus eliminating the acoustic basis for delay comparison and achieving pseudo-delay cancellation, ultimately improving the user experience of simultaneous interpretation. It also displays a speaker information set, including text content corresponding to the original speech of at least one speaker. Based on the speaker information set, a speaker selection instruction is triggered to determine the target speaker, giving users the right to actively choose. This enhances listening flexibility and scenario applicability, further improving the user experience of simultaneous interpretation. It also effectively solves the problem of speech masking in multi-person conversation scenarios, improving the scenario applicability of simultaneous interpretation and ultimately enhancing the user experience. Furthermore, by displaying text content corresponding to the original speech of at least one speaker, it greatly improves the accuracy and convenience for users when selecting the target speaker, thus enhancing the user experience.
[0057] The delay elimination method for simultaneous interpreting provided by the present invention will be described below through various embodiments. Figures 1-5 The present invention describes a delay elimination method for simultaneous interpreting.
[0058] Figure 1 This is one of the flowcharts illustrating the delay elimination method for simultaneous interpreting provided by the present invention, such as... Figure 1 As shown, the delay elimination method for simultaneous interpretation includes the following steps 110, 120, 130 and 140.
[0059] Step 110: Obtain mixed audio including the original language speech of at least one speaker.
[0060] In this embodiment of the invention, the execution subject of the simultaneous interpretation delay elimination method can be an electronic device, which may include, but is not limited to: simultaneous interpretation headsets (such as smart headsets), conference control centers, translation terminals, smartphones, tablets, cloud servers, desktop computers, laptops, etc.
[0061] Here, the number of speakers can be one or more, thus accommodating both solo presentations and more complex meeting scenarios such as group discussions and debates. Different speakers correspond to different original language pronunciations. These original language pronunciations are the pronunciations of the original language, which is the language before translation.
[0062] Here, mixed audio refers to the original audio signal containing multiple sound sources, recorded in a real-world scenario using audio acquisition devices such as microphones. In one embodiment, in a simultaneous interpretation scenario, mixed audio may include: the original language speech of one or more speakers (e.g., the English speech of a conference presenter).
[0063] Furthermore, the mixed audio also includes ambient noise and / or the original speech of non-target speakers. It should be understood that ambient noise and the original speech of non-target speakers together constitute background noise. For example, background noise may further include the conversation of non-target speakers, applause in the venue, air conditioning noise, equipment prompts, etc.
[0064] In one specific embodiment, the acquisition of mixed audio can be accomplished using an audio acquisition device integrated into the user device. In one embodiment, the audio acquisition device can be a simultaneous interpretation headset with physical noise reduction capabilities; the headset integrates a microphone array capable of acquiring ambient sound from different directions, providing a data basis for subsequently distinguishing the spatial locations of different speakers.
[0065] In one embodiment, the mixed audio is acquired through a microphone array, thereby mixing the audio into multi-channel audio.
[0066] In one embodiment, the mixed audio is one or more frames of audio data, and the number of frames in the multi-frame audio data is less than a preset number of frames, thereby ensuring word-level masking of the target speaker, that is, masking the target speaker as soon as a word is spoken, or even before the target speaker has finished speaking, thereby avoiding a large difference between the playback time of the target language speech and the speaking time of the original language speech, that is, avoiding a mismatch between the target language speech and the background audio, avoiding a poor listening experience for the user, and ultimately improving the user experience of simultaneous interpretation.
[0067] Step 120: Separate the original speech of the target speaker from the mixed audio to obtain background audio that masks the original speech of the target speaker.
[0068] Wherein, the target speaker is one of the at least one speakers, and the original speech of the target speaker is the speech to be simultaneously interpreted.
[0069] Here, the target speaker refers to one or more specific speakers in the mixed audio that the user wants to hear the translated content from. The number of target speakers can be one or more.
[0070] The target speaker is determined based on the following method:
[0071] Display speaker information set; the speaker information set includes the text content corresponding to the original speech of the at least one speaker;
[0072] The target speaker is determined based on the speaker selected by the speaker selection instruction; the speaker selection instruction is an instruction triggered based on the speaker information set.
[0073] Here, speaker selection instructions refer to instructions issued by the user through an interactive interface (such as a touch screen, physical buttons, or voice commands) to specify one or more target speakers.
[0074] Here, the speaker information set refers to a selectable set of information presented to the user after analyzing at least one speaker, such as a speaker information table.
[0075] Considering potential issues in practical applications—such as when multiple unfamiliar speakers are present, even if users see labels like "Speaker 1 (Male)" and "Speaker 2 (Female)," they may still be unable to immediately determine which speaker they truly want to listen to—this approach directly displays snippets of the speaker's speech, providing users with the most direct and unambiguous basis for judgment.
[0076] In one embodiment, the speaker's original speech is quickly transcribed into text using a lightweight, real-time automatic speech recognition engine, and the text is dynamically displayed next to the corresponding speaker option.
[0077] Furthermore, the speaker information set also includes a content summary or keywords corresponding to the original speech of at least one speaker. In one embodiment, the original speech of the speaker is quickly transcribed into text using a lightweight, real-time automatic speech recognition engine. The text is then subjected to rapid natural language processing to extract a content summary or keywords, thereby presenting it to the user in a more concise manner.
[0078] The automatic speech recognition engine is characterized by speed priority. Its goal is not to obtain 100% accurate text recordings, but to quickly and roughly accurately identify the text content being spoken.
[0079] It should be understood that the embodiments of the present invention grant users the right to actively choose, solving the problem that traditional speech processing technology is unable to accurately and flexibly lock onto specific processing objects in complex multi-person conversation scenarios. This makes the simultaneous interpretation system no longer limited to a solo mode that can only handle a single speaker, but can easily cope with complex scenarios such as multi-person discussions, debates, and roundtable conferences. Users can freely switch the listening object (target speaker) according to the meeting agenda and personal interests, thereby greatly enhancing listening flexibility and scenario applicability, and thus improving the user experience of simultaneous interpretation.
[0080] It should be understood that users are no longer passively receiving pre-set translations from the system, but can choose what they listen to. This active choice significantly enhances user participation and satisfaction, giving users complete control, improving the personalized experience, and thus enhancing the user experience of simultaneous interpretation.
[0081] It should be understood that speaker separation technology encounters difficulties when handling multiple targets and requiring dynamic switching between them. This invention, by providing a speaker information set and acquiring user-triggered speaker selection commands, accurately conveys the user's intent to the backend speech separation model. This achieves stable and reliable masking of one or more specific speakers in complex acoustic environments, effectively solving the speech masking problem in multi-person conversation scenarios, thereby improving the applicability of simultaneous interpretation and ultimately enhancing the user experience of simultaneous interpretation.
[0082] It should be understood that the embodiments of the present invention directly show what is being said, changing the user's choice from guessing to confirmation, which greatly improves the accuracy of identifying the target speaker in scenarios with multiple strangers, similar voices, or where direct observation is not possible, thereby improving the accuracy of identifying the target speaker and thus enhancing the user experience.
[0083] It should be understood that users no longer need to engage in voice-to-speaker matching; they can make a judgment simply by glancing at the text. In fast-paced discussions with frequent speaker switching, this WYSIWYG selection method greatly reduces the cognitive burden on users, making the selection of the target speaker smoother. This significantly improves the efficiency and convenience of user selection, thereby enhancing the user experience.
[0084] It should be understood that even if a user chooses to listen to only one target speaker, the real-time text content of other speakers displayed on the interface allows the user to understand what others are discussing at any time, thereby gaining better control over the overall meeting, not missing any potentially interesting topics, and thus improving the user experience.
[0085] In one specific embodiment, the pure original speech of the target speaker is extracted from the mixed audio, thereby improving the accuracy of simultaneous interpretation and enhancing the accuracy of target speech generation. Furthermore, it ensures that the original speech of the target speaker is completely masked in the background audio, ensuring that the user is completely unaware of the original speech of the target speaker, thus improving the effect of delay cancellation and ultimately enhancing the user experience. Accordingly, by subtracting or suppressing the original speech of the target speaker from the mixed audio, a pure background audio that retains all environmental details can be obtained.
[0086] Here, background audio refers to the audio portion remaining after the target speaker's speech component has been stripped or suppressed from the original mixed audio. It should be understood that background audio is not muted, but rather retains all other environmental sounds (such as ambient noise and the original speech of non-target speakers) to ensure a natural and undistorted final sound, thereby improving the user experience.
[0087] Step 130: Simultaneous interpretation is performed on the original speech of the target speaker to obtain the target speech of the target speaker.
[0088] Here, simultaneous interpretation refers to the process of machine simultaneous interpretation, which is to translate the speech of one source language into the speech of another target language in real time.
[0089] Here, target language speech refers to the speech of the target language that the audience ultimately needs to hear after the translation task is completed (e.g., Chinese speech).
[0090] In one specific embodiment, the original speech of the target speaker is input into a simultaneous interpreting model to obtain the target language speech output by the simultaneous interpreting model. This simultaneous interpreting model can be an end-to-end model, directly mapping the input original language speech signal to the target language speech signal, thereby reducing the intermediate steps and cumulative latency of traditional cascaded systems (speech recognition-machine translation-speech synthesis). It should be noted that, since this embodiment of the invention solves the subjective latency problem, the simultaneous interpreting model can adopt a strategy that focuses more on translation quality (e.g., waiting for a complete semantic group to finish before starting translation), without excessively pursuing extremely low latency, thereby improving the accuracy of simultaneous interpreting.
[0091] Step 140: The target language speech is fused with the background audio to obtain the output audio.
[0092] Here, the output audio is the audio output to the user. In one embodiment, the output audio is played through a simultaneous interpretation device (such as simultaneous interpretation headphones) to play the output audio to the audience. It should be noted that the simultaneous interpretation device can block out all external sounds, thereby avoiding hearing the original speech of the target speaker. This can be achieved through physical noise reduction, so that the audience can only hear the output audio, thus ensuring the effectiveness of delay cancellation.
[0093] Specifically, the translated target language speech is used as the foreground voice, and the preserved ambient sound is used as the background audio. The two are superimposed to obtain the original language speech output, excluding the target speaker. In other words, the target language speech and background audio are reconstructed and integrated into a single audio output, with the goal of improving the overall quality, naturalness, and clarity of the final output audio.
[0094] In one specific embodiment, the target language speech and background audio are input into an acoustic fusion model to obtain the output audio of the acoustic fusion model. In one embodiment, the acoustic fusion model is trained based on two sample audio samples and their corresponding output audio labels.
[0095] In one embodiment, a trainable weighted formula can be used to fuse the target language speech with the background audio to ensure that the synthesized sound is clear, natural, and undistorted. For example, the weighted formula is as follows:
[0096] ;
[0097] In the formula, This indicates the final output audio; This represents a trainable parameter matrix; Represents the speech of the target language; This indicates background audio.
[0098] For ease of understanding, a specific embodiment is described herein. Exemplarily, user Li Ming wears a pair of intelligent simultaneous interpretation earphones and participates in an international artificial intelligence summit held in a large conference hall. The keynote speaker is an English-speaking expert, and there are also the voices of other audience members and environmental noises at the scene. After Li Ming puts on the earphones, the microphone array on the earphones starts to work, collects all the sounds in the venue, and forms a multi-channel mixed audio; this mixed audio contains the English speech of the keynote speaker, the slight coughing sound of the audience beside, and the sound of closing the door coming from afar. A interface pops up on the mobile App (application program) connected to the intelligent simultaneous interpretation earphones, showing that the speaker 1 on the main podium is detected. Li Ming clicks to select speaker 1 as the target speaker. Immediately, the voice separation module is started. The module generates a mask matrix based on the voiceprint characteristics and the sound source spatial orientation characteristics of speaker 1; through this mask matrix, two channels of audio are separated: one is the pure original English voice of the keynote speaker, and the other is the background audio, which contains the coughing sound of the audience and the sound of closing the door, but the original English voice of the keynote speaker has been completely blocked. That pure original English voice of the keynote speaker is sent into the end-to-end simultaneous interpretation engine built into the earphones. After the engine monitors that the keynote speaker has finished speaking a complete short sentence (such as “So today, we're going to talk about the future of AI”), it translates it into the Chinese voice “那么今天,我们将要讨论人工智能的未来” completely and smoothly. The translated Chinese voice is intelligently fused with the background audio separated before. Finally, what Li Ming hears through the earphones is that, against the background of the real ambient sound of the scene, a clear and smooth Chinese female voice (the voice of the interpreter) is broadcasting. He can't hear the original English voice of the keynote speaker at all. Based on this, although there is objectively a delay in machine translation, since Li Ming can't hear the original English voice as a comparison anchor point, he only feels that there is a natural pause between sentences in the Chinese translation, and there is no tearing feeling and delay feeling of the original voice and the translated voice fighting against each other, and he obtains a smooth and immersive experience like listening to a local radio, thus improving the user experience.
[0099] Exemplarily, as Figure 2 shown, this mixed audio is first sent into the selectable speaker shielding system. The selectable speaker shielding system separates the original voice of the designated speaker (the original language voice of the target speaker) and outputs it to the end-to-end machine simultaneous interpretation system below for translation, and generates the background audio with the original voice of the designated speaker blocked, and outputs this background audio to the subsequent acoustic fusion module; after receiving the pure original voice, the end-to-end machine simultaneous interpretation system performs the translation task and outputs the target language voice of the designated speaker; finally, the acoustic fusion module receives the translated target language voice and the shielded background audio, and fuses the two to generate the final output audio.
[0100] It should be understood that existing technologies strive to reduce objective delays, but cannot eliminate them entirely. This invention departs from this approach by masking the original spoken language, which serves as an anchor point for delay perception, directly eliminating the possibility of users making time comparisons. This completely resolves the negative impact of delay at the subjective experience level, fundamentally eliminating the listener's subjective perception of delay.
[0101] It should be understood that since translation quality is no longer sacrificed for extremely low latency, it can be improved. For example, simultaneous interpreting models can wait for more complete semantic units (such as a phrase or clause) to be input before translating. This allows simultaneous interpreting models to better understand the context, resulting in more accurate, fluent, and grammatically correct translations that conform to the target language, thereby improving the accuracy of simultaneous interpreting and ultimately enhancing the user experience. In other words, it indirectly improves the overall quality of simultaneous interpreting.
[0102] It should be understood that while masking the original voice of the target speaker, all other background ambient sounds are carefully preserved. This allows listeners to still feel the atmosphere of the scene (such as applause, laughter, and other ambient sounds) while listening to the translation, avoiding the isolated listening experience of traditional simultaneous interpretation. This greatly enhances the sense of presence and immersion, thereby improving the user experience of simultaneous interpretation. In other words, it significantly enhances the naturalness and immersion of the auditory experience.
[0103] It should be understood that the root cause of the poor experience for listeners in machine simultaneous interpretation is not the objective computational delay itself, but rather the listener's ability to compare the translated audio with the speaker's original audio, which serves as a time anchor, thus subjectively perceiving a delay. Based on this finding, this invention proposes a delay pseudo-cancellation method, which actively and selectively blocks the original language speech of the target speaker, allowing listeners to receive only the translated target language speech and natural background ambient sound, thereby eliminating the acoustic basis for delay comparison and eradicating the subjective sense of delay.
[0104] The simultaneous interpreting delay elimination method provided in this invention separates the original speech of the target speaker from mixed audio, obtaining background audio that masks the original speech of the target speaker. This masks the original speech, which serves as an anchor point for delay perception, directly eliminating the possibility of time comparison for the user. Therefore, it completely solves the negative impact of delay at the subjective experience level, ultimately achieving pseudo-delay elimination to improve the user experience of simultaneous interpreting. Furthermore, while masking the original speech of the target speaker, it retains all other background ambient sounds, greatly enhancing the sense of presence and immersion, thus improving the user experience of simultaneous interpreting. Since translation quality is no longer sacrificed for extremely low latency, simultaneous interpreting the original speech of the target speaker can yield highly accurate target speech, thereby improving the accuracy of simultaneous interpreting and ultimately enhancing the user experience of simultaneous interpreting. Simultaneously, the target speech is compared with the background audio... The system integrates and outputs audio, allowing listeners to receive only the translated target language speech and natural background ambient sound. This eliminates the acoustic basis for delay comparison, achieving pseudo-delay cancellation and ultimately improving the user experience of simultaneous interpreting. Furthermore, it displays a speaker information set, including text content corresponding to at least one speaker's original language speech. Based on the speaker information set, a speaker selection command is triggered to determine the target speaker, granting users the right to actively choose. This enhances listening flexibility and scenario applicability, further improving the user experience of simultaneous interpreting. It also effectively solves the problem of speech masking in multi-person conversation scenarios, improving the scenario applicability of simultaneous interpreting and ultimately enhancing the user experience. Moreover, by displaying text content corresponding to at least one speaker's original language speech, it greatly improves the accuracy and convenience for users when selecting the target speaker, thus enhancing the user experience.
[0105] Based on any of the above embodiments, in this method, the speaker information set further includes the identification information of the at least one speaker.
[0106] Here, identification information refers to the information used to uniquely distinguish different speakers in the speaker information set. This identification information can be a simple number (such as "Speaker 1"), or more detailed descriptive information (such as "Platform, Male Voice", "Current Speaker"), or it can be an ID (Identification).
[0107] For example, in an international business roundtable forum, three guests (American expert A, French expert B, and Japanese expert C) speak in their respective native languages, coordinated by a Chinese moderator. Mr. Wang, an audience member, wants to follow the views of American expert A and Japanese expert C. Mr. Wang opens the app that comes with his simultaneous interpretation headset. The app displays a list of speakers detected by the system in real time. Mr. Wang selects two checkboxes for expert A and expert C on the app, which issues a speaker selection command, identifying expert A and expert C as the target speakers. When American expert A speaks, Mr. Wang's headset plays a Chinese translation of A's speech, while expert A's original English and the voices of others (private conversations between B and C) are muted or treated as faint background noise. When the discussion shifts to Japanese expert C, the system automatically recognizes the change in speaker. Since expert C is also on Mr. Wang's selection list, the headset seamlessly switches, playing a Chinese translation of C's speech and muting C's original Japanese. If French expert B were to interrupt at this point, since expert B was not selected, Mr. Wang would not hear any translation from expert B in his headphones. Of course, expert B's original voice could also be muted to ensure that Mr. Wang's listening focus is not interrupted.
[0108] For example, text is pushed to the user interface in real time and associated with the corresponding speaker's identification information. For instance, the speaker information set might be "(Speaker 1: Current speech segment; Speaker 2: Current speech segment)". As the speaker continues speaking, the text content within the speaker information set scrolls or refreshes like subtitles, always displaying the latest content segment. After seeing this real-time content segment, the user can immediately understand the identity of each speaker and the topic of their discussion, allowing them to easily select the speaker of interest and trigger the speaker selection command.
[0109] The simultaneous interpretation delay elimination method provided in this invention, through the above-mentioned approach, grants users the right to actively choose, enhances listening flexibility and scenario applicability, and thus improves the user experience of simultaneous interpretation; it also effectively solves the problem of voice blocking in multi-person conversation scenarios, thereby improving the scenario applicability of simultaneous interpretation and ultimately improving the user experience of simultaneous interpretation.
[0110] Based on any of the above embodiments, in this method, the identification information of the at least one speaker is determined in the following manner:
[0111] The sound source is located in the mixed audio to obtain the spatial location features of the sound source of the at least one speaker, and the voiceprint features are extracted from the mixed audio to obtain the voiceprint features of the at least one speaker.
[0112] The spatial orientation features of the sound source of the at least one speaker and the voiceprint features of the at least one speaker are input into the target association model to obtain the speech feature identification vector of the at least one speaker output by the target association model.
[0113] Based on the speech feature identification vector of the at least one speaker, the identification information of the at least one speaker is determined; the speech feature identification vector of the at least one speaker corresponds one-to-one with the identification information of the at least one speaker.
[0114] Here, sound source localization is a technique that uses a microphone array to determine the physical location of a sound source. In one specific embodiment, by analyzing the time difference of arrival (TDOA) or phase difference of sound at different microphones, the spatial orientation characteristics of the sound source, such as the azimuth and elevation angles at which the sound is transmitted, can be calculated. These spatial orientation characteristics are used to represent the physical location of the sound source.
[0115] Specifically, real-time sound source localization is performed on the mixed audio to obtain the sound source spatial orientation features of at least one speaker, so as to distinguish the physical location of different speakers and form a dynamic speaker spatial orientation matrix (including information such as azimuth angle, pitch angle, angular velocity, etc.), thereby generating the sound source spatial orientation features of at least one speaker.
[0116] Here, voiceprint feature extraction is a technique for extracting biometric features from speech signals that can uniquely identify the speaker. It analyzes the unique acoustic characteristics of a human voice, such as timbre, pitch, and formants, to generate a high-dimensional voiceprint feature.
[0117] Specifically, the voiceprint features of at least one speaker are separated and extracted from the mixed audio.
[0118] In one specific embodiment, mixed audio is input into a pre-trained voiceprint recognition model to obtain the voiceprint features of at least one speaker output by the voiceprint recognition model.
[0119] The target association model is used to fuse the spatial location features and voiceprint features of the same speaker's sound source to obtain a speech feature identification vector.
[0120] In one embodiment, the target association model is a feature fusion model, typically constructed from a deep neural network. Instead of simply concatenating two features, this target association model acts like an intelligent decision center, learning and understanding the relationship between a location and a voiceprint, ultimately outputting a more discriminative internal identity representation that integrates both types of information—a voice feature identifier vector.
[0121] In one embodiment, the target association model is constructed using multiple layers of nonlinear neural networks.
[0122] It's important to note that this target association model addresses a key question: "Is the loud male voice I heard on the left the same loud male voice I'm hearing in the middle now?" If the voiceprint features match, but the spatial location has reasonably changed (e.g., moving from left to center), the target association model will determine that it's the same person who has moved, and will maintain the original voice feature vector. If the spatial locations are the same or similar, but the voiceprint features have changed significantly, the target association model will determine that it's the same location but a different person speaking, and will generate a new voice feature vector. If the spatial locations are different, and the voiceprint features are also different, then it is undoubtedly a new person.
[0123] Here, the speech feature identifier vector is an internally used, highly condensed, and robust speaker identity vector. It is more reliable than a single voiceprint or location information and serves as a digital identity card that the system creates for each individual speaker.
[0124] It should be noted that this speech feature identification vector is the identification vector corresponding to the acquisition time of the original language speech. That is, it generates the speech feature identification vector of the original language speech frame, which is used to identify the speaker's sound source spatial location features and voiceprint features corresponding to the original language speech frame.
[0125] After identifying at least one speaker, a speaker information set can be generated and then provided to the user in a synchronized manner, so that the user can select the speaker to be muted as needed, that is, specify the original language speech of the simultaneous interpretation (the speech to be muted).
[0126] For example, at a company product launch, Mr. Zhang is giving a speech in the center of the stage (position A). Midway through, CTO Mr. Li comes on stage from the left (position B) to add a few technical details; Mr. Li's voice is somewhat similar to Mr. Zhang's. Mr. Zhang then moves to the right (position C) of the stage to continue his speech. During Mr. Zhang's speech at position A, his position A features and voiceprint Z features are extracted, fused to generate a voice feature identifier vector Vector_Z, and displayed on the app as "Mr. Zhang (center)". If Mr. Li speaks at location B, and only voiceprint recognition is used, the system might mistakenly identify Mr. Zhang as the speaker because Mr. Li's voiceprint is similar to Mr. Zhang's. If only sound source localization is used, the system only knows that the sound comes from B, but does not know who it is. However, by extracting the location B feature and the voiceprint L feature, the target association model receives [location B] and [voiceprint L]. Even if the voiceprint L is somewhat similar to the voiceprint Z, the model determines that this is a combination of a new voiceprint and a new location. Therefore, it decisively creates a new voice feature identifier vector Vector_L for it and adds an item "Mr. Li (left side)" to the App. Mr. Zhang moves to position C to continue his speech. If the system only uses sound source localization, it will think that this is a new speaker because it cannot associate position C with position A. However, by extracting the features of position C and the voiceprint Z feature, the target association model receives [position C] and [voiceprint Z]. Although the position has changed, the model recognizes that the voiceprint Z is known (associated with Vector_Z). Therefore, the model determines that this is a combination of a known voiceprint and a new position, that is, the same person has moved. It will not create a new identity, but will update the position information associated with Vector_Z and dynamically update "Mr. Zhang (center)" to "Mr. Zhang (right)" on the App.
[0127] It should be understood that the embodiments of the present invention, through a dual verification mechanism of spatial location and biometric voiceprint, overcome the inherent limitations of single-feature verification. It avoids misidentifying a person due to similar voiceprints or failing to recognize a person due to their location changing. Therefore, the accuracy of speaker identification information is improved, resulting in a qualitative leap even in complex real-world scenarios. This significantly enhances the accuracy and robustness of speaker recognition, thereby improving the accuracy of identifying the target speaker and ultimately enhancing the user experience.
[0128] It should be understood that in noisy environments, it may be difficult to extract high-quality, long-duration, clean voiceprints. In this embodiment of the invention, even if the voiceprint features are somewhat blurry, as long as the spatial location information is clear, the system can still make an accurate judgment. This makes the embodiments of the invention more adaptable to real-world environments (i.e., robust), combining information from two different dimensions to refer to the same person, thereby greatly improving the accuracy and reliability of speaker identification in complex environments such as multi-person, dynamic, and noisy situations. This significantly reduces the stringent requirements for the quality of a single feature, avoiding over-reliance on the accuracy of identification based on a single feature.
[0129] It should be understood that the embodiments of the present invention can intelligently handle highly dynamic situations such as speaker movement and multiple people taking turns speaking, ensuring continuous locking of the target speaker's identity. This is crucial for ensuring the continuity of simultaneous interpretation and avoiding confusion for users during selection and listening. It achieves continuous and stable tracking of speakers in dynamic environments, thereby improving the accuracy of target speaker identification and thus enhancing the user experience of simultaneous interpretation.
[0130] The simultaneous interpreting delay elimination method provided in this invention combines two different dimensions of information to refer to the same speaker, thereby greatly improving the accuracy and reliability of speaker identification in complex environments such as multi-person, dynamic, and noisy situations. This improves the accuracy of identifying the target speaker and ultimately enhances the user experience of simultaneous interpreting.
[0131] Based on any of the above embodiments, in this method, step 120 includes:
[0132] Generate a mask matrix for reconstructing the original speech of the target speaker;
[0133] Based on the mask matrix and the mixed audio, speech reconstruction is performed to obtain the original speech of the target speaker;
[0134] The original speech of the target speaker in the mixed audio is masked to obtain the background audio.
[0135] Here, the mask matrix can be understood as an intelligent, semi-transparent digital acoustic filter that operates in the time-frequency domain of audio.
[0136] In one embodiment, the size of the mask matrix is exactly the same as that of the spectrogram. Each element in the mask matrix has a value between 0 and 1, corresponding to each time-frequency cell in the spectrogram. This element value represents a probability, that is, how likely it is that the energy of the time-frequency cell belongs to the target speaker. In other words, the mask matrix is a time-frequency probability matrix with values between 0 and 1, marking the probability that each time point and each frequency unit belongs to the target speaker.
[0137] Furthermore, if the element value is 1, it means that the energy at this time point and frequency is 100% from the target speaker; if the element value is 0, it means that it is 100% from the background sound. This either 0 or 1 masking mechanism can significantly reduce the artifacts such as sound distortion and noise after separation, ensuring that pure original speech and pure background audio are obtained.
[0138] Here, speech reconstruction refers to the process of extracting and reconstructing the original speech of the target speaker from mixed audio based on the guidance provided by the mask matrix. In one specific embodiment, the spectrogram of the mixed audio is multiplied element-wise with the mask matrix to obtain the original speech of the target speaker.
[0139] Here, masking refers to the operation of removing the target speaker's original speech from the original mixed audio. In one embodiment, the most straightforward method is spectral subtraction, which involves subtracting the spectrogram of the target speaker's original speech reconstructed in the previous step from the spectrogram of the mixed audio.
[0140] For example, such as Figure 3 As shown, a target speaker mask matrix is generated to reconstruct the original speech of the target speaker. The multi-channel audio input (mixed audio) and the target speaker mask matrix are input to the mask application module to synthesize the original speech estimate of the selected speaker (the original speech of the target speaker), which is to perform speech signal reconstruction. Then, the original speech estimate of the selected speaker is extracted from the multi-channel audio input through a masking operation, and other speech information is retained to obtain the background audio output. The background audio includes necessary auxiliary audio information in the environment (such as background speech, noise, etc.) to prevent the synthesized sound from being distorted. Finally, a natural background audio output that masks the speech of the target speaker is obtained.
[0141] It should be understood that by operating at the fine time-frequency unit level through a masking matrix, pixel-level acoustic separation can be achieved. This ensures that the reconstructed original speech is very clean, while the background audio left after masking is also extremely complete, with low distortion and high fidelity. This improves the accuracy of simultaneous interpretation and ensures that the output audio is completely free of the target speaker's original speech, thereby improving the delay cancellation effect of simultaneous interpretation and ultimately enhancing the user experience.
[0142] The simultaneous interpreting delay elimination method provided in this invention improves the accuracy of speech separation and speech masking through the above-mentioned method. That is, it ensures that the reconstructed original speech is very clean, while the background audio left after masking is also extremely complete, with little distortion and high fidelity, thereby improving the accuracy of simultaneous interpreting. It also ensures that the output audio is completely free of the original speech of the target speaker, thereby improving the delay elimination effect of simultaneous interpreting and thus improving the user experience of simultaneous interpreting.
[0143] Based on any of the above embodiments, in this method, generating a mask matrix for reconstructing the original speech of the target speaker includes:
[0144] The mixed audio and the speech feature vector of the target speaker are input into the speech separation model to obtain the mask matrix output by the speech separation model.
[0145] The target speaker's voice feature identifier vector is obtained by fusing the target speaker's sound source spatial location features and voiceprint features.
[0146] Here, the speech feature identifier vector is a robust internal identifier that integrates spatial location and voiceprint information. In this embodiment of the invention, it acts as an instruction or condition, telling the speech separation model to find the voice with the feature corresponding to this condition from the given mixed audio.
[0147] The speech separation model is obtained by training based on sample mixed audio.
[0148] Here, the speech separation model refers to a deep neural network model used to separate specific sound sources from mixed signals. Specifically, this speech separation model is a conditional speech separation model because it receives an additional conditional input (i.e., a speech feature identifier vector) to guide its separation behavior.
[0149] In one specific embodiment, the acoustic features of the mixed audio and the speech feature vector of the target speaker are input into a speech separation model to obtain the mask matrix output by the speech separation model. In one embodiment, the acoustic features can be features obtained by performing a Fourier transform on the mixed audio; for example, the acoustic features can be an amplitude spectrum or a Mel spectrum.
[0150] In one embodiment, the mask matrix is an ideal binary mask, where each element is either 0 or 1. A value of 1 indicates that the energy of that time-frequency unit is mainly contributed by the target speaker; a value of 0 indicates that the energy of that time-frequency unit is mainly contributed by the background audio. It should be understood that using this hard mask with values of either 0 or 1 aims to maximize the suppression of the target speaker's original speech, achieving cleaner and more thorough separation, ensuring that almost no original speech components of the target speaker remain in the background audio, thereby achieving the best delay pseudo-cancellation effect.
[0151] In one embodiment, to make the mask matrix output by the speech separation model more accurate, the speech separation model is trained using a set of paired audio data of a target speaker with samples and a target speaker with masks. The speech of the target speaker with samples is processed according to the aforementioned process to obtain the sample background audio of the target speaker with masks. The loss (such as Mel loss) between the sample background audio and the real background audio of the target speaker with masks is calculated. The speech separation model is optimized by minimizing this loss during the training process.
[0152] It should be understood that the speech separation model is not a model that blindly separates all sound sources, but rather a guided or conditional separation process. By using the speech feature vector representing the target speaker's identity as one of the key inputs to the speech separation model, it can be directed to operate only on the user-selected target speaker, thereby achieving accurate speech separation and ensuring both selectivity and accuracy in speech separation.
[0153] The simultaneous interpreting delay cancellation method provided in this invention uses the speech feature identifier vector representing the target speaker's identity as one of the key inputs to the speech separation model. This ensures that the reconstructed original speech is very clean, while the background audio left after masking is also extremely complete, with low distortion and high fidelity. This improves the accuracy of simultaneous interpreting and ensures that the output audio is completely free of the target speaker's original speech, thereby improving the delay cancellation effect of simultaneous interpreting and ultimately enhancing the user experience of simultaneous interpreting.
[0154] To facilitate understanding of the above embodiments, as follows: Figure 4 As shown, firstly, multi-channel audio input (mixed audio) is acquired through a pre-configured microphone array (such as a headset for simultaneous interpretation or a conference room microphone combination); secondly, the sound source localization module analyzes the multi-channel audio input in real time and calculates the speaker spatial orientation vector (sound source spatial orientation feature) of each main sound source; simultaneously, the speaker voiceprint feature extraction module extracts the speaker voiceprint features from the multi-channel audio input to obtain a voiceprint feature vector that can uniquely identify the speaker; then, the speaker spatial orientation vector and the voiceprint feature vector are input into the dynamic target association module (target association model) to obtain the speaker speech information (speech feature identification vector) output by the dynamic target association module; based on the speaker speech information, The system generates identification information for each speaker, forming a list of selectable speakers (speaker information set). Then, the listener or system selects one or more target speakers from the list of selectable speakers to determine the target for subsequent separation. Next, the mixed audio and the speech feature vectors of the user-selected target speakers are input into a conditional separation network (speech separation model). The network dynamically generates a target speaker mask matrix based on the input, covering all time-frequency units of the audio. Finally, signal reconstruction is performed based on the target speaker mask matrix and the multi-channel audio input, reconstructing the original speech output (original language speech) of the target speaker and the background audio output after subtracting the target speaker's speech.
[0155] Based on any of the above embodiments, in this method, step 130 includes:
[0156] The original speech of the target speaker is encoded to obtain speech coding features;
[0157] The input features based on the speech coding feature conversion are input into the speech feature generation model to obtain the speech features output by the speech feature generation model;
[0158] The speech features are decoded to obtain the target language speech of the target language.
[0159] Here, speech coding refers to converting the input raw speech into a high-dimensional feature representation that can be processed by a computer, such as Mel spectrum, hidden layer features extracted by convolutional neural networks, or temporal coding vectors. These features can condense the time-frequency and semantic information of speech. Speech coding features are the output of speech coding, carrying a set of feature vectors that sufficiently reconstruct the multidimensional semantic and acoustic information of the original speech, and serve as the input to subsequent machine learning models.
[0160] In one specific embodiment, the original speech of the target speaker is input into a speech encoder to obtain the speech coding features output by the speech encoder. Based on this, the time-domain speech signal is converted into high-dimensional, time-aligned speech coding features, typically a sequence of Mel spectrum or hidden layer activation vectors.
[0161] The input features are those adapted to the speech feature generation model. This adaptation process may include linear transformation, normalization, and positional encoding embedding, so that the input features are consistent with the input space during the training of the speech feature generation model.
[0162] In one specific embodiment, speech coding features are input to an adapter to obtain input features output by the adapter. Based on this, the aforementioned speech coding features are converted into a feature format acceptable to the speech feature generation model through one or more adapter layers, thereby improving the processing efficiency and accuracy of the speech feature generation model.
[0163] The speech feature generation model is used to generate speech features of the target language based on the speech coding features of the original language. In one embodiment, the speech feature generation model is built on a Large Language Model (LLM). This speech feature generation model directly predicts the speech feature representation of the corresponding target language from the speech coding features of the source language by leveraging powerful contextual understanding and language generation capabilities.
[0164] Here, speech decoding refers to converting the speech features of the target language output by the speech feature generation model back into a playable speech signal (waveform). This speech decoding typically relies on a vocoder or a neural network vocoder.
[0165] In one specific embodiment, speech features are input to a speech decoder to obtain the target language speech of the target language output by the speech decoder.
[0166] For example, such as Figure 5 As shown, the original speech of the speaker (the original language speech of the target speaker) is input into a speech encoder to obtain the speech coding features output by the speech encoder. The speech coding features are then input into an adapter to obtain the input features output by the adapter. The input features are then input into a large language model to obtain the speech features output by the large language model. Finally, the speech features are input into a speech decoder to obtain the speaker's target language speech output by the speech decoder. In other words, the system receives the original speech of a specified speaker, translates it, and outputs the speaker's target language speech with minimal delay, striving to achieve natural and fluent simultaneous interpretation. This machine simultaneous interpretation system includes a speech encoder that converts the input speech into high-dimensional features, then an adapter that converts it into features understandable by a large language model, inputting these features into the large language model to obtain the target language speech features, and finally a speech decoder that restores the speaker's target language speech.
[0167] It should be understood that the end-to-end design reduces the intermediate steps of traditional recognition, translation, and synthesis. Combined with the subjective delay elimination method of the embodiments of the present invention, it enables the audience to obtain an immediate and imperceptible delay experience, that is, further reduce the overall system delay.
[0168] The simultaneous interpretation delay elimination method provided in this embodiment of the invention achieves end-to-end simultaneous interpretation through the above-described manner, thereby reducing the delay of simultaneous interpretation and improving the user experience of simultaneous interpretation.
[0169] Based on the above embodiments, this invention proposes a delay pseudo-elimination scheme. From a psychoacoustic and technical intervention perspective, it "pseudo-eliminates" the perceived delay by suppressing the audience's perception of the speaker's original voice, thereby reducing the negative impact of time lag in machine simultaneous interpretation. This method differs from the traditional approach of simply improving translation algorithms; it proposes actively intervening in the original voice at the audio level. This can eliminate the audience's subjective sense of delay without altering the objective computational delay, thus subjectively improving the simultaneous interpretation experience.
[0170] The delay elimination device for simultaneous interpretation provided by the present invention will be described below. The delay elimination device for simultaneous interpretation described below can be referred to in correspondence with the delay elimination method for simultaneous interpretation described above.
[0171] Figure 6 This is a schematic diagram of the delay cancellation device for simultaneous interpretation provided by the present invention, as shown below. Figure 6 As shown, the simultaneous interpretation delay elimination device includes: an audio acquisition module 610, a speech separation module 620, a simultaneous interpretation module 630, and an audio fusion module 640.
[0172] The audio acquisition module 610 is used to acquire mixed audio including the original language speech of at least one speaker.
[0173] The speech separation module 620 is used to separate the original speech of the target speaker from the mixed audio to obtain background audio that masks the original speech of the target speaker; the target speaker is one of the at least one speakers, and the original speech of the target speaker is the speech to be interpreted simultaneously.
[0174] The simultaneous interpretation module 630 is used to simultaneously interpret the original speech of the target speaker to obtain the target speech of the target speaker.
[0175] The audio fusion module 640 is used to fuse the target language speech with the background audio to obtain the output audio;
[0176] The target speaker is determined based on the following modules:
[0177] The information display module is used to display a speaker information set; the speaker information set includes the text content corresponding to the original speech of the at least one speaker;
[0178] The speaker determination module is used to determine the target speaker based on the speaker indicated by the speaker selection instruction; the speaker selection instruction is an instruction triggered based on the speaker information set.
[0179] Figure 7 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 7As shown, the electronic device may include: a processor 710, a communications interface 720, a memory 730, and a communications bus 740, wherein the processor 710, the communications interface 720, and the memory 730 communicate with each other through the communications bus 740. The processor 710 can invoke logical instructions in the memory 730 to execute a delay cancellation method for simultaneous interpretation. The method includes: acquiring a mixed audio containing the original language speech of at least one speaker; separating the original language speech of a target speaker from the mixed audio to obtain background audio that masks the original language speech of the target speaker; the target speaker is one of the at least one speakers, and the original language speech of the target speaker is the speech to be simultaneously interpreted; performing simultaneous interpretation on the original language speech of the target speaker to obtain the target language speech of the target speaker; fusing the target language speech with the background audio to obtain an output audio; wherein the target speaker is determined based on: displaying a speaker information set; the speaker information set including text content corresponding to the original language speech of the at least one speaker; determining the target speaker based on a speaker selection instruction; the speaker selection instruction being an instruction triggered based on the speaker information set.
[0180] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0181] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the simultaneous interpretation delay cancellation method provided by the above methods. The method includes: acquiring a mixed audio including the original language speech of at least one speaker; separating the original language speech of a target speaker from the mixed audio to obtain background audio that masks the original language speech of the target speaker; the target speaker is one of the at least one speakers, and the original language speech of the target speaker is the speech to be simultaneously interpreted; performing simultaneous interpretation on the original language speech of the target speaker to obtain the target language speech of the target speaker; fusing the target language speech with the background audio to obtain an output audio; wherein the target speaker is determined based on the following method: displaying a speaker information set; the speaker information set including text content corresponding to the original language speech of the at least one speaker; determining the target speaker based on the speaker indicated by a speaker selection instruction; the speaker selection instruction being an instruction triggered based on the speaker information set.
[0182] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a delay cancellation method for simultaneous interpretation provided by the methods described above. The method includes: acquiring a mixed audio comprising the original speech of at least one speaker; separating the original speech of a target speaker from the mixed audio to obtain background audio that masks the original speech of the target speaker; the target speaker being one of the at least one speakers, and the original speech of the target speaker being the speech to be simultaneously interpreted; simultaneously interpreting the original speech of the target speaker to obtain the target speech of the target speaker; and fusing the target speech with the background audio to obtain an output audio; wherein the target speaker is determined based on: displaying a speaker information set; the speaker information set including text content corresponding to the original speech of the at least one speaker; determining the target speaker based on a speaker selection instruction; and the speaker selection instruction being an instruction triggered based on the speaker information set.
[0183] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0184] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0185] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for delay cancellation of simultaneous interpretation, characterized in that, include: Acquire mixed audio including the original speech of at least one speaker; The original speech of the target speaker is separated from the mixed audio to obtain background audio that masks the original speech of the target speaker; the target speaker is one of the at least one speakers, and the original speech of the target speaker is the speech to be interpreted simultaneously; Simultaneous interpretation is performed on the original speech of the target speaker to obtain the target speech of the target speaker. The target language speech is fused with the background audio to obtain the output audio; the output audio is the audio output to the user. The target speaker is determined based on the following method: Display speaker information set; the speaker information set includes the text content corresponding to the original speech of the at least one speaker; The target speaker is determined based on the speaker selected by the speaker selection instruction; the speaker selection instruction is an instruction triggered based on the speaker information set.
2. The delay cancellation method for simultaneous interpretation according to claim 1, characterized in that, The speaker information set also includes the identification information of the at least one speaker.
3. The delay-removal method for simultaneous interpretation according to claim 2, characterized in that, The identification information of the at least one speaker is determined based on the following method: The sound source is located in the mixed audio to obtain the spatial location features of the sound source of the at least one speaker, and the voiceprint features are extracted from the mixed audio to obtain the voiceprint features of the at least one speaker. The spatial orientation features of the sound source of the at least one speaker and the voiceprint features of the at least one speaker are input into the target association model to obtain the speech feature identification vector of the at least one speaker output by the target association model. The target association model is used to fuse the spatial location features of the sound source and the voiceprint features of the same speaker to obtain the desired result. Speech feature identifier vector; Based on the speech feature identification vector of the at least one speaker, the identification information of the at least one speaker is determined; the speech feature identification vector of the at least one speaker corresponds one-to-one with the identification information of the at least one speaker.
4. The delay-elimination method for simultaneous interpretation according to any one of claims 1 to 3, characterized in that, The step of separating the original speech of the target speaker from the mixed audio to obtain background audio that masks the original speech of the target speaker includes: Generate a mask matrix for reconstructing the original speech of the target speaker; Based on the mask matrix and the mixed audio, speech reconstruction is performed to obtain the original speech of the target speaker; The original speech of the target speaker in the mixed audio is masked to obtain the background audio.
5. The delay-removal method for simultaneous interpretation according to claim 4, characterized in that, The generation of the mask matrix for reconstructing the original speech of the target speaker includes: The mixed audio and the speech feature vector of the target speaker are input into the speech separation model to obtain the mask matrix output by the speech separation model; the speech feature vector of the target speaker is obtained by feature fusion based on the sound source spatial orientation features and voiceprint features of the target speaker; The speech separation model is obtained by training based on sample mixed audio.
6. The method for delay elimination in simultaneous interpretation according to any one of claims 1 to 3, characterized in that, The simultaneous interpretation of the original speech of the target speaker to obtain the target speech of the target speaker includes: The original speech of the target speaker is encoded to obtain speech coding features; The input features, converted from the speech coding features, are input into the speech feature generation model to obtain the speech features output by the speech feature generation model; the speech feature generation model is used to generate speech features of the target language based on the speech coding features of the original language speech; the input features are features adapted to the speech feature generation model. The speech features are decoded to obtain the target language speech of the target language.
7. A delay cancellation device for simultaneous interpretation, characterized in that, include: An audio acquisition module is used to acquire mixed audio including the original speech of at least one speaker; The speech separation module is used to separate the original speech of the target speaker from the mixed audio to obtain background audio that masks the original speech of the target speaker; The target speaker is one of the at least one speakers, and the original speech of the target speaker is the speech to be simultaneously interpreted. The simultaneous interpretation module is used to simultaneously interpret the original speech of the target speaker to obtain the target speech of the target speaker. An audio fusion module is used to fuse the target language speech with the background audio to obtain output audio; the output audio is the audio output to the user. The target speaker is determined based on the following modules: The information display module is used to display a speaker information set; the speaker information set includes the text content corresponding to the original speech of the at least one speaker; The speaker determination module is used to determine the target speaker based on the speaker indicated by the speaker selection instruction; the speaker selection instruction is an instruction triggered based on the speaker information set.
8. An electronic device comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the delay elimination method for simultaneous interpretation as described in any one of claims 1 to 6. 9.A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed by a processor, it implements the delay elimination method for simultaneous interpretation as described in any one of claims 1 to 6.
10. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the delay elimination method for simultaneous interpretation as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Simultaneous interpretation method and system, storage medium and electronic device
CN116095266A
High-quality voice conversion method based on VITS and reserved background voice
CN117037821A
Simultaneous interpretation system, method and equipment and readable storage medium
CN117119108A