A method and device for semantic completion in end-to-end far-field simultaneous interpretation

By using edge microphone arrays and generative completion networks for semantic guidance, the problems of low signal-to-noise ratio and loss of speech details in edge translation devices in environments such as large exhibitions are solved, and efficient simultaneous interpretation is achieved on resource-constrained devices.

CN122135729APending Publication Date: 2026-06-02GUANGZHOU UNIVERSITY

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGZHOU UNIVERSITY
Filing Date
2026-04-01
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing edge translation devices face challenges in complex acoustic environments such as large exhibitions, including low signal-to-noise ratio in far-field sound pickup, damage to speech details by traditional noise reduction algorithms, and a contradiction between computing power and algorithm complexity, resulting in unclear hearing and inaccurate translation.

Method used

An end-side microphone array is used for adaptive beamforming and blind reverberation suppression. A generative completion network is constructed to perform frequency domain prediction using semantic feedback and to perform decoding by combining confidence scores, thereby achieving semantic completion of speech signals.

Benefits of technology

Effective suppression of far-field reverberation and noise on resource-constrained edge devices improves the accuracy and fluency of simultaneous interpretation in noisy environments, solving the problems of unclear hearing and inaccurate translation of weak far-field signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135729A_ABST
    Figure CN122135729A_ABST
Patent Text Reader

Abstract

This invention relates to the field of audio signal processing, and provides a method and device for semantic completion in edge-side far-field simultaneous interpretation, comprising: S1: acquiring the original noisy speech signal in the far-field environment through the microphone array of the edge device, and performing adaptive beamforming to lock the location of the target sound source; S2: performing time-frequency analysis on the residual speech spectrogram to detect low signal-to-noise ratio missing frame regions caused by far-field attenuation or background noise masking; S3: obtaining the semantic context vector of the translated text in the current session, and constructing a generative completion network based on semantic feedback; S4: using the residual speech spectrogram as the main input, and injecting the semantic context vector as conditional guidance information into the generative completion network; S5: inputting the reconstructed complete speech spectrogram into the edge-side translation model. This invention introduces "semantic-guided acoustic restoration," which can solve the problem of "unclear hearing and inaccurate translation" of weak far-field signals on edge devices, and improve the usability of simultaneous interpretation in noisy environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio signal processing technology, specifically to a method and device for semantic completion in end-to-end far-field simultaneous interpretation. Background Technology

[0002] With increasingly frequent international exchanges, large-scale international exhibitions, academic conferences, and transnational forums have become commonplace. In these settings, attendees often need to rely on simultaneous interpretation equipment to understand the speakers' content. Traditional simultaneous interpretation depends on human interpreters, which is not only costly but also difficult to meet the demands of large-scale concurrent access.

[0003] In recent years, automatic simultaneous interpretation (AIS) technology has become increasingly widespread. However, existing edge translation devices (such as smart translation headsets and translators) face significant challenges in real-world large-scale exhibition scenarios, exhibiting the following notable shortcomings: (1) The "signal-to-noise ratio" bottleneck of far-field sound pickup: In large exhibitions or lecture halls, the distance between the audience and the speaker is usually far (e.g., 5-20 meters), and the energy of the speech signal attenuates severely during air propagation (especially the high-frequency components carrying semantic information). At the same time, the scene is often accompanied by noisy human voices and severe room reverberation. Although existing edge devices are equipped with noise reduction algorithms, they often have difficulty extracting clear target speech from background noise when faced with weak far-field signals with low signal-to-noise ratio.

[0004] (2) “Destructive” repair of traditional noise reduction algorithms: Currently, mainstream speech enhancement algorithms (such as spectral subtraction and Wiener filtering) mainly focus on noise suppression. When processing weak far-field signals, in order to suppress strong noise, the speech spectrum is often over-smoothed, resulting in blurred formant structures and even erroneously removing weak consonants (such as / s / , / t / , / f / ). This kind of “signal damage” may be acceptable to the human ear, but it is fatal to machine translation models that rely heavily on acoustic details, and can easily lead to “missed translations” or “semantic inversions”.

[0005] (3) The contradiction between edge computing power and algorithm complexity: Existing high-performance speech restoration models (such as large models based on Diffusion) usually have huge computational requirements, making it difficult to run in real time on battery-powered edge devices (such as AR glasses and headphones). Lightweight models often have limited restoration capabilities and cannot handle the extremely complex acoustic environment of exhibitions. Therefore, there is an urgent need for a simultaneous interpretation method that can run on resource-constrained edge devices, effectively suppress far-field reverberation and noise, and "intelligently complete" missing speech details based on semantic logic, in order to solve the technical problem of "not being able to hear clearly and not being able to translate accurately" in noisy environments. Summary of the Invention

[0006] To address the aforementioned technical problems, this invention aims to provide a method and device for semantic completion in end-to-end far-field simultaneous interpretation. To solve these problems, this invention employs the following technical solution: A semantic completion method for end-to-end far-field simultaneous interpretation includes the following steps: S1: Acquire the original noisy speech signal in the far field environment through the microphone array of the end device, perform adaptive beamforming to lock the location of the target sound source, and perform blind reverberation suppression processing to obtain the residual speech spectrogram after removing reverberation interference. S2: Perform time-frequency analysis on the residual speech spectrogram to detect low signal-to-noise ratio missing frame regions caused by far-field attenuation or background noise masking; S3: Obtain the semantic context vector of the translated text in the current session and construct a generative completion network based on semantic feedback; S4: The residual speech spectrogram is used as the main input, and the semantic context vector is injected into the generative completion network as the conditional guidance information. The semantic logic is used to perform frequency domain prediction and generation of the missing frame region to reconstruct the complete speech spectrogram. S5: Input the reconstructed complete speech spectrogram into the end-side translation model, decode it by combining the confidence score of the completed region, and output the target language translation.

[0007] Preferably, in S3, obtaining the semantic context vector of the translated text in the current session includes... Extract the historical target language text output by the simultaneous interpretation system in the previous time window; A lightweight language model is used to encode historical target language text and predict the probability distribution of source language phonemes that may correspond to the currently missing speech segment. The phoneme probability distribution is mapped to a high-dimensional semantic feature vector, which serves as the semantic context vector.

[0008] Preferably, in S4, the generative completion network adopts a U-Net architecture that includes an encoder, a bottleneck layer, and a decoder; The encoder is used to extract acoustic features from the residual speech spectrogram; the bottleneck layer uses a cross-attention mechanism to fuse semantic context vectors into the acoustic features to constrain the spectral texture generated by the decoder to conform to linguistic logic; the decoder output is a complete spectrogram that fills in the missing frame regions.

[0009] Preferably, in S4, the generative completion network adopts a generative adversarial network structure.

[0010] Preferably, in S5, decoding in conjunction with the confidence score of the completed region includes: the generative completion network outputs a generated confidence map corresponding to each time-frequency unit while outputting the reconstructed spectrogram; in the attention mechanism layer of the end-side translation model, the generated confidence map is introduced as a mask weight to reduce the attention weight of low-confidence completed regions during translation decoding, thereby reducing the generation of hallucination translation.

[0011] Preferably, in S1, adaptive beamforming and blind reverberation suppression are performed jointly in the frequency domain; the late reverberation component is estimated by a weighted prediction error algorithm and subtracted from the microphone array signal, retaining the direct sound component containing the main semantic information.

[0012] Preferably, the completion method runs on edge devices with limited computing resources. Both the generative completion network and the edge translation model are processed by INT8 integer quantization. Furthermore, the generative completion network adopts a dynamic activation mechanism, which is activated only when the time proportion of the detected missing frame region exceeds a preset threshold. Otherwise, the residual speech spectrogram is directly input into the edge translation model.

[0013] Preferably, frequency domain prediction of missing frame regions is based on an auditory masking effect model, and time-frequency units with a signal-to-noise ratio lower than the auditory perception threshold are marked as regions to be repaired.

[0014] Preferably, when the edge device detects continuous low signal-to-noise ratio speech input, it automatically triggers the ultra-far field enhancement mode to reduce the decoding beamwidth of the edge translation model, prioritizing the fluency of the translation result over the richness of the vocabulary.

[0015] An end-to-end far-field simultaneous interpretation semantic completion device, used to execute the aforementioned end-to-end far-field simultaneous interpretation semantic completion method, comprising: A multi-microphone array acquisition module is used to acquire far-field speech and perform beamforming. The digital signal processing unit is used to perform blind reverberation suppression and missing frame region detection; The neural network acceleration unit is used to run generative completion networks and edge translation models, and to perform spectrogram reconstruction and translation decoding based on semantic feedback. The audio output module is used to play the target language translation.

[0016] The present invention has the following beneficial effects: This invention innovatively introduces a "semantic-guided acoustic restoration" mechanism, which can effectively solve the problem of "unclear hearing and inaccurate translation" of weak far-field signals on resource-constrained edge devices without the need to add expensive sound pickup hardware, and significantly improve the availability of simultaneous interpretation in noisy environments. Attached Figure Description

[0017] The present invention will be further described with reference to the accompanying drawings, but the embodiments in the drawings do not constitute any limitation on the present invention. For those skilled in the art, other drawings can be obtained based on the following drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of the structure of an end-side far-field simultaneous interpretation semantic completion method according to the present invention; Detailed Implementation The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] In the description of this invention, it should be noted that the terms "vertical," "upper," "lower," "horizontal," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limiting this invention.

[0020] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or a connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0021] As shown in the figure, a semantic completion method for end-to-end far-field simultaneous interpretation includes the following steps: S1: Acquire the original noisy speech signal in the far field environment through the microphone array of the end device, perform adaptive beamforming to lock the location of the target sound source, and perform blind reverberation suppression processing to obtain the residual speech spectrogram after removing reverberation interference. S2: Perform time-frequency analysis on the residual speech spectrogram to detect low signal-to-noise ratio missing frame regions caused by far-field attenuation or background noise masking; S3: Obtain the semantic context vector of the translated text in the current session and construct a generative completion network based on semantic feedback; S4: The residual speech spectrogram is used as the main input, and the semantic context vector is injected into the generative completion network as the conditional guidance information. The semantic logic is used to perform frequency domain prediction and generation of the missing frame region to reconstruct the complete speech spectrogram. S5: Input the reconstructed complete speech spectrogram into the end-side translation model, decode it by combining the confidence score of the completed region, and output the target language translation.

[0022] In a preferred embodiment of the present invention, step S3 involves obtaining the semantic context vector of the translated text in the current session, including... Extract the historical target language text output by the simultaneous interpretation system in the previous time window; A lightweight language model is used to encode historical target language text and predict the probability distribution of source language phonemes that may correspond to the currently missing speech segment. The phoneme probability distribution is mapped to a high-dimensional semantic feature vector, which serves as the semantic context vector.

[0023] In a preferred embodiment of the present invention, in S4, the generative completion network adopts a U-Net architecture that includes an encoder, a bottleneck layer, and a decoder. The encoder is used to extract acoustic features from the residual speech spectrogram; the bottleneck layer uses a cross-attention mechanism to fuse semantic context vectors into the acoustic features to constrain the spectral texture generated by the decoder to conform to linguistic logic; the decoder output is a complete spectrogram that fills in the missing frame regions.

[0024] In a preferred embodiment of the present invention, in S4, the generative completion network adopts a generative adversarial network structure.

[0025] In a preferred embodiment of the present invention, in step S5, decoding in conjunction with the confidence score of the completed region includes: the generative completion network outputs a generated confidence map corresponding to each time-frequency unit while outputting the reconstructed spectrogram; in the attention mechanism layer of the end-side translation model, the generated confidence map is introduced as a mask weight to reduce the attention weight of low-confidence completed regions during translation decoding, thereby reducing the generation of hallucination translation.

[0026] In a preferred embodiment of the present invention, in S1, adaptive beamforming and blind reverberation suppression are performed jointly in the frequency domain; the late reverberation component is estimated by a weighted prediction error algorithm and subtracted from the microphone array signal, retaining the direct sound component containing the main semantic information.

[0027] In a preferred embodiment of the present invention, the completion method runs on an edge device with limited computing resources. Both the generative completion network and the edge translation model are processed by INT8 integer quantization. The generative completion network adopts a dynamic activation mechanism, which is activated only when the time proportion of the detected missing frame region exceeds a preset threshold. Otherwise, the residual speech spectrogram is directly input into the edge translation model.

[0028] In a preferred embodiment of the present invention, frequency domain prediction of missing frame regions is based on an auditory masking effect model, and time-frequency units with a signal-to-noise ratio lower than the auditory perception threshold are marked as regions to be repaired.

[0029] In a preferred embodiment of the present invention, when the end-side device detects continuous low signal-to-noise ratio speech input, it automatically triggers an ultra-far-field enhancement mode to reduce the decoding beamwidth of the end-side translation model and prioritize the fluency of the translation result over the richness of the vocabulary.

[0030] An end-to-end far-field simultaneous interpretation semantic completion device, used to execute the aforementioned end-to-end far-field simultaneous interpretation semantic completion method, comprising: A multi-microphone array acquisition module is used to acquire far-field speech and perform beamforming. The digital signal processing unit is used to perform blind reverberation suppression and missing frame region detection; The neural network acceleration unit is used to run generative completion networks and edge translation models, and to perform spectrogram reconstruction and translation decoding based on semantic feedback. The audio output module is used to play the target language translation.

[0031] Existing simultaneous interpretation technologies often face challenges in complex acoustic environments such as large exhibitions or academic conferences, including long pickup distances, high background noise, and severe reverberation. Traditional noise reduction algorithms often destroy high-frequency details of speech while suppressing noise, resulting in missing speech signal bands and low signal-to-noise ratios. Furthermore, existing edge translation models struggle to handle incomplete speech input, leading to frequent omissions and mistranslations, which seriously affect the accuracy and continuity of cross-language communication.

[0032] To address the aforementioned problems, this invention proposes a semantic completion method for edge-side far-field simultaneous interpretation. This method first utilizes an edge-side microphone array for adaptive beamforming and blind reverberation suppression to directionally lock onto the far-field target sound source and remove environmental reverberation interference, obtaining the residual speech signal. Subsequently, a generative completion network based on semantic context feedback is constructed, using the incomplete speech spectrogram as input and employing the semantic logic of the translated text to predict and fill spectral holes caused by far-field attenuation or noise masking. Finally, the repaired complete speech features are input into the edge-side translation model, combined with confidence-weighted decoding of the completed region, to output the target language translation. Compared with existing technologies, this invention innovatively introduces a "semantic-guided acoustic restoration" mechanism, effectively solving the problem of "unclear hearing and inaccurate translation" of weak far-field signals on resource-constrained edge devices without adding expensive sound pickup hardware, significantly improving the usability of simultaneous interpretation in noisy environments.

[0033] This invention provides a method and device for semantic completion in edge-side far-field simultaneous interpretation. By employing an edge-side processing scheme of "physical enhancement followed by semantic completion", the semantic context of the translated content is used to guide the repair of the speech signal, thereby improving the usability of translation with low computational power consumption.

[0034] To address the aforementioned technical problems, this invention provides a method for semantic completion in end-to-end far-field simultaneous interpretation, comprising the following steps: Physical layer enhancement (beamforming and dereverberation): Raw noisy speech signals from the far-field environment are acquired using a microphone array on the end-side device. First, adaptive beamforming techniques (such as the MVDR algorithm) are used to spatially locate the target speaker, suppressing interference noise from the sides and back. Next, blind reverberation suppression processing (such as the weighted prediction error (WPE) algorithm) is performed to subtract late reverberation components from the received signal, obtaining a residual speech spectrogram that preserves direct sound characteristics.

[0035] Missing region detection: Fine-grained time-frequency analysis is performed on the residual speech spectrogram. Based on the auditory masking effect model or a preset signal-to-noise ratio threshold, "low signal-to-noise ratio missing frame regions" that are lost due to far-field transmission attenuation or transient noise (such as applause, screams) are identified.

[0036] Semantic-guided intelligent completion: Construct a generative completion network (such as a lightweight U-Net or GAN) running on the edge. Unlike traditional repair methods that rely solely on acoustic features, this invention introduces a "semantic feedback mechanism": Obtain the translated text or source language recognition results output by the simultaneous interpretation system in the previous time window; use a lightweight language model to encode these historical texts into semantic context vectors, predict the phoneme probabilities that the currently missing speech segment may correspond to in linguistic logic; inject this semantic context vector as conditional guidance information into the completion network, guiding the network to "draw" the missing spectral texture in the frequency domain.

[0037] Confidence-weighted translation: The reconstructed complete speech spectrogram is input into the edge translation model. To prevent the completion network from producing erroneous speech (illusion), the completion network simultaneously outputs a "confidence map." During decoding, the translation model reduces the attention weight of the completed region based on this confidence map, prioritizing the use of high-confidence real signal regions, thereby outputting an accurate and fluent target language translation.

[0038] The components, modules, mechanisms, and devices in this invention that are not described in detail are all general standard parts or components known to those skilled in the art. Their structures and principles can be learned by those skilled in the art through technical manuals or conventional experimental methods.

[0039] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit the scope of protection of the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the essence and scope of the technical solutions of the present invention.

Claims

1. A semantic completion method for end-to-end far-field simultaneous interpretation, characterized in that, Includes the following steps: S1: Acquire the original noisy speech signal in the far field environment through the microphone array of the end device, perform adaptive beamforming to lock the location of the target sound source, and perform blind reverberation suppression processing to obtain the residual speech spectrogram after removing reverberation interference. S2: Perform time-frequency analysis on the residual speech spectrogram to detect low signal-to-noise ratio missing frame regions caused by far-field attenuation or background noise masking; S3: Obtain the semantic context vector of the translated text in the current session and construct a generative completion network based on semantic feedback; S4: The residual speech spectrogram is used as the main input, and the semantic context vector is injected into the generative completion network as the conditional guidance information. The semantic logic is used to perform frequency domain prediction and generation of the missing frame region to reconstruct the complete speech spectrogram. S5: Input the reconstructed complete speech spectrogram into the end-side translation model, decode it by combining the confidence score of the completed region, and output the target language translation.

2. The method for semantic completion in end-to-end far-field simultaneous interpretation according to claim 1, Its characteristic is that, in S3, obtaining the semantic context vector of the translated text in the current session includes... Extract the historical target language text output by the simultaneous interpretation system in the previous time window; A lightweight language model is used to encode historical target language text and predict the probability distribution of source language phonemes that may correspond to the currently missing speech segment. The phoneme probability distribution is mapped to a high-dimensional semantic feature vector, which serves as the semantic context vector.

3. The method for semantic completion in end-to-end far-field simultaneous interpretation according to claim 1, Its characteristic is that, in S4, the generative completion network adopts a U-Net architecture that includes an encoder, a bottleneck layer, and a decoder; The encoder is used to extract acoustic features from the residual speech spectrogram; the bottleneck layer adopts a cross-attention mechanism to fuse semantic context vectors into the acoustic features, so as to constrain the spectral texture generated by the decoder to conform to linguistic logic. The decoder output fills in the missing frame region with a complete spectrogram.

4. The method for semantic completion in end-to-end far-field simultaneous interpretation according to claim 1, characterized in that, In S4, the generative completion network adopts a generative adversarial network structure.

5. The method for semantic completion in end-to-end far-field simultaneous interpretation according to claim 1, characterized in that, In S5, decoding by combining the confidence score of the completed region includes: the generative completion network outputs the generated confidence map corresponding to each time-frequency unit while outputting the reconstructed spectrogram; in the attention mechanism layer of the end-side translation model, the generated confidence map is introduced as a mask weight to reduce the attention weight of low-confidence completed regions during translation decoding, so as to reduce the generation of hallucination translation.

6. The method for semantic completion in end-to-end far-field simultaneous interpretation according to claim 1, characterized in that, In S1, adaptive beamforming and blind reverberation suppression are performed jointly in the frequency domain; the late reverberation component is estimated by a weighted prediction error algorithm and subtracted from the microphone array signal, retaining the direct sound component containing the main semantic information.

7. The method for semantic completion in end-to-end far-field simultaneous interpretation according to claim 1, characterized in that, The completion method runs on edge devices with limited computing resources. Both the generative completion network and the edge translation model are processed by INT8 integer quantization. The generative completion network adopts a dynamic activation mechanism, which is activated only when the time proportion of the detected missing frame region exceeds a preset threshold. Otherwise, the residual speech spectrogram is directly input into the edge translation model.

8. The method for semantic completion in end-to-end far-field simultaneous interpretation according to claim 1, characterized in that, Frequency domain prediction of missing frame regions is based on the auditory masking effect model, and time-frequency units with a signal-to-noise ratio lower than the auditory perception threshold are marked as regions to be repaired.

9. The method for semantic completion in end-to-end far-field simultaneous interpretation according to claim 1, characterized in that, When the edge device detects continuous low signal-to-noise ratio speech input, it automatically triggers the ultra-far field enhancement mode, reduces the decoding beamwidth of the edge translation model, and prioritizes the fluency of the translation result rather than the richness of the vocabulary.

10. A semantic completion device for end-to-end far-field simultaneous interpretation, characterized in that, A method for performing semantic completion of end-side far-field simultaneous interpretation as described in any one of claims 1-9 includes: A multi-microphone array acquisition module is used to acquire far-field speech and perform beamforming. The digital signal processing unit is used to perform blind reverberation suppression and missing frame region detection; The neural network acceleration unit is used to run generative completion networks and edge translation models, and to perform spectrogram reconstruction and translation decoding based on semantic feedback. The audio output module is used to play the target language translation.