Voice silencing system for enhanced privacy during telecommunication

The voice-privacy system addresses voice overheard risks in public spaces by using near-field cancellation and array interference, achieving significant SPL reduction for enhanced privacy.

WO2026093927A1PCT designated stage Publication Date: 2026-05-07AN SOFTWARE GMBH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
AN SOFTWARE GMBH
Filing Date
2025-10-29
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Voice interactions in public spaces are susceptible to being overheard due to latency issues in commodity devices and the complexity of hardware-assisted platforms with microphone and speaker arrays.

Method used

A voice-privacy system utilizing near-field cancellation and array-based destructive interference, with a predictor estimating future speech portions to align emissions, and employing echo-cancellation to protect the recorded voice channel, suitable for both software-only and hardware-assisted implementations.

Benefits of technology

Effectively minimizes acoustic pressure within a local control region and reduces audibility for bystanders, achieving at least 5-15 dB SPL reduction and ensuring privacy during voice interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000010_0000
    Figure 00000010_0000
  • Figure 00000011_0000
    Figure 00000011_0000
  • Figure 00000012_0000
    Figure 00000012_0000
Patent Text Reader

Abstract

A voice-privacy system configured to reduce audibility of a user's speech by emitting a predictive anti-voice signal. A microphone captures a voice signal x(t); a processing circuit predicts future samples over a prediction horizon and computes a phase-aligned, gain-matched inversion â(t). At least one output transducer emits â(t) to reduce sound pressure level at bystander locations. An error microphone may sense residual acoustic energy and enable computation of a cancellation-efficiency metric used by a controller to adapt emission parameters for improved privacy. Embodiments include acoustic echo cancellation, safety limiting and comfort-noise, and adaptive switching to multi-transducer destructive-interference configurations. In some embodiments, the prediction horizon is selected responsive to latency, including values less than 10 ms when device latency permits, and may be greater (e.g., tens to hundreds of milliseconds) depending on measured latency.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] VOICE SILENCING SYSTEM FOR ENHANCED PRIVACY DURING TELECOMMUNICATION

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS

[0003] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 713,090, filed October 29, 2024. The entire contents of the provisional application are incorporated by reference herein.

[0004] TECHNICAL FIELD

[0005] The disclosure relates to active acoustic control for voice privacy, including near-field cancellation and array-based destructive interference, in both software-only and hardware-assisted implementations.

[0006] BACKGROUND

[0007] Voice interactions in public spaces risk being overheard. Commodity devices (e.g., smartphones, smartglasses) provide microphones and speakers but often incur record / playback latencies on the order of tens to -100 ms through the operating-system audio stack, which complicates time-aligned anti-voice emission. Hardware-assisted platforms with microphone and speaker arrays can provide improved spatial control at the cost of additional components.

[0008] SUMMARY

[0009] A voice-privacy system operates in two complementary modes. In a near-field mode, one or more output transducers positioned proximate to the user's mouth emit an anti-voice signal time-aligned and gain-matched to minimize acoustic pressure within a local control region. In an array mode, two or more output transducers are driven with complex weights to produce a beam pattern having one or more nulls steered toward bystanders. A predictor estimates a future portion of the user's speech with a horizon selected responsive to measured end-to-end latency, including values less than 10 ms when device latency permits and, in some embodiments, greater values depending on measured latency enabling aligned emission despite processing and propagation delays. Echo-cancellation protects the recorded voice channel. Implementations include software-only deployments on smartphones and smartglasses using built-in transducers and OS audio stacks, and hardware-assisted platforms employing microphone / speaker arrays and dedicated processors.

[0010] BRIEF DESCRIPTION OF THE DRAWINGS

[0011] FIG. 1 is a system-level block diagram including microphone(s), a processing unit including a controller, and output(s).

[0012] FIG. 2 is a signal-processing pipeline for capture, configurable prediction horizon H(t), inversion / EQ, latency compensation, safety limiting, comfort-noise insertion, and emission.

[0013] FIG. 3 is an alignment illustration showing the prediction horizon H(t) selected to counter device end-to-end latency such that emission is substantially concurrent (e.g., H(t) less than about 10 ms when supported by platform latency).

[0014] FIG. 4A is a specialized acoustic echo cancellation (AEC) architecture for a single capture channel.

[0015] FIG. 4B illustrates a multi-input / multi-output (MIMO) AEC for arrays of emitters and microphones.

[0016] FIG. 5 is a control diagram for a self-tuning loop that minimizes residual acoustic energy using feedback r(t) from an error microphone.

[0017] FIG. 6 illustrates example device contexts and output-transducer selection.

[0018] FIG. 7 depicts array-based destructive-interference with a beamformer / controller computing weights to steer one or more nulls toward a bystander region while preserving a passband toward capture microphones.

[0019] FIG. 8 illustrates a near-field control region and an example single-transducer cancellation embodiment; an optional error microphone (not shown) may be positioned as described herein.

[0020] FIG. 9 illustrates a software-only smartphone / smartglasses implementation path through an OS audio stack, including a latency estimator and a controller that adjusts prediction horizon and emission timing responsive to measured end-to-end latency.

[0021] FIG. 10 illustrates a hardware-assisted array with a dedicated signal-processing circuit (e.g., a DSP) and low-latency links. DETAILED DESCRIPTION

[0022] Definitions. As used herein, an "anti-voice signal" may be denoted a(t) and is also referred to as \bar{x}(t); the two notations are used interchangeably. A "near-field control region" is a bounded volume (e.g., within a few centimeters, and in some embodiments within up to about 10 cm, of the user's mouth) in which the system seeks to minimize acoustic pressure or a weighted measure thereof; "bystander location" is a region outside the control region where audibility reduction is desired. As used herein, a "processing circuit" broadly includes any combination of hardware and software that performs the described signal processing, including but not limited to CPUs, GPUs, DSPs, and NPUs, whether general-purpose or dedicated. A "controller" is logic (hardware, software, or a combination) that schedules emission with a prediction horizon and adaptively sets gain and delay so that sound pressure level measured at a bystander location is reduced relative to no emission; in near-field embodiments, such controller may be referred to as a "near-field controller."

[0023] System Overview (FIG. 1). One or more microphones capture the user's speech x(t) and provide it to a processing unit. The processing unit generates a real-time anti-voice signal \bar{x}(t) and implements a controller that schedules emission with a prediction horizon and adaptively sets gain and delay. The controller drives one or more output transducers positioned proximate to the mouth in a near-field mode, and can also configure an array mode when multiple emitters are available. A mode selector, predictor, and alignment logic (not separately depicted in FIG. 1) reside within the processing unit and apply per-mode delay / gain calibration using stored profiles.

[0024] Near-Field Mode (FIG. 8). One or more output transducers positioned proximate to the mouth emit \bar{x}(t+\Delta). In some embodiments a single transducer reduces hardware complexity while achieving cancellation within the local control region; in other embodiments multiple closely located transducers improve coupling, bandwidth, or robustness while remaining within a near-field operating regime. An error microphone, optionally placed inside, near, or outside the control region (not shown in FIG. 8 for clarity), provides residual measurements r(t) for adaptive updates. Optional acoustic ducting / baffling increases coupling. Array-Based Destructive-Interference Mode (FIG. 7). An array of N>2 emitters outputs signals s_i(t) = w_i * \bar{x}(t) for complex weights w_i computed by a beamformer / controller (e.g., MVDR / LCMV / GSC) from geometry and bystander-direction estimates. As depicted in FIG. 7, the weights are applied to the emitter array to steer one or more nulls toward a bystander region while preserving a passband oriented toward capture microphones as needed.

[0025] Prediction and Alignment (FIGS. 2 3). To overcome device and propagation delays, the predictor estimates a future portion of the user's speech with a horizon H(t) sufficient to compensate record / playback latency in the deployment platform. Depending on measured end-to-end latency, H(t) may be less than about 10 ms when platform latency permits, or greater (e.g., tens to about 100 ms) on higher-latency devices. Alignment logic schedules emission accordingly.

[0026] Acoustic Echo Cancellation (FIGS. 4A 4B). In near-field mode, a single-input / single-output AEC (FIG. 4 A) uses the emitted anti -voice or driver signal as a reference and subtracts an estimated echo from the recorded mic signal v(t) to produce a clean recorded voice channel e(t); the error e(t) adapts the filter (e.g., LMS / NLMS / RLS). When multiple near-field emitters are used, a multi-input AEC is employed; in array configurations, a MIMO AEC (FIG. 4B) models paths from each emitter to each capture microphone and outputs clean recorded channels, protecting the recorded voice from contamination by anti-voice emissions. In some embodiments, separate coefficient sets model near-field emission coupling versus environmental echo paths.

[0027] Self-Tuning and Mode Selection (FIG. 5). As shown in FIG. 5, an error microphone provides a residual signal r(t) used by the controller to adapt delay and gain, and a comparator / estimator compares the emitted or predicted anti-voice with the captured voice and / or residual to generate timing / gain / prediction updates and a cancellation-efficiency metric (e.g., a ratio or difference of SPLs). A controller minimizes residual energy measured in the control region and / or at intended null locations, schedules emission with a prediction horizon, adaptively sets gain and delay to reduce SPL at bystander locations, updates beamformer weights or near-field timing / gain, and selects or blends modes based on device context (e.g., headset active, handset distance, available emitters). Safety Limiting and Comfort Noise. To maintain safe operation and reduce perceptible artifacts to bystanders, the processing circuit may limit emitted SPL according to a safety threshold and inject a shaped comfort-noise floor when deep cancellation would otherwise create unnatural quieting or pumping.

[0028] Personalized Equalization. In some embodiments, the processing circuit blends the anti-voice signal with a personalized equalization profile derived from the user's head-related transfer function (HRTF) to improve coupling and perceptual masking while maintaining cancellation efficacy.

[0029] Implementation Notes. The processing circuit may utilize fixed-point arithmetic on resource-constrained devices to reduce power and memory usage while meeting real-time constraints; alternatively or additionally, general-purpose processors (CPU / GPU / NPU) may execute portions of the pipeline exposed by an operating system.

[0030] Performance Metrics. Performance can be quantified by: (i) SPL reduction at a defined location (e.g., at least about 5 15 dB reduction at a point 1 m from the user in representative conditions), and / or (ii) a cancellation-efficiency metric computed as a ratio or difference between a reference SPL and a residual SPL measured by an error microphone or derived from system sensors. These metrics can be used by a controller to adapt parameters toward improved privacy.

[0031] Transparency / Interlock. A safety interlock may disable emission in response to regulatory constraints or user selection of a transparency mode.

[0032] Computer-Readable Media. The functionality described herein may be embodied in non-transitory computer-readable media storing instructions that, when executed by one or more processors of the processing circuit, cause performance of methods disclosed herein, including prediction, inversion, scheduling, AEC, beamforming, safety limiting, and self-tuning.

[0033] Software-Only Implementations (FIG. 9). The system executes on a smartphone or smartglasses using built-in microphones and speakers accessed through an operating-system audio stack. As shown in FIG. 9, (i) a latency estimator measures capture-to-playback end-to-end latency (e.g., via timestamps from the OS audio path and driver), and a controller uses the measured latency to dynamically adjust the prediction horizon H(t) and emission timing; H(t) may be limited by a configured upper bound (for example, less than about 10 ms when platform latency permits); and (ii) the recorded mic stream is processed by AEC using the emission drive as a reference to output a clean recorded voice channel to the app / OS.

[0034] Hardware- Assisted Implementations (FIG. 10). The system employs microphone and speaker arrays driven by a dedicated signal-processing circuit (e.g., a DSP) with low-latency digital audio interfaces an example of the processing circuit described herein; distributed arrays may span multiple devices (earbuds, handset, wearable).

[0035] INDUSTRIAL APPLICABILITY

[0036] The system enables private voice interaction in public or shared spaces for calls, voice messages, and Al assistant queries.

[0037] EMBODIMENT NOTES

[0038] Features described for one mode may be combined with the other unless context dictates otherwise. "Comprising" is used inclusively.

Claims

CLAIMS1. A voice-privacy system configured to reduce audibility of a user's voice by emitting an anti-voice signal using at least one output transducer, comprising:(a) at least one microphone positioned to capture a user's utterances and generate a captured voice signal x(t);(b) a processing circuit configured to generate, in real time, an anti-voice signal a(t) that is a predictive, phase-aligned, and gain-matched inversion of at least a portion of x(t);(c) at least one output transducer configured to emit a(t); and(d) a controller configured to schedule emission of a(t) with a prediction horizon of less than 10 ms and to adaptively set gain and delay so that a sound pressure level (SPL) of the user's voice measured at a bystander location is reduced relative to speaking without emission of a(t), wherein the at least one microphone is operatively coupled to provide x(t) to the processing circuit, the processing circuit is operatively coupled to provide a(t) to the controller, and the controller is operatively coupled to drive the at least one output transducer to emit a(t).

2. The system of claim 1, wherein the controller dynamically adjusts said prediction horizon and emission timing responsive to measured end-to-end latency while maintaining the less-than-10-ms horizon.

3. The system of any of claims 1 2, further comprising an acoustic echo cancellation (AEC) subsystem configured to prevent contamination of a recorded voice channel by the emitted anti-voice signal.

4. The system of any of claims 1 3, wherein the controller implements a self-tuning mechanism that compares the emitted or predicted anti-voice signal with the captured voice signal to update timing, gain, or prediction parameters to improve cancellation.

5. The system of any of claims 1 4, wherein the controller selects which available speaker of a device to emit a(t) based on a context including device configuration or proximity so as to improve cancellation effectiveness.

6. A method of reducing audibility of a user's voice, comprising: capturing a voice signal x(t); predicting future samples of x(t) over a prediction horizon of less than 10 ms; computing an anti-voice signal a(t) as a phase-aligned and gain-matched inversion; and emitting a(t) from at least one transducer so as to reduce SPL at a bystander location while protecting a recorded voice channel using acoustic echo cancellation.

7. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause performance of the method of claim 6.

8. The system of any of claims 1 5, further comprising an error microphone positioned outside a near-field control region, the processing circuit being configured to adapt delay and gain by minimizing a residual r(t) measured by an error microphone.

9. The system of any of claims 1 5, wherein a near-field control region is defined as a bounded volume within 10 cm of the user's mouth.

10. A voice-privacy system comprising at least two output transducers and a controller configured to drive the at least two output transducers with weights to produce destructive interference at a bystander location.

11. The system of claim 10, wherein the controller employs a beamformer selected from minimum variance distortionless response (MVDR), linearly constrained minimum variance (LCMV), or generalized sidelobe canceller (GSC) architectures to steer a null toward a bystander region.

12. The system of any of claims 1 5 and 10 11, wherein end-to-end latency is maintained sufficiently low to enable emission of the anti-voice signal substantially concurrently with the user’s utterance.

13. The system of any of claims 1 5 and 10 12, wherein the processing circuit limits emitted SPL to comply with a safety threshold and applies a comfort-noise floor to reduce perceptible artifacts to bystanders.

14. The system of any of claims 1 5 and 10 13, wherein the processing circuit blends the anti-voice signal with a personalized equalization profile derived from the user's head-related transfer function (HRTF).

15. The system of any of claims 1 5 and 10 14, wherein the processing circuit comprises at least one of a CPU, GPU, DSP, or NPU, and uses fixed-point arithmetic for low-power operation.

16. The system of any of claims 1 5 and 10 15, wherein the AEC subsystem uses a multi-path adaptive filter with separate coefficients for near-field emission and environmental echoes.

17. The system of any of claims 1 5 and 10 16, further comprising a safety interlock that disables emission responsive to regulatory constraints or user selection of a transparency mode.

18. The system of any of claims 1 5 and 10 17, wherein SPL measured at a position 1 m from the user is reduced by at least 10 dB relative to no emission of a(t).

19. The system of any of claims 8 18, wherein a controller computes a cancellation-efficiency metric based on a ratio or difference between a reference sound pressure level and a residual sound pressure level measured by an error microphone, and adaptively adjusts emission parameters to maximize said efficiency metric.

20. A method comprising: capturing a voice signal x(t); computing an anti-voice signal a(t); and driving at least two output transducers with weights to produce destructive interference that reduces SPL at a bystander location, and optionally computing or updating beamformer weights using MVDR, LCMV, or GSC techniques.

21. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause performance of the method of claim 20.

Citation Information

Patent Citations

  • Sensor Fusion to Improve Speech / Audio Processing in a Mobile Device

    US20130332156A1

  • Adaptive speech intelligibility control for speech privacy

    US20210183402A1