Signal Processing Apparatus for Selective Voice Suppression

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In environments like sports venues and concert halls, it is challenging to distinguish between public shouts and private conversations using existing sound capturing techniques, as both types of sounds are often recorded from the same area, leading to difficulties in determining whether a sound is private speech that needs to be masked.

Innovation Solution

A signal processing apparatus with a detection unit for voice detection, a determination unit for inter-channel similarity analysis, and a suppression unit to mask private speech based on similarity thresholds, ensuring that only private conversations are suppressed while maintaining public shouts for an enhanced sense of presence.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Illumination intensity

If sound capturing is performed in spectator stands to capture shouts and create sense of presence, then the sense of presence is enhanced, but private conversations are also captured and need to be masked

Engineering Contradiction:
Improvesense of presenceVSAvoidprivate speech exposure
Core Design Contradiction:
Illumination intensityVSObject-affected harmful factors

Solution Approach 1:

The patent segments the audio signal processing by analyzing each microphone channel individually for voice detection, then comparing signals from multiple channels to distinguish private speech from public shouts. This segmentation allows selective masking of only the harmful private conversations while preserving the beneficial public event sounds.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary analysis process that compares audio signals between multiple microphone channels. This intermediary step (similarity determination) acts as a mediator to identify whether a detected voice belongs to private conversation or public event, enabling accurate discrimination without direct human intervention.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If voice detection is performed on audio signals from multiple microphones, then private speech can be identified, but the complexity of determining similarity between signals increases

Engineering Contradiction:
Improveprivate speech detection accuracyVSAvoidsignal processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent changes the parameter of comparison from individual signal analysis to inter-channel similarity analysis. By transforming the problem into comparing audio signals across different microphone channels rather than analyzing each signal in isolation, the system achieves more accurate private speech detection with manageable processing complexity.

Inventive Principle:
Principle #35Parameter changes

3Object-affected harmful factors

If masking process is applied to all detected speech, then private conversations are protected, but public shouts are also suppressed and sense of presence is reduced

Engineering Contradiction:
Improveprivate speech protectionVSAvoidsense of presence
Core Design Contradiction:
Object-affected harmful factorsVSIllumination intensity

Solution Approach 1:

The patent applies local quality by treating different audio sources differently based on their identification. Instead of uniform masking of all detected speech, the system applies selective masking only to signals identified as private conversations while leaving public shouts unchanged. This localized approach protects privacy without degrading the audio quality of public event sounds.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11024330B2Signal processing apparatus, signal processing method, and storage medium
Publication Date: 2021.06.01 CANON KK
  • US11024330B2 patent drawing
  • US11024330B2 patent drawing
  • US11024330B2 patent drawing

AI summary

A signal processing apparatus includes a detection unit configured to perform a voice detection process on each of a plurality of audio signals captured by a plurality of microphones arranged at mutually different positions, a determination unit configured to determine a degree of similarity between two or more of the plurality of audio signals in which voice is detected by the detection unit, and a suppression unit configured to perform a process of suppressing the voice contained in at least one of the two or more audio signals, in response to a determination that the degree of similarity between the two or more audio signals is less than a threshold by the determination unit.