Audio Signal Conversion Using Deep Neural Networks for Directional Isolation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Non-directional microphones in smartphones struggle to effectively isolate sound from a specific direction, often capturing significant surrounding noise along with the desired sound.

Innovation Solution

An audio signal processing apparatus that utilizes a deep neural network and convolutional neural networks to convert audio signals from non-directional microphones into unidirectional signals, allowing for the selective collection of sound from either the front or back direction by recognizing sound direction and adjusting the output accordingly.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a non-directional microphone is used to collect sound, then the microphone can capture sound from all directions, but surrounding noise is also collected at a large level

Engineering Contradiction:
Improveomnidirectional sound collectionVSAvoidsurrounding noise
Core Design Contradiction:
Adaptability or versatilityVSObject-affected harmful factors

Solution Approach 1:

The patent replaces the mechanical/directional approach of traditional microphones with a signal processing system using deep neural networks. Instead of relying on physical microphone directionality, the system uses AI algorithms to analyze audio signals and extract directional information, substituting mechanical design with intelligent software processing to achieve noise reduction while maintaining omnidirectional capture capability

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the parameters of audio signal processing by using deep neural networks to transform the audio signal characteristics. The system adjusts frequency, time, and spatial parameters through AI-based processing to separate desired sound from noise, dynamically modifying signal parameters to achieve directional selectivity from omnidirectional input

Inventive Principle:
Principle #35Parameter changes

2Object-affected harmful factors

If a unidirectional microphone is used to collect sound from a specific direction, then surrounding noise is reduced, but the device complexity increases

Engineering Contradiction:
Improvesurrounding noise reductionVSAvoidmicrophone system complexity
Core Design Contradiction:
Object-affected harmful factorsVSDevice complexity

Solution Approach 1:

The patent makes the non-directional microphone system multi-functional by integrating deep neural network processing that enables both omnidirectional capture and directional noise reduction. The same hardware configuration serves multiple purposes: capturing all-direction sound when needed and selectively filtering directional noise when needed, eliminating the requirement for separate unidirectional microphone systems

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent creates a virtual unidirectional microphone effect through software by using deep neural networks to synthesize directional audio characteristics from omnidirectional input. Instead of physically copying unidirectional microphone hardware, the system digitally replicates the directional selectivity function through AI-based signal transformation

Inventive Principle:
Principle #26Copying

3Object-affected harmful factors

If multiple microphones are used to achieve directional sound collection, then sound isolation improves, but the ease of operation decreases

Engineering Contradiction:
Improvesound isolationVSAvoiduser configuration simplicity
Core Design Contradiction:
Object-affected harmful factorsVSEase of operation

Solution Approach 1:

The patent enables the audio system to automatically determine and adjust directional parameters without user intervention. The deep neural network autonomously analyzes the acoustic environment, identifies sound sources, and configures optimal directional filtering parameters, making the system self-adjusting and eliminating the need for users to manually configure multiple microphones or adjust settings

Inventive Principle:
Principle #25Self-service

4Measurement precision

If traditional signal processing methods are used to convert audio signals, then the processing speed is fast, but the measurement precision of sound direction is insufficient

Engineering Contradiction:
Improvesound direction recognition accuracyVSAvoidsignal processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary training of deep neural networks offline before actual use. During this pre-processing phase, the system learns optimal directional recognition patterns and parameters from training data. When deployed, the pre-trained model rapidly processes audio signals with high precision, as the complex learning and parameter optimization work was already completed in advance, reducing real-time processing time while maintaining accuracy

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240276147A1Audio signal processing apparatus, audio signal processing method, and electronic device
Publication Date: 2024.08.15 SONY GROUP CORP
  • US20240276147A1 patent drawing
  • US20240276147A1 patent drawing
  • US20240276147A1 patent drawing

AI summary

The present technique makes it possible to satisfactorily collect sound coming from a predetermined direction using a non-directional microphone.An audio signal conversion unit converts an audio signal obtained by collecting sound by the non-directional microphone into a unidirectional audio signal. For example, the audio signal conversion unit is configured with a deep neural network. In this case, for example, the deep neural network is trained to learn to minimize a difference between an acoustic feature amount extracted from an audio signal converted by the deep neural network and an acoustic feature amount extracted from a unidirectional audio signal obtained by collecting sound by a unidirectional microphone.