Non-verbal Sound Event Recognition via Frame Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current sound identification systems are ineffective in recognizing non-verbal sound events and scenes, such as a baby crying or a gun shooting, as they typically focus on transcribing speech and fail to accurately classify and distinguish non-verbal audio events and environments.

Innovation Solution

A method that processes audio data by extracting multiple acoustic features from frames of audio signals, using techniques like signal processing algorithms and regression methods, including training neural networks to classify sound classes and recognize non-verbal sound events and scenes, even in the presence of overlapping audio events.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If sound identification systems focus on transcribing speech, then speech recognition accuracy is improved, but recognition of non-verbal sound events deteriorates

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidnon-verbal sound event recognition
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The audio signal is divided into multiple frames, and each frame is independently processed to extract acoustic features. This segmentation allows the system to analyze different portions of the audio signal separately, enabling accurate identification of both speech and non-verbal sound events in their respective temporal contexts.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system employs a universal acoustic feature extraction and classification framework that handles multiple types of audio events (speech, non-verbal sounds, environmental scenes) through the same processing pipeline. The acoustic features and classification algorithms are designed to be applicable across diverse sound types, making the system multi-functional rather than specialized for only speech recognition.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If multiple acoustic features are extracted from each frame, then classification precision is improved, but processing complexity increases

Engineering Contradiction:
Improveclassification precisionVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces traditional manual feature extraction methods with neural network-based automatic feature learning. The neural network automatically learns and extracts relevant acoustic features from the raw audio frames, substituting complex manual signal processing mechanisms with a data-driven approach that achieves high classification precision while managing processing complexity through automated feature representation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If sound class scores are processed for multiple frames, then recognition accuracy of non-verbal events is improved, but processing time increases

Engineering Contradiction:
Improverecognition accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

Acoustic features are extracted and sound class scores are computed for each frame in advance, before final event recognition decisions are made. This preliminary processing of acoustic features and scoring allows the system to have ready-to-use classified data when performing temporal pattern matching across multiple frames, reducing the computational burden and time required during the final recognition stage.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11587556B2Method of recognising a sound event
Publication Date: 2023.02.21 META PLATFORMS TECHNOLOGIES LLC
  • US11587556B2 patent drawing
  • US11587556B2 patent drawing
  • US11587556B2 patent drawing

AI summary

A method for recognising at least one of a non-verbal sound event and a scene in an audio signal comprising a sequence of frames of audio data, the method comprising: for each frame of the sequence: processing the frame of audio data to extract multiple acoustic features for the frame of audio data; and classifying the acoustic features to classify the frame by determining, for each of a set of sound classes, a score that the frame represents the sound class; processing the sound class scores for multiple frames of the sequence of frames to generate, for each frame, a sound class decision for each frame; and processing the sound class decisions for the sequence of frames to recognise the at least one of a non-verbal sound event and a scene.