Environmental Sound Recognition via Server-Side Model Aggregation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Mobile devices often struggle to accurately recognize environmental sounds due to limited storage capacity and processing power, and fail to distinguish between similar sound environments, leading to inaccurate sound recognition in new surroundings.

Innovation Solution

A system and method where client devices communicate with a server to share and aggregate sound models, using databases with sound models and labels to generate and match input sound models, improving recognition accuracy through crowd-sourcing and confidence level determination.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a mobile device stores and processes sound environmental data locally, then it can recognize ambient sounds, but the device's limited storage capacity and processing power prevent accurate recognition of a broad scope of environmental sounds

Engineering Contradiction:
Improvesound recognition accuracyVSAvoidstorage capacity
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The system divides the sound recognition functionality between the mobile device (client) and the remote server. The client device captures ambient sound and generates a sound model, while the server stores a comprehensive database of sound models and performs the complex matching operations. This segmentation allows the device to maintain lightweight local storage while achieving accurate recognition through server resources.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a server as an intermediary between the mobile device and the sound recognition database. The server acts as a mediator that receives sound models from multiple devices, aggregates them into a comprehensive database, and provides recognition services back to clients. This intermediary enables devices to access broad scope sound recognition without requiring large local storage capacity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If a mobile device recognizes only ambient sounds of its surroundings, then it can provide localized sound recognition, but the device fails to accurately recognize ambient sounds of new surroundings

Engineering Contradiction:
Improvesound environment adaptabilityVSAvoidsound recognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The system performs preliminary actions by pre-collecting and aggregating sound models from multiple sources and locations into the server database before the client device needs recognition services. This pre-aggregation of diverse environmental sound data enables the device to adapt to new surroundings accurately without requiring retraining or additional data collection when moving to new environments.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The server database is designed to store a universal collection of sound models representing diverse environmental sounds from multiple locations and contexts. This universal database serves all client devices regardless of their specific location or environment, enabling any device to recognize sounds from new surroundings with high accuracy by comparing against the comprehensive aggregated database.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If mobile devices operate independently without sharing data, then they maintain device autonomy, but they cannot distinguish and accurately recognize different sound environments

Engineering Contradiction:
Improvesound environment differentiation accuracyVSAvoidsystem architecture complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges sound models from multiple independent devices into a unified server database. Each device contributes its collected sound data to the server, which aggregates and combines these models into a comprehensive database. This merging enables all devices to benefit from the collective data, improving their ability to differentiate between various sound environments while maintaining their operational independence.

Inventive Principle:
Principle #5Merging (Combining)

4Measurement precision

If a device uses a local database of sound models, then it can perform sound recognition, but the limited database scope reduces recognition accuracy for diverse environmental sounds

Engineering Contradiction:
Improvesound recognition accuracyVSAvoidsound model diversity
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system creates copies of sound models from multiple sources and stores them in the server database. Instead of each device maintaining a unique limited database, the server aggregates copies of sound models from various environments and devices, creating a diverse comprehensive database that can be accessed by all clients, thereby improving both accuracy and adaptability.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS9443511B2System and method for recognizing environmental sound
Publication Date: 2016.09.13 QUALCOMM INC
  • US9443511B2 patent drawing
  • US9443511B2 patent drawing
  • US9443511B2 patent drawing

AI summary

A method for recognizing an environmental sound in a client device in cooperation with a server is disclosed. The client device includes a client database having a plurality of sound models of environmental sounds and a plurality of labels, each of which identifies at least one sound model. The client device receives an input environmental sound and generates an input sound model based on the input environmental sound. At the client device, a similarity value is determined between the input sound model and each of the sound models to identify one or more sound models from the client database that are similar to the input sound model. A label is selected from labels associated with the identified sound models, and the selected label is associated with the input environmental sound based on a confidence level of the selected label.