Background Audio Semantic Entity Detection Without Task Interruption

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users often fail to capture and explore supplemental information about topics of interest heard in various audio environments due to the inconvenience of interrupting their tasks to look up information.

Innovation Solution

A computing device with machine-learned models analyzes audio signals in the background to identify semantic entities, which are then displayed on the device's screen, allowing users to access supplemental information upon request.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If a user manually looks up supplemental information about topics heard in audio signals, then information completeness is improved, but user convenience deteriorates due to task interruption

Engineering Contradiction:
Improveinformation completenessVSAvoiduser convenience
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The system performs preliminary action by automatically capturing and analyzing audio signals to identify semantic entities and retrieve supplemental information in advance, before the user needs it. The computing device continuously monitors audio inputs, extracts meaningful entities, and prepares information displays, eliminating the need for users to manually interrupt tasks to look up information.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements self-service by autonomously performing the entire information retrieval workflow without user intervention. The computing device automatically captures audio, identifies semantic entities through machine learning models, queries databases for supplemental information, and displays results - all without requiring the user to manually search or interrupt their current tasks.

Inventive Principle:
Principle #25Self-service

2Loss of information

If audio analysis is performed continuously in the background, then information availability is improved, but computational resource consumption increases

Engineering Contradiction:
Improveinformation availabilityVSAvoidcomputational resource consumption
Core Design Contradiction:
Loss of informationVSUse of energy by moving object

Solution Approach 1:

The system applies periodic action by analyzing audio signals at specific intervals or triggered by certain conditions rather than continuously processing all audio data. The machine learning models are activated periodically or event-driven, processing audio segments only when semantic entities are detected or at predetermined time intervals, reducing overall computational load while maintaining information availability.

Inventive Principle:
Principle #19Periodic action

Solution Approach 2:

The system uses partial action by selectively analyzing only portions of audio signals that contain potential semantic entities rather than processing the entire audio stream. The machine learning models focus on detecting and analyzing relevant segments, applying computational resources only where needed to identify meaningful information while ignoring irrelevant audio portions.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12387719B2Systems and methods for identifying and providing information about semantic entities in audio signals
Publication Date: 2025.08.12 GOOGLE LLC
  • US12387719B2 patent drawing
  • US12387719B2 patent drawing
  • US12387719B2 patent drawing

AI summary

Systems and methods for determining identifying semantic entities in audio signals are provided. A method can include obtaining, by a computing device comprising one or more processors and one or more memory devices, an audio signal concurrently heard by a user. The method can further include analyzing, by a machine-learned model stored on the computing device, at least a portion of the audio signal in a background of the computing device to determine one or more semantic entities. The method can further include displaying the one or more semantic entities on a display screen of the computing device.