Media Content Steering via Comparative Voice Intent Parsing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing media playback devices and streaming services lack the ability to interpret voice commands containing comparative-type instructions, such as 'more' or 'less,' leading to inaccurate recommendations and inefficient searches, which result in wasted computing power and poor user experience.

Innovation Solution

A media content steering system that identifies and processes voice commands to provide media content items with attributes different from the currently playing content, by parsing utterances into intents and slots, and retrieving media content items with facet types and values that match the user's request, ensuring relative differences within constrained attributes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If existing media playback devices and streaming services use simplistic voice command interpretation, then device complexity is reduced, but measurement precision of user intent deteriorates

Engineering Contradiction:
Improvevoice command processing systemVSAvoiduser intent recognition
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The voice command processing system is segmented into multiple specialized components: an utterance parser that divides commands into intents and slots, a constraint extractor that identifies comparative instructions, a media content analyzer that examines current playback attributes, and a recommendation generator that synthesizes results. This segmentation allows each component to specialize in specific tasks, improving overall precision without proportionally increasing overall system complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediary data structures (intents, slots, and constraints) that mediate between the raw voice command and the media content selection. These intermediaries transform unstructured speech into structured representations that can be systematically processed, bridging the gap between simple device architecture and sophisticated intent recognition capabilities.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If the system performs multiple searches to find satisfactory media content, then reliability of user satisfaction improves, but productivity of the system deteriorates

Engineering Contradiction:
Improveuser satisfactionVSAvoidsearch efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system performs preliminary analysis of the voice command to extract constraints and comparative instructions before conducting the media content search. By pre-processing the user intent and identifying specific attributes to modify (e.g., tempo, mood, genre), the system narrows the search space upfront, ensuring reliable user satisfaction with a single targeted search rather than multiple trial searches.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system incorporates feedback loops where the current media content attributes are analyzed and compared against the extracted constraints. This feedback mechanism allows the recommendation generator to adjust selections based on the relationship between current playback and user preferences, improving reliability while reducing the need for iterative searches.

Inventive Principle:
Principle #23Feedback

3Ease of operation

If the system returns static playlists that are not personalized, then ease of operation is improved, but adaptability to user preferences deteriorates

Engineering Contradiction:
Improvemedia content deliveryVSAvoidpersonalization capability
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The system transitions from static playlist delivery to dynamic content selection by incorporating real-time analysis of user voice commands and current playback state. The recommendation generator dynamically adjusts media content selections based on extracted constraints and comparative instructions, allowing the system to adapt to user preferences while maintaining simple operation through natural voice interaction.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes parameters of media content selection by extracting constraints from voice commands that specify desired modifications to current playback attributes (e.g., 'more upbeat,' 'slower tempo'). These parameter changes enable personalization without requiring complex user interfaces, as the voice command naturally expresses the desired parameter adjustments.

Inventive Principle:
Principle #35Parameter changes

4Measurement precision

If the system interprets comparative-type instructions accurately, then measurement precision of media content attributes improves, but device complexity increases

Engineering Contradiction:
Improvemedia content attribute matchingVSAvoidvoice command processing system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The complex task of interpreting comparative instructions is segmented into specialized sub-tasks: parsing utterances into intents and slots, extracting constraints from comparative keywords, analyzing current media content attributes, and generating recommendations that satisfy constraints. This segmentation improves measurement precision for each sub-task while distributing complexity across modular components.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates universal data structures (intents, slots, constraints) that can handle various types of comparative instructions across different media content attributes. This multi-functionality allows a single processing framework to accurately interpret diverse user preferences (tempo, mood, genre, energy) without requiring separate specialized systems for each attribute type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12057114B2Media content steering
Publication Date: 2024.08.06 SPOTIFY
  • US12057114B2 patent drawing
  • US12057114B2 patent drawing
  • US12057114B2 patent drawing

AI summary

A media content steering solution is provided to identify a user query to steer playback of media content that is currently playing or has been played. The user steering query can include a voice request for playing media content that is relatively different from the media content being currently played or having been played. The media content steering solution analyzes the utterance of the user query and uses it to identify such different content that satisfies the user intent contained in the user query.