Contextual Smart Switching via Multi-Modal Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing multimedia processing systems fail to automatically detect and respond to undesired content, such as advertisements or profane language, in video and audio streams, which can distract users and pose safety risks, especially during tasks like driving, as they require manual intervention that can divert attention and lead to accidents.

Innovation Solution

A contextual stream switching system that utilizes multi-modal machine learning to analyze audio and video streams, identify patterns, and predict undesired content, then autonomously switch to alternative streams or adjust volume based on user preferences and historical data, incorporating geolocation sensors to assess contextual risk factors.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual intervention is used to detect and respond to undesired content, then user control and flexibility are maintained, but user distraction and safety risks increase during critical tasks

Engineering Contradiction:
Improvesafety during critical tasksVSAvoidmanual intervention requirement
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system performs automatic detection and response to undesired content without requiring user intervention. The machine learning model autonomously identifies advertisements, profane language, and other unwanted content in audio和视频 streams, and automatically switches streams or adjusts volume, allowing the system to serve itself rather than requiring continuous user control

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual user actions with an automated machine learning-based system. Instead of requiring users to manually detect and respond to undesired content, the system uses computational algorithms to perform detection and response actions, substituting mechanical human intervention with an automated computational system

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If automated stream switching is implemented, then user distraction is reduced and safety is improved, but system complexity increases

Engineering Contradiction:
Improvesafety during critical tasksVSAvoidautomated system structure
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The machine learning model serves multiple functions: it detects various types of undesired content including advertisements, profane language, and other unwanted material across different media streams. This multi-functional approach consolidates what would otherwise require multiple separate systems into a single unified solution, managing complexity through versatility

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system introduces a machine learning model as an intermediary between the user and the multimedia streams. This intermediary automatically processes and filters content, making decisions about stream switching and volume adjustment, thereby reducing the complexity of direct user-system interaction while maintaining safety and reliability

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If multi-modal machine learning is used to analyze content, then detection accuracy is improved, but computational resource consumption increases

Engineering Contradiction:
Improveundesired content detection accuracyVSAvoidcomputational energy consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system applies multi-modal analysis selectively rather than continuously to all content. By using the machine learning model to analyze only relevant segments of audio和视频 streams where undesired content is likely to occur, the system achieves high detection accuracy while avoiding the excessive computational resource consumption that would result from analyzing every moment of every stream in detail

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12063416B2Contextual smart switching via multi-modal learning mechanism
Publication Date: 2024.08.13 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12063416B2 patent drawing
  • US12063416B2 patent drawing
  • US12063416B2 patent drawing

AI summary

The present invention may include a computer receives multimedia data. The computer parses the multimedia data into an audio stream. The computer analyzes the audio stream to identify recognized patterns. The computer calculates a probability of an undesired content based on the recognized patterns and taking an action based on determining the probability is above a threshold.