Media Object Identification via Multi-Modal Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies lack the ability to effectively identify objects, particularly people, in media such as video, images, and audio, and fail to provide relevant information about these objects.

Innovation Solution

A device and method that utilize image/video recognition and audio recognition to identify objects in media, displaying identification information, with features like facial and voice recognition, and the ability to capture and process media files, providing biographical information, links, and recommendations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If facial recognition technology is used to verify access to buildings and computers, then security verification is improved, but the ability to identify unknown individuals in crowded places deteriorates

Engineering Contradiction:
Improvesecurity verificationVSAvoididentification in crowded places
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent applies multi-functionality by integrating multiple recognition technologies (facial recognition, voice recognition, object recognition) into a single media identification system. This allows the system to adapt to different identification scenarios - using facial recognition for verified access control and voice recognition for identifying unknown individuals in crowded places, thus resolving the contradiction between security verification reliability and adaptability to different environments

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If current facial recognition technology is used, then verification of known individuals is improved, but identification of all objects in media deteriorates

Engineering Contradiction:
Improvefacial recognition accuracyVSAvoidobject identification scope
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent merges multiple recognition technologies (facial recognition, voice recognition, and general object recognition) into a unified media identification system. This combination allows the system to maintain high facial recognition accuracy for known individuals while simultaneously expanding capability to identify various objects including unknown persons, places, and things in media, thus resolving the contradiction between precision and scope

Inventive Principle:
Principle #5Merging (Combining)

3Measurement precision

If Song IDentity system is used to identify songs, then music identification is improved, but identification of people and information about them deteriorates

Engineering Contradiction:
Improvesong identification accuracyVSAvoidpeople identification capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent extends the specialized song identification capability of systems like Song IDentity into a universal media identification system. By integrating voice recognition and facial recognition technologies, the system maintains accurate music identification while adding the capability to identify people, places, and objects in various media types, thus resolving the contradiction between specialized precision and general versatility

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS7787697B2Identification of an object in media and of related media objects
Publication Date: 2010.08.31 SONY GROUP CORP
  • US7787697B2 patent drawing
  • US7787697B2 patent drawing
  • US7787697B2 patent drawing

AI summary

A method obtains media on a device, provides identification of an object in the media via image/video recognition and audio recognition, and displays on the device identification information based on the identified media object.