Automated Media Accompaniment Suggestion via Visual Object Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users face difficulties in manually selecting appropriate audio tracks to enhance their media content, such as slideshows or videos, due to the vast number of options and the limitations of mobile devices' compact form factor and user interface.
Innovation Solution
The technology automatically suggests audio, video, or other media accompaniments based on identified objects in the media content, using visual or aural features to determine relevant keywords and arrange audio tracks accordingly, thereby streamlining the process and enhancing user experience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If users manually select audio tracks from a large number of options, then they can find appropriate accompaniments for their media content, but the process becomes time-consuming and complex due to the vast number of available sound effects and music tracks
Solution Approach 1:
The system performs automatic audio track selection based on visual content analysis, allowing the media content itself to 'select' appropriate accompaniments through automated object recognition and keyword generation, eliminating the need for manual user selection while maintaining adaptability to different content types
Solution Approach 2:
The patent introduces an intermediary system that bridges media content and audio tracks through automated analysis. This intermediary generates keywords from visual content and uses them to search and select appropriate audio tracks, resolving the contradiction by providing adaptability without requiring direct manual user interaction
2Adaptability or versatility
If users manually browse and select audio tracks on mobile devices, then they can choose suitable accompaniments, but the compact form factor and user interface make the process unwieldy and time-consuming
Solution Approach 1:
The mobile device automatically performs audio track selection based on analyzing visual content, eliminating the need for users to manually browse and select tracks through the device's limited interface. The system serves itself by generating keywords from media content and automatically matching them with appropriate audio tracks
Solution Approach 2:
The patent extracts the complex audio selection task from the user's manual operation and transfers it to automated system processing. By extracting keywords from visual content and using them to search audio databases, the system removes the burden of manual browsing while maintaining the ability to find suitable audio tracks
3Productivity
If the system automatically suggests audio tracks based on identified objects, then the selection process is streamlined and time is reduced, but the system complexity increases due to object recognition and keyword generation requirements
Solution Approach 1:
The patent segments the audio track selection process into distinct automated stages: visual content analysis, object identification, keyword generation, and audio track matching. This segmentation allows each component to be handled by specialized algorithms, improving overall productivity while managing system complexity through modular processing
Solution Approach 2:
The system performs preliminary actions by pre-generating keywords from visual content and pre-searching for matching audio tracks before user interaction is needed. This preliminary automated processing significantly improves selection efficiency by having recommendations ready in advance, while the complexity is managed through automated scripting and algorithmic processing
Data Source
AI summary
The disclosed technology includes automatically suggesting audio, video, or other media accompaniments to media content based on identified objects in the media content. Media content may include images, audio, video, or a combination. In one implementation, one or more images representative of the media content may be extracted. A visual search may be run across the images to identify objects or characteristics present in or associated with the media content. Keywords may be generated based on the identified objects and characteristics. The keywords may be used to determine suitable audio tracks to accompany the media content, for example by performing a search based on the keywords. The determined tracks may be presented to a user, or automatically arranged to match the media content. In another implementation, an aural search may be run across samples of the audio data to similarly identify objects and characteristics of the media content.


