Video Recommendation System Using Voice-to-Text Sentiment Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video recommendation technologies fail to accurately provide product information to users by not analyzing the content of videos directly and extracting relevant product information from them, making it difficult for consumers to access product details effectively.
Innovation Solution
A video recommendation system that collects and analyzes videos related to products by converting voice data to text, extracting noun keywords, and performing sentiment analysis to identify relevant videos and provide partial video sections related to the product, allowing for accurate sentiment determination and easier product information access.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If video recommendation is based on user evaluations and viewing preferences, then user engagement is improved, but product information accuracy deteriorates
Solution Approach 1:
The patent segments video analysis into multiple dimensions: user engagement metrics (views, likes, comments) and product information extraction (voice content, keywords, sentiment). By separating these analysis streams, the system can recommend videos based on user preferences while simultaneously extracting accurate product information from video content through voice-to-text conversion and keyword analysis.
Solution Approach 2:
The patent introduces voice-to-text conversion technology as an intermediary to bridge the gap between video content and product information. This intermediary converts spoken content in videos into text, enabling accurate extraction of product names, features, and sentiments without relying solely on user evaluations, thus maintaining both user engagement and product information accuracy.
2Measurement precision
If video content analysis is performed to extract product information, then product information accuracy is improved, but system complexity increases
Solution Approach 1:
The patent applies preliminary action by pre-processing video content through voice-to-text conversion and storing transcribed text data before recommendation generation. This preliminary processing extracts product information in advance, creating a ready-to-use text database that simplifies subsequent recommendation operations and reduces real-time computational complexity.
Solution Approach 2:
The system performs self-service by automatically converting video voice content to text and extracting product keywords without manual intervention. This automated content analysis reduces the need for manual video tagging and product information entry, lowering operational complexity while maintaining high product information accuracy.
3Loss of information
If entire videos are provided to users, then complete product information is available, but user time consumption increases
Solution Approach 1:
The patent extracts and highlights specific product-related segments from entire videos by converting voice content to text and identifying product keywords. This extraction allows the system to present only the relevant portions of videos that contain product information, enabling users to access complete product details without watching entire videos, thus reducing time consumption while maintaining information completeness.
Solution Approach 2:
The patent adds a text dimension to video content by converting voice to text and creating searchable text transcripts alongside video files. This dimensional transformation allows users to search and navigate to specific product information segments within videos, bypassing the need to watch entire videos while ensuring complete product information accessibility.
Data Source
AI summary
Disclosed is a method for recommending a video by a video recommendation system, comprising: collecting and storing in a database of the video recommendation system videos related to products being sold and video information of the videos; converting voice included in each of the videos to text; obtaining words from the converted text and a time stamp for each of the words; extracting noun keywords in the text and identifying frequencies of the noun keywords, by analyzing morphemes of the text; performing a sentiment analysis on sentences composed of the words in the text; receiving a selection of one of the products; identifying videos associated with the selected product from among the videos stored in the database based on the noun keywords and the frequencies of the noun keywords; providing videos according to a predetermined criterion among the identified videos, based on a result of the sentiment analysis; and if one of the provided videos is selected, providing a partial video in a time section associated with the selected product.


