Video Recommendation Voiceprint Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The accuracy of video recommendations in smart voice televisions is low because they rely on user profiles generated from all users, rather than individual user preferences, leading to non-personalized content suggestions.

Innovation Solution

A method that involves receiving a video recommendation request, determining a voice or face image with the greatest similarity to the user's input, and sending targeted video information based on a user profile associated with the identified voice or face, with confidence thresholds ensuring accurate recognition and personalized recommendations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If user profiles are generated from data of all users, then the system can provide video recommendations, but the recommendation accuracy is low because it does not reflect individual user preferences

Engineering Contradiction:
Improverecommendation accuracyVSAvoiduser identification system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the aggregated user data into individual user profiles by identifying and separating data based on unique user identifiers (voiceprints, face images, or account information). This allows the system to generate personalized recommendations for each user rather than providing generic recommendations based on all users' combined data, thereby improving recommendation accuracy while maintaining manageable system complexity through structured data organization.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If voiceprint recognition is used to identify users, then personalized recommendations can be provided, but the system complexity increases due to voice storage and comparison requirements

Engineering Contradiction:
Improveuser identification accuracyVSAvoidvoice data storage requirements
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent creates compact voiceprint templates that serve as simplified copies of actual user voices. Instead of storing and processing complete voice recordings, the system extracts essential acoustic features and stores these condensed representations. This significantly reduces the quantity of data that needs to be stored and processed while maintaining high user identification accuracy, as the voiceprint templates capture the unique characteristics needed for recognition without requiring full voice samples.

Inventive Principle:
Principle #26Copying

3Reliability

If multiple recognition methods (voice and face) are implemented, then user identification reliability improves, but the system complexity and processing time increase

Engineering Contradiction:
Improveuser identification reliabilityVSAvoidrecognition processing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements a dynamic recognition system that adaptively selects between voiceprint and face image recognition methods based on the current situation. The system can switch between recognition modes depending on factors such as user availability, environmental conditions, and confidence levels. This dynamic approach maintains high identification reliability by using the most appropriate method for each scenario while minimizing processing time by avoiding unnecessary recognition steps.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent introduces an intermediary layer that manages multiple recognition methods through a unified interface. This intermediary component coordinates between different recognition algorithms, handles confidence threshold evaluations, and determines when to use voice versus face recognition. By mediating between multiple recognition systems, the patent maintains reliability through comprehensive verification while reducing overall processing time through intelligent method selection and parallel processing capabilities.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10694247B2Method and apparatus for recommending video
Publication Date: 2020.06.23 BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
  • US10694247B2 patent drawing
  • US10694247B2 patent drawing
  • US10694247B2 patent drawing

AI summary

Embodiments of the present disclosure disclose a method and apparatus for recommending a video. The method includes: receiving a video recommendation request sent by a terminal device, the video recommendation request including a first voice, the first voice being a voice inputted by a user requesting a video recommendation; determining, from user voices stored in a server, a second voice having a greatest similarity with the first voice; and sending information of a target video to the terminal device according to a user profile corresponding to the second voice, if a first confidence recognizing the user as a user corresponding to the second voice being greater than or equal to a first threshold. The method and apparatus for recommending a video according to the embodiments of the present disclosure have a high accuracy of video recommendation.