Voiceprint Recognition for Personalized Multimedia Preview

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Smart televisions lack personalized service delivery to multiple family members, as they provide the same content to all users without distinguishing individual preferences or identities.

Innovation Solution

A method and apparatus that utilize voiceprint recognition to generate a voiceprint characteristic vector from user input, matching it with pre-trained models to identify user identity and select relevant multimedia files for personalized content recommendations, including preview information based on usage history and timbre preferences.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If voiceprint recognition is implemented to identify user identity, then personalized service capability is improved, but device complexity increases

Engineering Contradiction:
Improvepersonalized service capabilityVSAvoiddevice complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The smart television system integrates multiple functions including voiceprint recognition, user identity identification, and personalized content recommendation into a single unified platform. The voiceprint recognition model serves as a universal component that can identify different family members and trigger personalized services automatically, making the television adaptable to multiple users without requiring separate identification devices for each user.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The voiceprint characteristic vector acts as an intermediary between the user's voice input and the user identity. Instead of directly comparing raw voice signals, the system extracts and processes voiceprint features through a recognition model, creating an intermediate representation that enables accurate user identification while maintaining system efficiency and manageability.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If personalized content recommendation is provided based on user identity, then user experience is improved, but information processing complexity increases

Engineering Contradiction:
Improveuser experience qualityVSAvoidinformation processing complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by pre-processing and storing voiceprint characteristic vectors and user profile information during the user registration phase. When a user speaks, the system compares their voiceprint against previously stored templates and retrieves pre-configured personalized content recommendations, avoiding the need for complex real-time analysis and reducing processing complexity during actual use.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements feedback mechanisms by accumulating user interaction data such as playback counts and retrieval frequencies for different multimedia files. This feedback information is used to refine and update user profiles and content recommendations over time, improving personalization accuracy while using simple counting and comparison operations rather than complex algorithms.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11006179B2Method and apparatus for outputting information
Publication Date: 2021.05.11 BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
  • US11006179B2 patent drawing
  • US11006179B2 patent drawing
  • US11006179B2 patent drawing

AI summary

A method and an apparatus for outputting information are provided. A specific embodiment of the method comprises: in response to receiving voice inputted by a user, generating a voiceprint characteristic vector based on the voice; inputting the voiceprint characteristic vector into a voiceprint recognition model to obtain identity information of the user; selecting, from a preset multimedia file set, a predetermined number of multimedia files matching the obtained identity information of the user as target multimedia files; and generating, according to the target multimedia files, preview information, and outputting the preview information. This embodiment realizes the multimedia preview information recommendation with pertinence.