Automated Video Performance Segment Labeling via Voiceprint Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The existing method for labeling performance segments in videos is inefficient and imprecise due to manual operation by editors, leading to delays in providing a 'watch-him-only' function based on acting roles, especially after live broadcasts.

Innovation Solution

A method that automatically labels performance segments by determining role features from multimedia files, decoding videos to identify matching data frames, and combining timestamps to generate segment information, enabling efficient and precise labeling of acting roles' appearances.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual labeling by operation editors is used, then the watch-him-only function can be provided, but the labeling precision and efficiency are low

Engineering Contradiction:
Improvelabeling precisionVSAvoidlabeling efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent replaces the manual mechanical labeling process with an automated system using voiceprint recognition technology. The server automatically identifies acting roles by extracting and matching voiceprint features from video audio data, eliminating the need for manual operation editor intervention and achieving both high precision and high efficiency in performance segment labeling

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service automated labeling where the server independently performs voiceprint extraction, feature matching, and segment identification without human intervention. The automatic labeling system serves itself by using pre-stored voiceprint templates to identify acting roles and generate performance segment labels autonomously

Inventive Principle:
Principle #25Self-service

2Productivity

If manual labeling by operation editors is used, then the watch-him-only function can be provided, but it takes several days to provide the function after live broadcast

Engineering Contradiction:
Improvelabeling efficiencyVSAvoidtime delay
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent replaces the slow manual labeling process with automated voiceprint recognition technology that can rapidly process video files. The server automatically extracts voiceprint features, matches them against stored templates, and generates performance segment labels in minutes rather than days, dramatically reducing the time delay between live broadcast and availability of the watch-him-only function

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs preliminary action by pre-extracting and storing voiceprint features from video files automatically. The voiceprint templates are prepared in advance and can be quickly matched when labeling is needed, enabling rapid deployment of the watch-him-only function shortly after live broadcast without requiring time-consuming manual review

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11625920B2Method for labeling performance segment, video playing method, apparatus and system
Publication Date: 2023.04.11 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US11625920B2 patent drawing
  • US11625920B2 patent drawing
  • US11625920B2 patent drawing

AI summary

Provided is a method for labeling a segment of a video, in a server. In the method, a multimedia file corresponding to an acting role is obtained. A role feature of the acting role is determined based on the multimedia file. A target video is decoded to obtain a data frame and a playing timestamp corresponding to the data frame, the data frame including at least one of a video frame and an audio frame. In the data frame of the target video, a target data frame that matches the role feature is identified. A segment related to performance of the acting role in the target video is automatically labeled based on a playing timestamp of the target data frame.