Video Processing With Synchronized Transcripts for Focus Group Data Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for analyzing focus group videos require manual review to extract data, which is inefficient and time-consuming.

Innovation Solution

A system that uses artificial intelligence and machine learning to generate a synchronized transcript of a video, allowing users to select text for automatic playback to the next speaker, clip segments, and concatenate clips into a highlight video.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual review is used to extract data from focus group videos, then data extraction can be performed, but the process is inefficient and time-consuming

Engineering Contradiction:
Improvedata extraction efficiencyVSAvoidtime required for manual review
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent replaces the mechanical manual review process with an automated computer-based system that uses audio analysis and transcript generation to extract data from focus group videos, eliminating the need for human reviewers to manually watch and analyze entire video recordings

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service data extraction by automatically generating transcripts and identifying key moments in focus group videos without requiring manual intervention, allowing users to simply upload videos and retrieve processed results automatically

Inventive Principle:
Principle #25Self-service

2Loss of information

If the entire video is processed to generate a complete transcript, then all information is captured, but the processing time and computational resources increase

Engineering Contradiction:
Improveinformation completenessVSAvoidtranscript generation time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent extracts only the essential audio information from focus group videos to generate transcripts, selectively processing only the relevant spoken content rather than analyzing the entire video file including visual elements, thereby reducing processing time while maintaining information completeness

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system segments the video processing into separate audio and visual components, processing only the audio track for transcript generation while leaving the visual processing for later or separate handling, which significantly reduces the initial processing time and resource requirements

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If automatic speaker identification is implemented, then speaker differentiation is improved, but the system complexity increases

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary audio analysis layer that processes speaker patterns and characteristics to identify and differentiate speakers before generating the transcript, acting as a mediator between raw audio data and the final transcript output, which improves accuracy without significantly increasing overall system complexity

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250267026A1Systems and methods for processing and utilizing video data
Publication Date: 2025.08.21 MERCURY ANALYTICS LLC
  • US20250267026A1 patent drawing
  • US20250267026A1 patent drawing
  • US20250267026A1 patent drawing

AI summary

A method includes receiving, from an entity, a request to organize a survey on a topic, based on the request, organizing a survey of a plurality of people, recording a video of the survey, obtaining a transcription of the video and linking the transcription of the video in time to the video to yield a processed video. The method can further include presenting, on a user interface to the entity based on the processed video, the video and the transcription of the video, wherein each word in the transcription of the video is selectable by the entity, receiving a selection of text by the entity from the transcription of the video and, based on the selection of the text, presenting a portion of the video at a time that is associated with when a participant in the video spoke the text. The user can also select a “clip to next speaker” option to generate a clip.