Lip-reading Algorithm for Video Transcription with Poor Audio Intelligibility
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In digital evidence management systems, accessing and deciphering audio content, particularly speech, from video recordings can be challenging due to poor intelligibility caused by factors like muting, microphone issues, or noise, making it difficult to transcribe crucial information.
Innovation Solution
A system comprising a computing device with a controller and memory that applies a lip-reading algorithm to video data portions with low intelligibility ratings and visible lips, generating text representative of detected lip movement to supplement or correct audio transcription.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If audio content is used for transcription, then speech information can be extracted, but intelligibility is poor due to muting, microphone issues, or noise
Solution Approach 1:
The patent uses lip movement as an intermediary to bridge the gap between unavailable audio content and speech information extraction. When audio is unintelligible, the system captures visual lip movements and uses them as a mediator to infer and reconstruct the spoken words, thereby recovering speech information that would otherwise be lost.
Solution Approach 2:
The patent transitions from a single dimension (audio-only transcription) to multiple dimensions by incorporating visual information from video footage. When audio fails, the system switches to analyzing lip movements in the video dimension, creating a multi-modal approach that compensates for audio deficiencies and improves overall speech recognition reliability.
2Loss of information
If lip-reading algorithm is applied to all video data, then text can be generated from lip movements, but processing time and computational resources increase
Solution Approach 1:
The patent applies lip-reading algorithms selectively rather than universally. The system first evaluates audio intelligibility and only applies the computationally intensive lip-reading process to video segments where audio fails. This partial application approach recovers speech information where needed while avoiding unnecessary processing of already-clear audio segments, thus reducing overall processing time.
Solution Approach 2:
The patent divides the video processing task into segments based on audio quality assessment. The system segments video data into portions with intelligible audio and portions with unintelligible audio, applying different processing strategies to each segment. This segmentation allows efficient resource allocation, applying complex lip-reading algorithms only to the necessary segments rather than the entire video.
3Measurement precision
If manual transcription of unclear audio is performed, then accurate text can be obtained, but time and labor resources are consumed
Solution Approach 1:
The patent implements an automated system that performs transcription without requiring manual human intervention. The lip-reading algorithm automatically analyzes video footage, detects lip movements, and generates text transcripts. This self-service capability eliminates the need for manual transcription of unclear audio segments, maintaining high accuracy while dramatically improving productivity by automating what was previously a labor-intensive process.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A device and method for generating text representative of lip movement is provided. One or more portions of video data are determined that include: audio with an intelligibility rating below a threshold intelligibility rating; and lips of a human face. A lip-reading algorithm is applied to the one or more portions of the video data to determine text representative of detected lip movement in the one or more portions of the video data. The text representative of the detected lip movement is stored in a memory. A transcript that includes the text representative of the detected lip movement may be generated. Captioned video data may be generated from the video data and the text representative of detected lip movement.