Lip-reading Algorithm for Video Transcription with Poor Audio Intelligibility

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In digital evidence management systems, accessing and deciphering audio content, particularly speech, from video recordings can be challenging due to poor intelligibility caused by factors like muting, microphone issues, or noise, making it difficult to transcribe crucial information.

Innovation Solution

A system comprising a computing device with a controller and memory that applies a lip-reading algorithm to video data portions with low intelligibility ratings and visible lips, generating text representative of detected lip movement to supplement or correct audio transcription.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If audio content is used for transcription, then speech information can be extracted, but intelligibility is poor due to muting, microphone issues, or noise

Engineering Contradiction:
Improvespeech informationVSAvoidaudio intelligibility
Core Design Contradiction:
Loss of informationVSReliability

Solution Approach 1:

The patent uses lip movement as an intermediary to bridge the gap between unavailable audio content and speech information extraction. When audio is unintelligible, the system captures visual lip movements and uses them as a mediator to infer and reconstruct the spoken words, thereby recovering speech information that would otherwise be lost.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transitions from a single dimension (audio-only transcription) to multiple dimensions by incorporating visual information from video footage. When audio fails, the system switches to analyzing lip movements in the video dimension, creating a multi-modal approach that compensates for audio deficiencies and improves overall speech recognition reliability.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Loss of information

If lip-reading algorithm is applied to all video data, then text can be generated from lip movements, but processing time and computational resources increase

Engineering Contradiction:
Improvespeech information recoveryVSAvoidvideo processing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent applies lip-reading algorithms selectively rather than universally. The system first evaluates audio intelligibility and only applies the computationally intensive lip-reading process to video segments where audio fails. This partial application approach recovers speech information where needed while avoiding unnecessary processing of already-clear audio segments, thus reducing overall processing time.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent divides the video processing task into segments based on audio quality assessment. The system segments video data into portions with intelligible audio and portions with unintelligible audio, applying different processing strategies to each segment. This segmentation allows efficient resource allocation, applying complex lip-reading algorithms only to the necessary segments rather than the entire video.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If manual transcription of unclear audio is performed, then accurate text can be obtained, but time and labor resources are consumed

Engineering Contradiction:
Improvetranscription accuracyVSAvoidtranscription efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent implements an automated system that performs transcription without requiring manual human intervention. The lip-reading algorithm automatically analyzes video footage, detects lip movements, and generates text transcripts. This self-service capability eliminates the need for manual transcription of unclear audio segments, maintaining high accuracy while dramatically improving productivity by automating what was previously a labor-intensive process.

Inventive Principle:
Principle #25Self-service

Data Source

PatentEP3711049B1Device and method for generating text representative of lip movement
Publication Date: 2021.10.27 MOTOROLA SOLUTIONS INC
  • EP3711049B1 patent drawingFigure 1
  • EP3711049B1 patent drawingFigure 2
  • EP3711049B1 patent drawingFigure 3

AI summary

A device and method for generating text representative of lip movement is provided. One or more portions of video data are determined that include: audio with an intelligibility rating below a threshold intelligibility rating; and lips of a human face. A lip-reading algorithm is applied to the one or more portions of the video data to determine text representative of detected lip movement in the one or more portions of the video data. The text representative of the detected lip movement is stored in a memory. A transcript that includes the text representative of the detected lip movement may be generated. Captioned video data may be generated from the video data and the text representative of detected lip movement.