Digital Note Compilation from Video Using Edge Detection and Audio Transcription

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional systems for generating digital summaries from presentations face challenges in accuracy, efficiency, and flexibility, particularly in merging handwritten content and digital audio from digital videos, often resulting in inaccurate and rigid summaries that fail to capture the dynamic flow of presentations.

Innovation Solution

The system employs an edge detection algorithm to track handwritten content, combines it with digital audio using optical character recognition and similarity index search, and auto-corrects handwritten text, generating intuitive and organized summaries that accurately reflect both written and spoken content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional systems generate digital summaries from digital video, then digital summaries can be produced, but accuracy and flexibility are poor due to inability to properly merge handwritten content and digital audio

Engineering Contradiction:
Improveaccuracy of digital summaryVSAvoidcomplexity of merging handwritten content and digital audio
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments the digital video into separate components: handwritten content on writing surfaces and digital audio tracks. By processing these segments independently and then merging them based on temporal alignment, the system achieves accurate summaries without overwhelming complexity. The transcription generator extracts handwritten text separately from audio transcription, then combines them with temporal positioning information.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces temporal positioning information and writing surface identification as intermediary elements that facilitate the merging of handwritten content and digital audio. These intermediaries act as bridges that align the two different content types in time and space, enabling accurate summary generation without direct complex integration of the source materials.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If conventional systems process digital video frames to extract content, then visual components can be summarized, but efficiency is reduced due to rigid processing approaches

Engineering Contradiction:
Improveefficiency of digital summary generationVSAvoidflexibility to capture dynamic presentation flow
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The system dynamically adapts to the presentation flow by continuously monitoring writing surface changes and audio content in real-time. Instead of rigid frame-by-frame processing, the system adjusts its processing based on detected changes in handwritten content and speech, capturing only relevant segments and aligning them temporally to improve efficiency while maintaining flexibility.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes processing parameters based on the detected state of the presentation. By monitoring temporal alignment between handwritten content and audio, the system adjusts its extraction and merging parameters dynamically, processing content more efficiently during stable states and adapting when changes are detected in the presentation flow.

Inventive Principle:
Principle #35Parameter changes

3Loss of information

If conventional systems generate transcriptions from digital audio, then spoken content can be captured, but integration with handwritten content is inaccurate and rigid

Engineering Contradiction:
Improvecompleteness of content captureVSAvoidcomplexity of merging transcription with handwritten content
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system merges digital audio transcription and handwritten content transcription by aligning them temporally and associating them with specific writing surfaces. This combination approach ensures both spoken and written content are captured without loss, while the temporal alignment mechanism simplifies the integration process by providing a natural ordering framework that reduces merging complexity.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS10929684B2Intelligently generating digital note compilations from digital video
Publication Date: 2021.02.23 ADOBE INC
  • US10929684B2 patent drawing
  • US10929684B2 patent drawing
  • US10929684B2 patent drawing

AI summary

The present disclosure relates to systems, non-transitory computer-readable media, and methods for intelligently merging handwritten content and digital audio from a digital video based on monitored presentation flow. In particular, the disclosed systems can apply an edge detection algorithm to intelligently detect distinct sections of the digital video and locations of handwritten content entered onto a writing surface over time. Moreover, the disclosed systems can generate a transcription of handwritten content utilizing digital audio. For instance, the disclosed systems can utilize an audio text transcript as input to an optical character recognition algorithm and auto-correct text utilizing the audio text transcript. Further, the disclosed systems can analyze short form text from handwritten script and generate long form text from audio text transcripts. The disclosed systems can accurately, efficiently, and flexibly generate digital summaries that reflect diagrams, handwritten text transcriptions, and audio text transcripts over different presentation time periods.