E-book Read-Along Synchronization via OCR and Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems fail to synchronize e-book digital content playback and display effectively across different devices, especially in suboptimal audio environments, and lack user-specific speech emulation for accurate animation and interactivity.

Innovation Solution

A system that uses optical character recognition (OCR) to calculate on-screen display coordinates, performs speech recognition to generate animation playback timing metadata, and employs a machine learning module for user-specific speech emulation, enabling synchronized playback and interactivity across various devices and environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional speech recognition is used to generate animation playout timing metadata, then the system can process audio data, but accuracy deteriorates in suboptimal audio environments

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidaudio environment quality
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The system uses feedback by comparing recognized speech against the known sequence of words in the e-book content. The speech recognition results are validated and corrected by referencing the expected word sequence, thereby maintaining high accuracy even in suboptimal audio environments where conventional speech recognition would fail.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The known word sequence of the e-book acts as an intermediary that mediates between the noisy audio input and the final speech recognition output. By using the expected word sequence as a reference framework, the system can accurately generate animation playout timing metadata despite poor audio conditions.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If OCR is performed on e-book content to determine XY coordinates for each word, then display synchronization is achieved, but processing time increases

Engineering Contradiction:
Improvedisplay coordinate accuracyVSAvoidcontent processing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs preliminary action by pre-calculating and storing the XY coordinates of each word in the e-book content before the actual playout. This preprocessing step allows the coordinates to be readily available during playback, eliminating the need for real-time coordinate calculation and thus reducing processing time during actual use.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If speech recognition correlates spoken words with XY coordinates to generate animation metadata, then animation synchronization is achieved, but system complexity increases

Engineering Contradiction:
Improveanimation synchronization reliabilityVSAvoidsystem architecture complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system merges multiple functions into a unified process: speech recognition, coordinate correlation, and animation metadata generation are combined into a single integrated workflow. This consolidation improves reliability by ensuring all steps work together seamlessly while managing complexity through unified architecture rather than separate independent systems.

Inventive Principle:
Principle #5Merging (Combining)

4Adaptability or versatility

If the system supports multiple display devices with different aspect ratios and resolutions, then device compatibility is improved, but coordinate calculation complexity increases

Engineering Contradiction:
Improvedisplay device compatibilityVSAvoidcoordinate emulation complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system implements universality by creating a renderer controller that can emulate multiple display characteristics (different aspect ratios, resolutions, and orientations) using a single unified coordinate calculation framework. This allows the same e-book content to be accurately displayed across diverse devices without requiring separate coordinate systems for each device type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11620252B2System for recorded e-book digital content playout
Publication Date: 2023.04.04 THE UTREE GRP PTY LTD
  • US11620252B2 patent drawing
  • US11620252B2 patent drawing
  • US11620252B2 patent drawing

AI summary

A system allows for audio playout of e-book content data using a playout electronic device recorded by a remote recording electronic device. The system may analyse the e-book content to infer XY on-screen display coordinates for each word of the e-book and speech recognition may correlate the timing of the spoken words to the XY coordinates. As such, a read along display animation may be generated by the digital display of the playout electronic device in time with the audio and the respective on-screen position of each word. Eye tracking may be employed by the playout electronic device for the display of a gaze position indicator to the recording electronic device in substantial real-time. The system may further employ machine learning to optimise a trained machine to output at least one prosodic features for user profile specific speech emulation using a speech emulator.