IMU-Supported Super-Resolution for Wearable OCR

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Performing optical character recognition (OCR) on images captured by head-mounted or wearable devices is challenging due to focusing issues, perspective corrections, small camera form factors, high f-numbers, low power consumption, and motion blurring, especially in low-light environments.

Innovation Solution

The use of positional sensors, such as Inertial Measurement Units (IMUs), and super-resolution image post-processing techniques to enhance image quality and compensate for camera movements, allowing for high-acuity OCR capabilities despite camera constraints.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Use of energy by moving object

If a small form factor camera is used in wearable devices, then device portability and power consumption are improved, but image resolution and focus quality deteriorate

Engineering Contradiction:
Improvepower consumptionVSAvoidimage resolution
Core Design Contradiction:
Use of energy by moving objectVSMeasurement precision

Solution Approach 1:

The patent transitions from spatial resolution improvement to temporal dimension by capturing multiple images over time. The system captures a sequence of images at different time points and combines them to achieve high resolution, effectively using the time dimension to compensate for the limited spatial capabilities of small cameras.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent implements continuous image capture during a time interval to accumulate useful image data. By continuously capturing images and combining them through alignment and fusion processes, the system maintains continuous useful action to overcome the limitations of individual low-resolution frames.

Inventive Principle:
Principle #20Continuity of useful action

2Ease of operation

If the camera is affixed to the user's head or body, then hands-free operation is improved, but motion blurring increases due to uncompensated camera movements

Engineering Contradiction:
Improvehands-free operationVSAvoidimage clarity
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent employs feedback mechanisms by using IMU sensors to detect camera motion and then applying corrective transformations to the captured images. The system continuously monitors motion through sensors and adjusts image alignment accordingly, creating a feedback loop that compensates for motion-induced degradation.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent introduces an intermediary processing stage between image capture and final output. The system uses intermediate image alignment and fusion processes, guided by IMU data, to correct motion artifacts before producing the final high-resolution image, effectively using intermediaries to bridge the gap between motion and image quality.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If images are captured in low-light scenes, then operational versatility is improved, but image noise and reduced spatial resolution increase

Engineering Contradiction:
Improveoperational versatilityVSAvoidspatial resolution
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent performs preliminary actions by capturing multiple images before final processing. The system collects a sequence of low-light images and performs preliminary alignment and fusion operations to accumulate signal information, preparing the data in advance for high-resolution reconstruction before OCR is applied.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent merges multiple low-quality images into a single high-quality image through alignment and fusion processes. By combining information from multiple captured frames, the system accumulates sufficient signal to overcome the noise and low resolution inherent in individual low-light captures.

Inventive Principle:
Principle #5Merging (Combining)

4Stability of the object's composition

If a large f-number camera is used, then depth of field is improved, but light gathering capability and image brightness deteriorate

Engineering Contradiction:
Improvedepth of fieldVSAvoidimage brightness
Core Design Contradiction:
Stability of the object's compositionVSIllumination intensity

Solution Approach 1:

The patent maintains continuous image capture over a time interval to accumulate sufficient light information. By continuously capturing images and combining them through fusion processes, the system accumulates enough photon information to compensate for the limited light gathering capability of large f-number cameras.

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The patent compensates for reduced light gathering in the spatial domain by transitioning to the temporal domain. The system captures multiple images over time and combines them through fusion, using the time dimension to accumulate sufficient light information that would be unavailable in a single exposure.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS20250104456A1Optical Character Recognition (OCR) Enhancement via Inertial Measurement Unit (IMU)-Supported Super-Resolution Imaging
Publication Date: 2025.03.27 APPLE INC
  • US20250104456A1 patent drawing
  • US20250104456A1 patent drawing
  • US20250104456A1 patent drawing

AI summary

Electronic devices, methods, and program storage devices for achieving improved optical character recognition (OCR) operations are disclosed. Performing OCR operations on captured images, e.g., images captured by cameras that are affixed to a user's body (e.g., from mixed reality devices, such as smart HMDs) requires a low-power, robust camera design. Obtaining high spatial resolution in such captured images faces many challenges. However, images with higher spatial resolution can be created by combining information extracted from multiple images captured by such devices, leveraging information obtained from positional sensors of such devices, and performing SR post-processing operations. Such higher spatial resolution images may then be used to enable high-acuity OCR capabilities. The solutions disclosed herein also compensate for the missing ability of such devices due to the lack of a vestibulo-ocular reflex (i.e., the human visual system's ability to use compensating eye movement to fixate and read text clearly, despite head movement).