Portable Terminal Image-to-Voice Conversion for Accessibility

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Visually impaired individuals cannot effectively utilize high-resolution cameras on portable terminals, as they lack the ability to interpret visual information, limiting their ability to identify signs and navigate environments independently.

Innovation Solution

A portable terminal apparatus and method that captures images, detects object areas, recognizes character information within those areas, and converts it into voice output, enabling users to select and hear information about the image through a touch screen and audio processing unit.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If a high-resolution camera is used in a portable terminal, then image quality is improved, but visually impaired users cannot utilize the camera function due to inability to interpret visual information

Engineering Contradiction:
Improveimage qualityVSAvoidusability for visually impaired users
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The patent replaces the visual interpretation mechanism with an auditory interpretation mechanism. The controller detects object areas in the captured image, recognizes character information within those areas, and converts it to voice output through an audio processing unit. This substitution enables visually impaired users to access camera functionality by replacing the visual sense with the auditory sense.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces an intermediary system consisting of the controller and audio processing unit that mediates between the camera's visual output and the user's auditory perception. The controller acts as the intermediary by processing the image data, detecting object areas, recognizing characters, and generating voice descriptions that convey the visual information to visually impaired users.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If object area detection and character recognition are added to enable voice output, then accessibility for visually impaired users is improved, but device complexity increases

Engineering Contradiction:
Improveaccessibility for visually impaired usersVSAvoidsystem complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent segments the image processing function into distinct modules: object area detection, character recognition, and voice conversion. The controller detects specific object areas within the image, recognizes character information only within those detected areas, and converts only the relevant recognized information to voice output. This segmentation allows the system to process only necessary portions of the image, reducing overall computational complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by not processing the entire image for character recognition. Instead, it first detects object areas and then performs character recognition only within those specific detected areas. This selective processing approach reduces the complexity of the recognition system compared to analyzing the entire image, while still providing sufficient information for visually impaired users.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS9971562B2Apparatus and method for representing an image in a portable terminal
Publication Date: 2018.05.15 SAMSUNG ELECTRONICS CO LTD
  • US9971562B2 patent drawing
  • US9971562B2 patent drawing
  • US9971562B2 patent drawing

AI summary

An apparatus for displaying an image in a portable terminal includes a camera to photograph the image, a touch screen to display the image and to allow selecting an object area of the displayed image, a memory to store the image, a controller to detect at least one object area within the image when displaying the image of the camera or the memory and to recognize object information of the detected object area to be converted into a voice, and an audio processing unit to output the voice.