Scanning System Image Recognition Vocal Output

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing scanning systems cannot effectively read and convey the content of original documents that do not contain characters, such as photographs or drawings, to users with visual impairments.

Innovation Solution

A scanning system that includes a generating section to scan documents, an image recognition section to perform image recognition on the scan data, and a speaking section to convert the recognition results into voice output for photographs, pictures, or drawings, allowing the system to speak words corresponding to these images.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If optical character recognition is used to read documents, then character content can be converted to speech, but documents without characters (photographs, drawings) cannot be recognized or conveyed

Engineering Contradiction:
Improvecontent conveyance capabilityVSAvoiddocument type compatibility
Core Design Contradiction:
Loss of informationVSAdaptability or versatility

Solution Approach 1:

The system integrates multiple recognition functions: optical character recognition for text documents and image recognition for photographs and drawings. This multi-functional approach allows the scanning system to handle various document types uniformly, converting both text and images into speech output, thereby resolving the limitation of character-only recognition

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system introduces an image recognition section as an intermediary component between the scanning section and speaking section. This intermediary processes scan data to identify photographs, pictures, and drawings, enabling the system to recognize and convey content from non-text documents through appropriate speech output

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If the system speaks all recognition results including characters and codes, then complete information is provided, but users cannot distinguish between different content types

Engineering Contradiction:
Improveinformation completenessVSAvoiduser comprehension
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The system applies different speech output characteristics based on content type: it speaks words corresponding to photographs, pictures, and drawings differently from character text. This localized differentiation in speech output allows users to easily distinguish between image content and text content while receiving complete information

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The speaking section is divided into specialized components: one for handling character/text recognition results and another for handling image recognition results (photographs, pictures, drawings). This segmentation enables distinct speech patterns for different content types, improving user comprehension while maintaining information completeness

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11336793B2Scanning system for generating scan data for vocal output, non-transitory computer-readable storage medium storing program for generating scan data for vocal output, and method for generating scan data for vocal output in scanning system
Publication Date: 2022.05.17 SEIKO EPSON CORP
  • US11336793B2 patent drawing
  • US11336793B2 patent drawing
  • US11336793B2 patent drawing

AI summary

A multi-function printer includes a generating section, an image recognition section, a speaking section, and a transmitting section. The generating section scans an original document to generate scan data. The image recognition section performs image recognition on the scan data. The speaking section causes a word corresponding to a recognition result of the image recognition section and corresponding to a drawing, a photograph, or the like included in the original document to be spoken from a speaker circuit. The transmitting section transmits the scan data generated by the generating section to a specific destination when a specific operation is performed during speaking performed by the speaking section or within a certain period of time after completion of the speaking performed by the speaking section.