Scanning System Image Recognition Vocal Output
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing scanning systems cannot effectively read and convey the content of original documents that do not contain characters, such as photographs or drawings, to users with visual impairments.
Innovation Solution
A scanning system that includes a generating section to scan documents, an image recognition section to perform image recognition on the scan data, and a speaking section to convert the recognition results into voice output for photographs, pictures, or drawings, allowing the system to speak words corresponding to these images.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If optical character recognition is used to read documents, then character content can be converted to speech, but documents without characters (photographs, drawings) cannot be recognized or conveyed
Solution Approach 1:
The system integrates multiple recognition functions: optical character recognition for text documents and image recognition for photographs and drawings. This multi-functional approach allows the scanning system to handle various document types uniformly, converting both text and images into speech output, thereby resolving the limitation of character-only recognition
Solution Approach 2:
The system introduces an image recognition section as an intermediary component between the scanning section and speaking section. This intermediary processes scan data to identify photographs, pictures, and drawings, enabling the system to recognize and convey content from non-text documents through appropriate speech output
2Loss of information
If the system speaks all recognition results including characters and codes, then complete information is provided, but users cannot distinguish between different content types
Solution Approach 1:
The system applies different speech output characteristics based on content type: it speaks words corresponding to photographs, pictures, and drawings differently from character text. This localized differentiation in speech output allows users to easily distinguish between image content and text content while receiving complete information
Solution Approach 2:
The speaking section is divided into specialized components: one for handling character/text recognition results and another for handling image recognition results (photographs, pictures, drawings). This segmentation enables distinct speech patterns for different content types, improving user comprehension while maintaining information completeness
Data Source
AI summary
A multi-function printer includes a generating section, an image recognition section, a speaking section, and a transmitting section. The generating section scans an original document to generate scan data. The image recognition section performs image recognition on the scan data. The speaking section causes a word corresponding to a recognition result of the image recognition section and corresponding to a drawing, a photograph, or the like included in the original document to be spoken from a speaker circuit. The transmitting section transmits the scan data generated by the generating section to a specific destination when a specific operation is performed during speaking performed by the speaking section or within a certain period of time after completion of the speaking performed by the speaking section.


