Braille Output System Using Multimodal Input Comparison
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current systems that rely on sign language interpreters to convert spoken language into sign language do not effectively provide braille output for individuals who are both hearing-impaired and vision-impaired, leading to incomplete or inaccurate information due to human errors in transcription.
Innovation Solution
A system and method that utilize a camera to detect image data from sign language and a microphone to detect audio data from spoken language, with a processor converting both into text and determining an optimal word by comparing confidence values of the image-based and audio-based text, enabling near-real-time conversion to braille for improved accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a sign language interpreter converts spoken language into sign language, then hearing-impaired individuals can receive speech content, but vision-impaired and hearing-impaired individuals cannot access the information accurately
Solution Approach 1:
The patent replaces the human sign language interpreter system with an automated computer vision and speech recognition system. The system uses cameras to detect sign language gestures, microphones to capture spoken language, and processors to convert both into text, eliminating human transcription errors and providing reliable real-time access for hearing-impaired individuals.
Solution Approach 2:
The patent introduces an intermediary automated system between the speaker and the hearing-impaired audience. This intermediary system captures both spoken and signed language, processes them through multiple recognition algorithms, and generates accurate text output that serves as a reliable bridge for information transmission.
2Ease of operation
If a human transcribes speech into braille, then vision-impaired individuals can access information, but the transcription may be incomplete or inaccurate due to human limitations
Solution Approach 1:
The patent replaces human transcription with an automated computer-based system that uses optical character recognition for sign language gestures and speech recognition for spoken words. This automated system processes both input streams simultaneously, generating accurate and complete braille text without the limitations of human memory and attention.
Solution Approach 2:
The system incorporates feedback mechanisms where the processed text is displayed on a braille display device, allowing real-time verification and correction. The system can detect and correct errors by comparing the transcribed text against the original audio and visual inputs, ensuring high accuracy in the braille output.
3Measurement precision
If the system processes both audio and image data, then accuracy is improved through comparison, but device complexity increases
Solution Approach 1:
The patent divides the processing system into separate functional modules: a camera system for capturing sign language, a microphone system for capturing spoken language, independent processing units for each input type, and a comparison module that integrates both streams. This segmentation allows each component to be optimized independently while working together to achieve high accuracy.
Solution Approach 2:
The patent creates a universal processing system that can handle multiple input modalities (audio and visual) through a common architecture. The same processor handles both audio data and image data, converting them to text through different recognition algorithms and then comparing the results. This multi-functional approach improves accuracy without proportionally increasing complexity.
Data Source
AI summary
A system for determining output text based on spoken language and sign language includes a camera configured to detect image data corresponding to a word in sign language. The system also includes a microphone configured to detect audio data corresponding to the word in spoken language. The system also includes a processor configured to receive the image data from the camera and convert the image data into an image based text word. The processor is also configured to receive the audio data from the microphone and convert the audio data into an audio based text word. The processor is also configured to determine an optimal word by selecting one of the image based text word or the audio based text word based on a comparison of the image based text word and the audio based text word.


