ATM Component Mapping With AI-Generated Multilingual Audio Guidance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current ATM systems require manual and inaccurate processes for measuring dimensions and generating audio scripts, necessitating separate coding for sighted and visually impaired users, and lack support for multiple languages and dialects.
Innovation Solution
A generative artificial intelligence-based system that analyzes ATM dimensions, generates accurate audio scripts in multiple languages, and adapts user interfaces to user preferences, using nested AI models for efficient and dynamic output generation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual processes are used to measure ATM dimensions and generate audio scripts, then the process can be performed with simple equipment, but the measurement accuracy and script generation accuracy are insufficient
Solution Approach 1:
The patent replaces manual measurement processes with an image capture device and AI-based processing system. The image capture device captures images of the ATM front face, and a generative AI model processes these images to automatically determine component positions and generate audio scripts, eliminating the need for manual measuring tools and processes while significantly improving measurement accuracy.
Solution Approach 2:
The system creates a digital copy (image) of the ATM front face and processes this copy to extract dimensional information and component positions. By working with the image representation rather than physical measurements, the system achieves high precision without requiring complex physical measurement equipment.
2Ease of operation
If separate coding is performed for sighted and visually impaired users, then user-specific functionality can be optimized, but development time and complexity increase
Solution Approach 1:
The system dynamically generates user interfaces and audio scripts based on the detected user type. The generative AI model adapts the ATM interface in real-time, creating sighted-user interfaces with visual displays for sighted users and audio-only interfaces for visually impaired users. This dynamic adaptation eliminates the need for separate static codebases while maintaining optimized functionality for each user type.
Solution Approach 2:
The patent implements a universal ATM system that can serve both sighted and visually impaired users through a single codebase. The generative AI model processes user inputs and automatically generates appropriate interface modes, allowing one system to perform multiple functions that previously required separate systems.
3Adaptability or versatility
If limited language options are provided, then system complexity is reduced, but adaptability to different user language preferences is insufficient
Solution Approach 1:
The system changes the language parameter dynamically based on user detection. The generative AI model identifies the user's preferred language and automatically generates interface content in that language. This parameter-based adaptation allows the system to support multiple languages without requiring complex multi-language configuration systems, as the language selection is automated based on user characteristics.
Data Source
AI summary
Arrangements for using generative artificial intelligence models for ATM process generation are provided. In some examples, a computing platform may receive, from at least one image or measurement capture device, dimension data associated with an ATM. The ATM may have a plurality of components arranged on a face of the ATM. The dimension data may be input to a generative artificial intelligence model and the model may be executed to output, based on the dimension data, a position or location of each ATM component relative to a reference point on the ATM. In some examples, the model may further output at least one audio script describing the location of each component. The model may output one or more translations of the at least one audio script. The computing platform may transmit or send the at least one audio script to the ATM for presentation during user interaction with the ATM.


