Visual Language Processing via Attention-on-Attention Mechanism
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current AI models face challenges in manufacturing due to limited sample sizes for product design and personalization, leading to slow adoption and reduced accuracy in quality inspection tasks, especially in additive manufacturing, where human expertise is still essential for accurate and flexible decision-making.
Innovation Solution
A visual language processing (VLP) modeling framework using an attention-on-attention (AonA) mechanism that captures human visual searching patterns by analyzing eye movements to generate computational attention features, enhancing AI models with interpretable and dynamic feature extraction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If traditional AI models are used for manufacturing inspection, then automation is improved, but accuracy and reliability deteriorate due to limited sample sizes
Solution Approach 1:
The patent introduces human visual attention mechanisms as an intermediary between raw image data and AI classification. Eye tracking data captures human experts' visual searching patterns, which serve as a mediator to guide the AI model's attention to relevant features. This intermediary layer allows the system to leverage human expertise without requiring large annotated datasets, thereby maintaining high accuracy while achieving automation.
Solution Approach 2:
The system performs multiple functions: it captures human visual attention patterns through eye tracking, processes manufacturing images, generates attention maps, and performs defect classification. By integrating these functions into a unified framework, the system achieves both automation and high accuracy simultaneously, overcoming the traditional trade-off between automation and reliability.
2Reliability
If human expertise is incorporated into AI models, then accuracy is improved, but device complexity increases
Solution Approach 1:
The patent replaces complex manual feature engineering and expert system rules with an attention mechanism driven by eye tracking data. Instead of building complex knowledge bases or hand-crafted feature extractors, the system uses neural network attention layers that automatically learn from human visual behavior. This substitution reduces structural complexity while maintaining or improving accuracy.
Solution Approach 2:
The system changes the parameters fed into the AI model from raw pixel values to attention-weighted features derived from eye tracking data. By transforming the input parameters to reflect human visual priorities, the model achieves higher accuracy with a simpler architecture, as the attention mechanism pre-processes and highlights relevant information before classification.
3Measurement precision
If more features are extracted from raw data, then measurement precision is improved, but loss of information increases due to feature selection challenges
Solution Approach 1:
The patent implements feedback loops where eye tracking data continuously informs the attention mechanism. Human experts' visual patterns provide feedback on which features are most relevant, and this feedback is used to dynamically adjust attention weights. This feedback mechanism ensures that feature extraction preserves critical information while filtering out noise, improving measurement precision without significant information loss.
Solution Approach 2:
The system performs preliminary feature selection and weighting based on eye tracking data before the main classification process. By pre-processing the image data to highlight regions of interest according to human visual patterns, the system reduces the dimensionality of the problem early in the pipeline, preserving only the most informative features and minimizing information loss during subsequent processing.
Data Source
AI summary
Disclosed are various embodiments for a visual language processing modeling framework via an attention-on-attention mechanism, which may be employed for object identification, classification, and the like. In association with a display of a user interface, an eye tracking via images captured by an imaging device is performed to programmatically detect eye movement and fixation relative to sub-regions of the user interface. Eye fixations on at least one of the sub-regions from the eye tracking. Visual cues are extracted from the user interface based at least in part on the eye fixations, the visual cues being in a sequence of identification. A visual language sentence is generated based at least in part on the visual cues as extracted. The visual language sentence of the visual cues in the sequence of identification is correlated to at least one decision using a visual language understanding routine.


