Smart Capture System for Cross-App Data Mashup and Input Suggestions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing user device systems face challenges in providing efficient input suggestions across multiple applications and screens, particularly in situations requiring data sharing and entry, due to limitations in image analysis, screen content interpretation, and lack of dynamism in action suggestions, leading to inefficient user experiences.
Innovation Solution
The implementation of a user device system that collects data from various sources, uses a data mashup model to identify and classify content types, determine relationships, and predict actions, providing suggestions through deep screen capture and logical tree structures to enhance user experience by automating data sharing and entry across applications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Difficulty of detecting and measuring
If image analysis-based approaches are used to extract screen content, then content can be captured from visual displays, but the system becomes complex and inference time increases to 3.98 seconds on CPU
Solution Approach 1:
The patent replaces complex image analysis-based mechanical extraction systems with an intelligent system that uses logical tree structures and relationship extraction. Instead of relying on pixel-level image processing, the system uses structured data representation and semantic understanding to extract and interpret screen content, thereby reducing computational complexity while maintaining extraction capability.
Solution Approach 2:
The patent changes the fundamental parameter of content extraction from pixel-based image analysis to structured data-based logical tree analysis. By transforming the input representation from images to structured data trees with semantic relationships, the system achieves faster processing and reduced complexity while maintaining the ability to extract meaningful screen content.
2Difficulty of detecting and measuring
If conventional screenshot-based boundary extraction is used, then screen content can be captured, but it lacks intelligence based on screen type and cannot handle ongoing conversations efficiently
Solution Approach 1:
The patent introduces dynamic adaptation to different screen types through intelligent classification. The system automatically identifies the type of screen (e.g., conversation interface, form interface, media player) and adjusts its extraction and interpretation strategies accordingly. This dynamic behavior enables efficient handling of ongoing conversations and various screen types without requiring manual configuration.
Solution Approach 2:
The patent introduces an intermediary classification layer between raw screenshot capture and content extraction. This intermediary component analyzes screen type characteristics and routes the extraction process through appropriate specialized handlers, enabling intelligent adaptation to different screen types while maintaining a unified capture mechanism.
3Extent of automation
If action suggestion models are trained on remote servers and pushed to devices, then predictions can be made, but the suggestions lack dynamism and do not consider user response or other device data
Solution Approach 1:
The patent implements feedback loops where user responses to suggested actions are collected and used to refine future suggestions. The system learns from user behavior patterns and adjusts its action prediction model dynamically, considering both historical data and real-time user preferences. This feedback mechanism enables the suggestions to become more accurate and personalized over time.
Solution Approach 2:
The patent performs preliminary analysis of multiple data sources including other device data, application states, and user context before generating action suggestions. By pre-processing and consolidating relevant information from across the device ecosystem, the system prepares a comprehensive context that enables more accurate and personalized action predictions rather than relying solely on remote model predictions.
4Adaptability or versatility
If users manually copy and paste data across applications, then information can be shared, but the process is time-consuming and requires constant switching between applications
Solution Approach 1:
The patent merges data extraction capabilities across multiple applications by analyzing screen content and identifying extractable information regardless of which application is currently active. The system consolidates data from various sources including notifications, clipboard, and application screens into a unified structure, enabling automatic population of forms and fields across different applications without requiring manual copy-paste operations.
Solution Approach 2:
The patent enables the system to automatically perform data extraction, consolidation, and transfer operations without requiring user intervention. The intelligent system autonomously identifies when data should be shared between applications, extracts the relevant information from screen content, and populates target fields automatically, thereby eliminating the time-consuming manual copy-paste process.
5Difficulty of detecting and measuring
If text-based or view-based techniques are used to retrieve information, then data can be accessed from input fields, but these techniques fail when screens contain mainly images or non-editable fields
Solution Approach 1:
The patent creates a universal extraction system that can handle multiple screen formats including text-based interfaces, image-dominated interfaces, and non-editable fields. The system uses a combination of optical character recognition, image analysis, and logical tree structure generation to extract information from diverse screen types, making the extraction capability format-agnostic and applicable to social media platforms, forms, and other varied interfaces.
Data Source
AI summary
Example systems and methods provide input suggestions to a user to improve user experience on user devices. The input suggestions can be fill information from another app on device to the present app being used by user, information for performing a search (without the user having to copy-paste data or entering the data manually), responses to a message/notification received by the user, information/content/data to be shared between apps (without switching between apps), and emojis/GIFs that can be used by the user. The method includes analyzing one or more content of one or more screen displayed on device, generating at least one of a logical tree structure and a data mashup model of the one or more analyzed content for each screen, and providing a recommendation to a user. The recommendation can be a connected action or an input suggestion.


