Operating System Visual Search With On-Device Machine Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users face difficulties in obtaining additional information associated with visual data across different applications and media files due to the limitations of text searching, screenshot capture, and screenshot cropping, often leading to irrelevant search results and computational inefficiencies.
Innovation Solution
A visual search interface at the operating system level that utilizes on-device machine-learned models for object detection, optical character recognition, and segmentation, enabling seamless data processing across applications without relying on server computing systems, thereby reducing computational costs and enhancing privacy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If text searching is used to find information across applications, then search functionality is provided, but search accuracy and relevance deteriorate due to difficulty in determining appropriate search words
Solution Approach 1:
The patent replaces traditional text-based mechanical search systems with vision-based machine learning models. Instead of requiring users to input text queries, the system captures images of displayed content and uses computer vision algorithms to automatically extract and search for information, substituting optical recognition for text input mechanisms
Solution Approach 2:
The system enables self-service by automatically capturing screenshots, performing optical character recognition, and executing searches without requiring user intervention for query formulation. The visual search interface autonomously processes displayed content and generates search results based on extracted visual information
2Loss of information
If screenshot capture is used to obtain visual data, then visual content can be captured, but computational resources and processing time increase due to server-based processing requirements
Solution Approach 1:
The patent segments the visual search system into on-device components (display capture, object detection model, segmentation model) that can process images locally without requiring complete server-based processing. This segmentation enables partial local processing that reduces computational resource consumption while maintaining search functionality
Solution Approach 2:
The patent introduces an intermediary layer of on-device machine learning models that bridge between screenshot capture and server processing. These models pre-process and analyze visual data locally, filtering and preparing information before potential server submission, thereby reducing the computational burden on both device and server
3Productivity
If visual search data is transmitted to server computing systems for processing, then comprehensive search results can be obtained, but user privacy and data security are compromised
Solution Approach 1:
The system performs self-service visual search processing through on-device machine learning models that can independently analyze captured images and generate search results without transmitting visual data to external servers. This self-contained processing capability maintains privacy while delivering functional search results
Solution Approach 2:
The patent segments search functionality into on-device processing components that handle privacy-sensitive visual data locally. By dividing the search system into local and remote components, the patent enables privacy-preserving processing for certain operations while maintaining access to comprehensive search capabilities when needed
4Measurement precision
If multiple user inputs are required for screenshot capture and cropping, then precise visual selection can be achieved, but ease of operation deteriorates due to complex input requirements
Solution Approach 1:
The system provides self-service by automatically capturing and processing visual content without requiring manual user inputs for screenshot capture and cropping. The display capture component autonomously acquires visual data from the screen, eliminating the need for users to perform precise manual selections
Solution Approach 2:
The patent replaces manual mechanical user interactions (clicking, dragging, cropping) with automated display capture technology. The system automatically captures and processes visual content from the display, substituting automated optical recognition for manual user manipulation
Data Source
AI summary
Visual search in an operating system of a computing device can process and provide additional information on the content being provided for display. The computing device can include an operating system that includes a visual search interface that obtains and processes display data associated with content currently being provided for display. The visual search interface can generate display data based on the current content provided for display, process the display data with one or more on-device machine-learned models, and provide additional information to the user. The visual search interface may transmit data associated with the display data to perform additional data processing tasks. Application suggestions may be determined and provided based on the visual search data.


