Operating System Visual Search With On-Device Machine Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face difficulties in obtaining additional information associated with visual data across different applications and media files due to the limitations of text searching, screenshot capture, and screenshot cropping, often leading to irrelevant search results and computational inefficiencies.

Innovation Solution

A visual search interface at the operating system level that utilizes on-device machine-learned models for object detection, optical character recognition, and segmentation, enabling seamless data processing across applications without relying on server computing systems, thereby reducing computational costs and enhancing privacy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If text searching is used to find information across applications, then search functionality is provided, but search accuracy and relevance deteriorate due to difficulty in determining appropriate search words

Engineering Contradiction:
Improvesearch functionalityVSAvoidsearch accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent replaces traditional text-based mechanical search systems with vision-based machine learning models. Instead of requiring users to input text queries, the system captures images of displayed content and uses computer vision algorithms to automatically extract and search for information, substituting optical recognition for text input mechanisms

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service by automatically capturing screenshots, performing optical character recognition, and executing searches without requiring user intervention for query formulation. The visual search interface autonomously processes displayed content and generates search results based on extracted visual information

Inventive Principle:
Principle #25Self-service

2Loss of information

If screenshot capture is used to obtain visual data, then visual content can be captured, but computational resources and processing time increase due to server-based processing requirements

Engineering Contradiction:
Improvevisual data acquisitionVSAvoidcomputational resources
Core Design Contradiction:
Loss of informationVSUse of energy by moving object

Solution Approach 1:

The patent segments the visual search system into on-device components (display capture, object detection model, segmentation model) that can process images locally without requiring complete server-based processing. This segmentation enables partial local processing that reduces computational resource consumption while maintaining search functionality

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary layer of on-device machine learning models that bridge between screenshot capture and server processing. These models pre-process and analyze visual data locally, filtering and preparing information before potential server submission, thereby reducing the computational burden on both device and server

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If visual search data is transmitted to server computing systems for processing, then comprehensive search results can be obtained, but user privacy and data security are compromised

Engineering Contradiction:
Improvesearch capabilityVSAvoidprivacy risk
Core Design Contradiction:
ProductivityVSObject-affected harmful factors

Solution Approach 1:

The system performs self-service visual search processing through on-device machine learning models that can independently analyze captured images and generate search results without transmitting visual data to external servers. This self-contained processing capability maintains privacy while delivering functional search results

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent segments search functionality into on-device processing components that handle privacy-sensitive visual data locally. By dividing the search system into local and remote components, the patent enables privacy-preserving processing for certain operations while maintaining access to comprehensive search capabilities when needed

Inventive Principle:
Principle #1Segmentation

4Measurement precision

If multiple user inputs are required for screenshot capture and cropping, then precise visual selection can be achieved, but ease of operation deteriorates due to complex input requirements

Engineering Contradiction:
Improvevisual selection accuracyVSAvoidoperation complexity
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system provides self-service by automatically capturing and processing visual content without requiring manual user inputs for screenshot capture and cropping. The display capture component autonomously acquires visual data from the screen, eliminating the need for users to perform precise manual selections

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical user interactions (clicking, dragging, cropping) with automated display capture technology. The system automatically captures and processes visual content from the display, substituting automated optical recognition for manual user manipulation

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20250335498A1Visual Search Interface in an Operating System
Publication Date: 2025.10.30 GOOGLE LLC
  • US20250335498A1 patent drawing
  • US20250335498A1 patent drawing
  • US20250335498A1 patent drawing

AI summary

Visual search in an operating system of a computing device can process and provide additional information on the content being provided for display. The computing device can include an operating system that includes a visual search interface that obtains and processes display data associated with content currently being provided for display. The visual search interface can generate display data based on the current content provided for display, process the display data with one or more on-device machine-learned models, and provide additional information to the user. The visual search interface may transmit data associated with the display data to perform additional data processing tasks. Application suggestions may be determined and provided based on the visual search data.