Multimodal Interface Prediction Model Using Intermediate Embeddings

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional machine learning techniques face difficulties in effectively leveraging multimodal interface features and achieving strong performance for user interface prediction and generation, especially when high-quality labeled data is unavailable, making it challenging to develop efficient and accurate machine-learned models for user interface understanding.

Innovation Solution

A computer-implemented method for training and utilizing machine-learned models involves obtaining interface data, determining intermediate embeddings from structural data, interface images, and textual content, and processing these with a machine-learned interface prediction model to obtain user interface embeddings, followed by pre-training tasks to generate pre-training outputs, enabling the model to predict and generate user interfaces efficiently.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional machine learning techniques are used for user interface prediction, then the system can process interface data, but the model fails to effectively leverage multimodal interface features and achieves weak performance

Engineering Contradiction:
Improveprediction accuracyVSAvoidmodel complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the user interface into multiple modalities (visual, textual, structural) and processes each through dedicated embedding layers. The interface image is processed through a visual embedding layer, textual content through a textual embedding layer, and structural metadata through a structural embedding layer, with each segment transformed into intermediate embeddings that are then aggregated.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a composite representation by combining intermediate embeddings from multiple modalities (visual, textual, structural) into a unified user interface embedding. This composite approach integrates heterogeneous data types through a fusion mechanism that aggregates the intermediate embeddings, enabling the model to leverage synergistic information from all modalities simultaneously.

Inventive Principle:
Principle #40Composite materials

2Measurement precision

If high-quality labeled data is used for training, then the model achieves strong performance, but such data is commonly unavailable for user interfaces

Engineering Contradiction:
Improvemodel performanceVSAvoidlabeled data availability
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent performs preliminary action by pre-training the machine-learned model on a large corpus of user interface data before fine-tuning on task-specific labeled data. The pre-training phase initializes the model with general user interface understanding, enabling it to achieve reasonable performance even when task-specific labeled data is limited or unavailable.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If multimodal interface features are leveraged, then the model achieves better performance, but the processing becomes prohibitively difficult

Engineering Contradiction:
Improveinterface prediction accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces intermediate embeddings as intermediary representations that bridge the gap between raw multimodal inputs and the final user interface embedding. Each modality (visual, textual, structural) is first transformed into intermediate embeddings through dedicated embedding layers, which then serve as intermediaries that are aggregated to form the final representation, simplifying the integration process.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20240169186A1Machine-Learned Models for User Interface Prediction and Generation
Publication Date: 2024.05.23 GOOGLE LLC
  • US20240169186A1 patent drawing
  • US20240169186A1 patent drawing
  • US20240169186A1 patent drawing

AI summary

Generally, the present disclosure is directed to user interface understanding. More particularly, the present disclosure relates to training and utilization of machine-learned models for user interface prediction and/or generation. A machine-learned interface Nprediction model can be pre-trained using a variety of pre-training tasks for eventual downstream task training and utilization (e.g., interface prediction, interface generation, etc.).