Multimodal Neural Network for Gesture-Based Image Editing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional digital content editing systems face issues with accuracy, efficiency, and flexibility due to rule-based methods, requiring excessive computing resources and user input, and being inflexible in adapting to different devices and applications.

Innovation Solution

A multimodal selection system that uses neural networks to select and edit digital images based on verbal and gesture inputs, combining natural language processing with computer vision to accurately identify and modify objects, and dynamically deploys components across various devices, allowing for flexible operation and extensibility.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If rule-based methods are used to identify and modify objects, then the system structure is simple, but the accuracy deteriorates when rules do not apply to particular images

Engineering Contradiction:
Improvesystem structureVSAvoidobject identification accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent replaces rule-based mechanical systems with neural network-based systems. Specifically, it uses a neural network to identify objects in digital images, replacing traditional rule-based object identification methods. This substitution enables the system to handle complex, nuanced image recognition tasks that rule-based systems cannot accurately perform, thereby improving measurement precision without significantly increasing overall system complexity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of operation

If conventional digital content editing systems are used, then basic editing functions are available, but excessive computing resources and user input are required

Engineering Contradiction:
Improveuser input requirementVSAvoidediting efficiency
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The patent implements self-service by enabling the system to automatically perform editing functions based on natural language commands. The neural network interprets user intent and automatically identifies objects and applies appropriate edits without requiring manual selection or multiple interaction steps. This reduces the burden of user input while maintaining high editing efficiency, as the system serves itself by autonomously completing the editing workflow.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary action by pre-processing and pre-identifying objects in the image before the user issues editing commands. The neural network is ready to quickly identify and segment objects when a command is given, eliminating the need for users to manually select objects first. This preliminary object identification prepares the system in advance, making the actual editing process faster and more efficient.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If conventional digital content editing systems are used, then basic editing capabilities are provided, but accuracy deteriorates in handling ill-defined or generalized commands

Engineering Contradiction:
Improvecommand interpretation flexibilityVSAvoidobject identification accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent applies parameter changes by transforming the neural network's internal parameters (weights and biases) through training on diverse datasets. This allows the network to adapt to various command types and image characteristics. The system can accurately interpret ill-defined or generalized commands because the neural network has learned flexible parameter representations during training, enabling it to generalize across different scenarios while maintaining high object identification accuracy.

Inventive Principle:
Principle #35Parameter changes

4Ease of manufacture

If conventional digital content editing systems are used, then fixed functionality is provided, but flexibility deteriorates in adapting to different devices and applications

Engineering Contradiction:
Improvesystem deployment simplicityVSAvoiddevice and application adaptability
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent implements universality by designing a neural network-based editing system that can function across multiple devices and applications. The core neural network architecture remains consistent, but it can be deployed on various platforms (mobile devices, desktops, web browsers) and integrated with different applications. This universal design allows the system to adapt to different devices and applications without requiring fundamental redesign, maintaining ease of deployment while achieving broad adaptability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11594077B2Generating modified digital images utilizing a dispersed multimodal selection model
Publication Date: 2023.02.28 ADOBE INC
  • US11594077B2 patent drawing
  • US11594077B2 patent drawing
  • US11594077B2 patent drawing

AI summary

The present disclosure relates to systems, methods, and non-transitory computer readable media for generating modified digital images based on verbal and/or gesture input by utilizing a natural language processing neural network and one or more computer vision neural networks. The disclosed systems can receive verbal input together with gesture input. The disclosed systems can further utilize a natural language processing neural network to generate a verbal command based on verbal input. The disclosed systems can select a particular computer vision neural network based on the verbal input and/or the gesture input. The disclosed systems can apply the selected computer vision neural network to identify pixels within a digital image that correspond to an object indicated by the verbal input and/or gesture input. Utilizing the identified pixels, the disclosed systems can generate a modified digital image by performing one or more editing actions indicated by the verbal input and/or gesture input.