Voice-Controlled Image Processing via NLP Conversion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face a time-consuming and cumbersome experience when processing images using software like Photoshop, requiring manual learning and input of instructions for retouch operations.

Innovation Solution

An image processing device and method that converts voice signals into image processing instructions and target areas using voice recognition, natural language processing, and image recognition technologies, allowing users to process images through voice commands without prior software knowledge.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If users manually learn and input instructions in traditional image processing software, then precise image processing control is achieved, but user time consumption and operational complexity increase significantly

Engineering Contradiction:
Improveease of image processing operationVSAvoidtime consumption for learning and manual input
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical operations (typing instructions, navigating menus) with voice-based acoustic input. Users speak natural language commands which are converted to processing instructions through voice recognition and NLP, eliminating the need to manually navigate software interfaces or learn complex command structures.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces voice recognition technology and natural language processing models as intermediaries between the user and the image processing system. These intermediaries translate spoken language into structured processing instructions, serving as a mediator that bridges natural human communication and machine-executable commands without requiring users to learn software-specific syntax.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If traditional manual instruction input methods are used, then precise control over image processing is achieved, but user experience and operational efficiency deteriorate

Engineering Contradiction:
Improveimage processing efficiencyVSAvoidcomplexity of instruction input process
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system replaces complex mechanical interaction (manual navigation, menu selection, parameter adjustment) with voice-based acoustic control. This substitution maintains processing precision while dramatically simplifying the interaction model, allowing users to issue multiple processing commands through natural speech without navigating through software layers.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The voice command system serves multiple functions simultaneously: it captures user intent, identifies target regions, determines processing operations, and adjusts parameters—all through a single unified interface. This multi-functionality consolidates what would otherwise require separate manual operations into one cohesive voice-based control mechanism.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP3667487B1Image processing apparatus and method
Publication Date: 2023.11.15 SHANGHAI CAMBRICON INFORMATION TECH CO LTD
  • EP3667487B1 patent drawingFigure 1~4
  • EP3667487B1 patent drawingFigure 5~6
  • EP3667487B1 patent drawingFigure 7~8

AI summary

The present disclosure discloses an image processing device including: a receiving module configured to receive a voice signal and an image to be processed; a conversion module configured to convert the voice signal into an image processing instruction and a target area according to a target voice instruction conversion model, in which the target area is a processing area of the image to be processed; and a processing module configured to process the target area according to the image processing instruction and a target image processing model. The examples may realize a functionality of inputting voice to process images, which may save users' time spent in learning image processing software prior to image processing, and improve user experience.