Gesture-Based Multimodal Content Creation Screens for Generative AI Editing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional text-based prompt input methods for content creation are inefficient, requiring sequential inputs and re-entering of information, leading to time-consuming searches and inconsistent collaborative work due to geographical separation, resulting in reduced quality and resource waste.
Innovation Solution
A content creation screen provision method that allows users to input multiple types of data (images, videos, text) via predefined gestures, enabling seamless collaboration and editing through generative AI, facilitating efficient content creation and editing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If text-based prompt input methods are used, then users can input information in a standardized format, but users are required to input various types of information (images, videos, text) in text form only, making it difficult to obtain refined deliverables that combine different types of information
Solution Approach 1:
The prompt input area is designed to accept multiple types of data including images, videos, text, and audio files directly, rather than requiring conversion to text only. This multi-functional input capability allows users to upload diverse media types seamlessly, resolving the contradiction between handling multiple data types and ease of operation.
Solution Approach 2:
The system introduces an intermediary processing layer that automatically converts and integrates different data types (images, videos, audio) into a unified prompt format that can be processed by the generative AI model. This mediator enables seamless handling of multiple formats without requiring manual text conversion from users.
2Ease of operation
If sequential, time-series inputs are used, then users can input information step by step, but when a large amount of accumulated information exists, users may spend significant time searching for specific prompts
Solution Approach 1:
The system automatically manages and organizes prompts in advance by maintaining a structured repository of historical prompts and their associated data. When users need to reference or modify previous prompts, the system retrieves them automatically based on contextual cues or user selections, eliminating the need for manual searching through accumulated information.
Solution Approach 2:
The system provides feedback mechanisms that display relevant historical prompts and their usage context, allowing users to quickly identify and select appropriate previous prompts. This feedback loop reduces search time by presenting relevant options based on current input context and user behavior patterns.
3Adaptability or versatility
If multiple users collaborate remotely, then diverse perspectives and expertise can be combined, but challenges in communication arise due to geographical separation or time zone differences, leading to inconsistent work and project delays
Solution Approach 1:
The system creates and maintains synchronized copies of the content creation state across all user devices in real-time. When one user modifies prompts or generated content, these changes are automatically replicated to all other collaborators' interfaces, ensuring everyone works with the same current information regardless of location or time zone, thus maintaining consistency in collaborative efforts.
4Ease of manufacture
If users are required to re-enter previously generated outputs for modification, then precise control over content can be achieved, but the inconvenience of re-entering previously generated outputs increases time consumption
Solution Approach 1:
The system automatically preserves and structures previously generated outputs in an editable format within the prompt input area. When users wish to modify generated content, the system retrieves the original output and presents it pre-formatted and ready for editing, eliminating the need to re-enter content manually while maintaining precise control over modifications.
Data Source
AI summary
A content creation screen provision method and system thereof are provided. The method may include receiving a predefined first gesture for a first plurality of objects displayed on a content creation screen of the computing device, displaying the first plurality of objects in a prompt input area on the content creation screen, transmitting, to a service server, a first prompt generation request for creating content related to the first plurality of objects in response to a third gesture for content creation, receiving first content generated by generative artificial intelligence based on a first prompt input from the service server, displaying the first content on the content creation screen, receiving second content generated by the generative AI based on a second prompt input from the service server, and displaying the second content on the content creation screen.


