Speech-Guided Image Generation Using Semantic Extraction and AI Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems face limitations in accuracy and consistency when generating Internet images and videos, suffer from data acquisition issues due to crawling prevention mechanisms, data quality and copyright problems, and pose risks to personal information security and recognition errors.

Innovation Solution

A system and method utilizing speech recognition, language understanding, and cloud image generation through stability diffusion algorithms to create personalized images and videos based on user speech inputs, ensuring copyright compliance and enhanced user experience.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If web crawling technology is used to search for Internet images and video resources, then data acquisition capability is improved, but crawling prevention mechanisms block data acquisition

Engineering Contradiction:
Improvedata acquisition capabilityVSAvoiddata acquisition success rate
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent introduces a speech recognition intermediary that converts user speech into text queries, which then guide the image search process. This intermediary layer allows the system to bypass traditional web crawling limitations by directly generating images based on speech inputs, rather than attempting to crawl protected Internet resources.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical web crawling system with an AI-based image generation system driven by speech recognition. Instead of mechanically crawling through protected web pages, the system uses speech-to-text conversion followed by AI image generation, substituting the entire crawling mechanism with a more reliable alternative.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If machine learning models are introduced to enhance speech recognition accuracy, then recognition accuracy is improved, but system complexity increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the speech recognition system into distinct functional modules: speech input module, preliminary processing module, digital signal conversion module, feature extraction module, speech recognition module, and postprocessing module. Each module handles a specific aspect of speech processing, which improves overall accuracy while making the complex system more manageable and maintainable.

Inventive Principle:
Principle #1Segmentation

3Adaptability or versatility

If Internet image resources are used for speech-based search, then content availability is improved, but copyright infringement risks increase

Engineering Contradiction:
Improvecontent availabilityVSAvoidcopyright infringement risk
Core Design Contradiction:
Adaptability or versatilityVSObject-affected harmful factors

Solution Approach 1:

The patent converts the potential harm of copyright infringement into a benefit by using AI to generate original images based on speech inputs. Instead of copying protected Internet content, the system generates unique images that inherently avoid copyright issues, turning the limitation into a competitive advantage for legal safety.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

4Ease of operation

If speech recognition systems process user information, then user interaction capability is improved, but personal information security risks increase

Engineering Contradiction:
Improveuser interaction capabilityVSAvoidpersonal information security risk
Core Design Contradiction:
Ease of operationVSObject-affected harmful factors

Solution Approach 1:

The patent extracts only the essential semantic information from user speech inputs while discarding personally identifiable information. The speech recognition system processes speech to extract intent and meaning without retaining or storing sensitive personal data, thereby maintaining user interaction capability while mitigating security risks.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20260044997A1Speech recognition-based image generation system and method
Publication Date: 2026.02.12 HYUNDAI MOTOR CO LTD
  • US20260044997A1 patent drawing
  • US20260044997A1 patent drawing
  • US20260044997A1 patent drawing

AI summary

Disclosed are a system and a method of generating an image based on speech recognition. The system of generating an image based on speech recognition includes: a speech recognition apparatus configured to acquire speech information of a user, and convert the acquired speech information into text type user requirement information; a language understanding apparatus electrically connected to the speech recognition apparatus, and configured to analyze the text type user requirement information and extract semantic information of the user; a cloud image generation apparatus connected to be communicable with the language understanding apparatus, and configured to generate image data based on the extracted semantic information of the user by using a stability diffusion algorithm; and a display apparatus connected to be communicable with the cloud image generation apparatus, and configured to receive an image generated from the cloud image generation apparatus, and decode and display the received image.