Text-to-Image Conversion via Neural Feature Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data processing technologies face challenges in efficiently compressing and transmitting image data, requiring significant computing and network resources, and existing neural network models for text-to-image and image-to-text conversions have low calculation accuracy and processing efficiency.
Innovation Solution
A method and device that determine character and text features from input text, generate initial and target visual features, and create a corresponding image, allowing for efficient data compression by converting images to text and vice versa, improving accuracy and processing efficiency through neural network models like transformers and CNNs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If image data is transmitted directly, then image quality is preserved, but computing resources and network resources are significantly increased
Solution Approach 1:
The patent creates a text representation copy of the image data that can be transmitted and processed with significantly lower resource consumption. Instead of transmitting the original image data, the system transmits and processes text features that represent the image content, achieving resource efficiency while maintaining functional equivalence for the intended application.
Solution Approach 2:
The patent transforms image data into text features by changing the data representation parameters. The visual information is converted into textual feature vectors that can be more efficiently compressed, transmitted, and processed, thereby reducing computing and network resource requirements while preserving the essential information needed for the application.
2Use of energy by moving object
If image data is subjected to lossy compression or lossless compression, then resource consumption is reduced, but processing accuracy and fidelity are compromised
Solution Approach 1:
The patent replaces traditional image compression mechanisms with a text-based processing system. Instead of compressing image pixels directly, the system converts images to text features, processes them through neural networks, and generates text representations that can be accurately converted back to images, achieving both efficiency and accuracy.
Solution Approach 2:
The patent introduces text features as an intermediary representation between the original image data and the processed output. This intermediary text representation layer enables efficient compression and processing while maintaining the ability to accurately reconstruct the image, thus resolving the trade-off between resource consumption and processing accuracy.
3Productivity
If existing neural network models are used for text-to-image and image-to-text conversions, then data processing is enabled, but calculation accuracy and processing efficiency are low
Solution Approach 1:
The patent segments the text-to-image and image-to-text conversion process into distinct modules with specialized functionality. The text-to-image module includes character feature extraction, text feature extraction, and visual feature generation components, while the image-to-text module includes image feature extraction and text generation components. This segmentation enables each module to be optimized for its specific task, improving both accuracy and efficiency.
Solution Approach 2:
The patent introduces multiple feature dimensions including character-level features, text-level features, and visual features that operate in different dimensional spaces. By processing information through these multiple dimensions simultaneously, the system achieves higher calculation accuracy and processing efficiency compared to single-dimensional approaches.
Data Source
AI summary
Embodiments of the present disclosure relate to a method, a device, and a computer program product for processing data. The method includes determining, based on acquired text, character features for a group of characters in the text and text features for the text. The method further includes determining initial visual features for the text based on the text features. The method further includes determining target visual features for the text based on the initial visual features, the character features, and the text features. The method further includes generating a target image corresponding to the text based on the target visual features. Through the method, the accuracy of conversion between text and an image is improved, the data processing efficiency is improved, and the data compression efficiency is further improved.


