Speech Synthesis Using Visual Feature Vectors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech synthesis techniques based on deep neural networks struggle to generate natural-sounding speech when reading books with images, as they primarily rely on linguistic information and lack the ability to incorporate visual and contextual elements, resulting in a difference in naturalness compared to human narration.
Innovation Solution
A speech synthesis apparatus and method that incorporates both linguistic and visual information, using a neural network with two inputs: a language vector derived from text and a visual feature vector extracted from image information, to generate synthesized speech that considers the visual elements of a book, such as illustrations and character descriptions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If speech synthesis is based only on linguistic information from text, then the synthesis process is simple and fast, but the naturalness and cadence of the synthesized speech deteriorates when reading books with images
Solution Approach 1:
The patent merges linguistic information from text with visual information from images by combining their respective feature vectors into a unified input for the speech synthesis model. This allows the system to leverage both text and image data to generate more natural-sounding speech that captures the intended cadence and emotion.
Solution Approach 2:
The patent introduces an intermediary processing stage where visual information from images is extracted and converted into a feature vector that can be integrated with linguistic features. This intermediary representation enables the speech synthesis model to indirectly incorporate visual context without directly processing raw image data.
2Reliability
If visual information from images is incorporated into speech synthesis, then the naturalness of synthesized speech improves, but the processing complexity and computational load increases
Solution Approach 1:
The patent extracts only the essential visual features from images that are relevant to speech synthesis, such as features related to character expressions, scene context, or emotional cues. This selective extraction reduces the amount of visual data that needs to be processed while retaining the information necessary to improve speech naturalness.
Solution Approach 2:
The patent performs preliminary processing of visual information by pre-extracting and encoding image features into a compact representation before they are integrated with linguistic information. This preliminary action prepares the visual data in advance, reducing the computational burden during the actual speech synthesis process.
3Ease of manufacture
If only text information is used for speech synthesis, then the system is easier to implement, but it cannot capture the emotional and contextual nuances present in illustrated books
Solution Approach 1:
The patent creates a multi-functional speech synthesis system that can process both text and image inputs. The unified model architecture allows the same system to handle pure text synthesis and text-with-image synthesis, making the system versatile without requiring separate processing pipelines for different input types.
Data Source
AI summary
A speech synthesis apparatus according to the present disclosure includes a memory and a processor coupled to the memory. The processor is configured to: obtain utterance information on subjects to be uttered, wherein the subjects to be uttered are texts contained in data on a book, obtain image information on images that are contained in the data on the book, obtain speech data corresponding to the subjects to be uttered; and generate, based on the obtained utterance information, the obtained image information, and the obtained speech data, a speech synthesis model for reading out a text associated with an image.


