Text-to-Speech File Size Prediction Before Conversion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current text-to-speech conversion technologies face challenges in predicting storage space requirements for converted audio files, particularly on portable devices, due to variations in file types and characteristics, leading to uncertainty in how much data can be stored or how much playing time is feasible on available storage.

Innovation Solution

A method and apparatus that predict the resultant attribute of a text file before conversion by determining the file type and size, using a calculator component to estimate the audio file size and playing time, allowing users to decide how much data fits within available storage space.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If text is split into phonemes for conversion, then speech output quality is improved, but computational complexity and storage space requirements increase

Engineering Contradiction:
Improvespeech output qualityVSAvoidcomputational complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary analysis of the text file characteristics (word count, character count, file type) before conversion to predict the resulting audio file size. This allows users to understand storage requirements in advance without actually performing the complex conversion process, thereby avoiding unnecessary computational complexity while maintaining speech quality through the phoneme method.

Inventive Principle:
Principle #10Preliminary action

2Manufacturing precision

If text is split into phonemes for conversion, then speech output quality is improved, but storage space requirements increase

Engineering Contradiction:
Improvespeech output qualityVSAvoidstorage space
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The system calculates and presents predicted audio file sizes based on text file characteristics before actual conversion occurs. Users can use these predictions to determine how much storage space is needed and make informed decisions about which files to convert, thereby managing storage space efficiently while maintaining high speech quality through the phoneme conversion method.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses different prediction models and calculation parameters for different file types (e.g., .txt, .doc, .pdf). By adjusting the prediction parameters based on the specific file type and its characteristics, the system provides accurate storage space predictions without requiring actual conversion, thus managing storage requirements while maintaining speech quality.

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If diphone method is used to split phrases, then synthetic speech quality is improved, but storage space requirements increase

Engineering Contradiction:
Improvesynthetic speech qualityVSAvoidstorage space
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The system performs preliminary calculation of predicted audio file sizes based on text characteristics before conversion. This allows users to understand the storage space requirements in advance for diphone method conversions, enabling them to make informed decisions about which files to convert and how to manage storage space while maintaining high synthetic speech quality.

Inventive Principle:
Principle #10Preliminary action

4Quantity of substance

If digitized speech method is used, then storage space requirements are reduced, but speech naturalness decreases

Engineering Contradiction:
Improvestorage spaceVSAvoidspeech naturalness
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The system calculates predicted audio file sizes based on text file characteristics before conversion occurs. This allows users to understand storage requirements in advance and make informed decisions about which files to convert using the digitized speech method (which has lower storage requirements) versus the phoneme method (which produces more natural speech), thereby balancing storage space constraints with speech naturalness requirements.

Inventive Principle:
Principle #10Preliminary action

5Adaptability or versatility

If text files of different types are converted, then conversion versatility is improved, but prediction accuracy becomes more difficult

Engineering Contradiction:
Improveconversion versatilityVSAvoidprediction accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The system applies different prediction models and calculation parameters for different file types (e.g., .txt, .doc, .pdf). Each file type has its own specific characteristics (word count, character count, formatting) that are taken into account when making predictions. This localized approach to prediction for each file type maintains high prediction accuracy while supporting versatile conversion capabilities across multiple file formats.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system adjusts prediction parameters and calculation methods based on the specific file type being converted. By changing the prediction model parameters according to file type characteristics, the system maintains accurate predictions across diverse file formats while preserving conversion versatility. Each file type receives customized prediction treatment based on its specific properties.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8145490B2Predicting a resultant attribute of a text file before it has been converted into an audio file
Publication Date: 2012.03.27 CERENCE OPERATING CO
  • US8145490B2 patent drawing
  • US8145490B2 patent drawing
  • US8145490B2 patent drawing

AI summary

An apparatus for predicting a resultant attribute of a text file before it has been converted to an audio file by a text-to-speech converter application. In accordance with an embodiment, the apparatus includes: a receiver component for receiving a text file and a request to determine a resultant attribute of the text file before it is converted to an audio file, by a text-to-speech converter component; a calculation component for determining a file type associated with the received text file and the size of the received text file; a calculation component for identifying an attribute associated with the determined file type; and a calculation component for determining from the identified attribute and the size of the received text file a resultant attribute of the text file before it is converted to an audio file by the text-to-speech converter component.