Automated Audio Description Generation via Fragment Stitching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems for creating audio descriptions of products or processes require human input or scripted content, limiting their efficiency and scalability in generating natural-sounding voiceovers for various products or processes.

Innovation Solution

A method that uses a plurality of human voice recordings, automatically selects and stitches together audio fragments based on electronically obtained attribute values, eliminating the need for human intervention and script input, to create natural-sounding voiceover descriptions for products or processes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If human voice recordings are manually selected and stitched, then the audio description quality is high, but the time consumption and labor requirements increase significantly

Engineering Contradiction:
Improveaudio description qualityVSAvoidtime consumption
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent segments the voiceover creation process into discrete audio fragments corresponding to different product attributes. Each fragment is pre-recorded with proper prosody and can be independently selected and stitched together, transforming a manual process into an automated assembly process that reduces time consumption while maintaining quality

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system automatically selects and stitches audio fragments based on electronically obtained attribute values without requiring manual human intervention. The automated selection process uses rules to match audio fragments with product attributes, enabling the system to serve itself and eliminate time-consuming manual labor

Inventive Principle:
Principle #25Self-service

2Productivity

If automated systems are used to generate audio descriptions, then efficiency and scalability improve, but the natural-sounding quality and context recognition may deteriorate

Engineering Contradiction:
ImproveefficiencyVSAvoidnatural-sounding quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent performs preliminary action by having humans record audio fragments with proper prosody and context in advance. These pre-recorded fragments are then automatically assembled, combining the natural quality of human recording with the efficiency of automated processing. The preliminary recording phase captures authentic human speech patterns that automated synthesis cannot replicate

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system copies and reuses authentic human voice recordings rather than attempting to synthesize speech. By selecting and stitching existing human-recorded audio fragments, the system preserves the natural quality and context recognition of human speech while achieving automated production efficiency

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If human input and script input are required, then the audio description can be customized, but the scalability and ability to generate thousands of descriptions is limited

Engineering Contradiction:
Improvecustomization capabilityVSAvoidscalability
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent creates a universal system where a single set of audio fragments can be used across multiple product descriptions. The fragments are organized by attribute categories and can be automatically combined in different configurations to generate descriptions for numerous products, enabling both customization and scalability simultaneously

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system dynamically assembles audio descriptions by automatically selecting and stitching audio fragments based on the specific product attributes. This dynamic assembly process allows the same foundational audio library to generate customized descriptions for thousands of different products without requiring manual recreation of each description

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS8112279B2Automatic creation of audio files
Publication Date: 2012.02.07 DEALER DOT COM
  • US8112279B2 patent drawing
  • US8112279B2 patent drawing
  • US8112279B2 patent drawing

AI summary

A method of building an audio description of a particular product of a class of products includes providing a plurality of human voice recordings, wherein each of the human voice recordings includes audio corresponding to an attribute value common to many of the products. The method also includes automatically obtaining attribute values of the particular product, wherein the attribute values reside electronically. The method also includes automatically applying a plurality of rules for selecting a subset of the human voice recordings that correspond to the obtained attribute values and automatically stitching the selected subset of human voice recordings together to provide a voiceover product description of the particular product. A similar method is used to build an audio description of a particular process.