Automated Audio Description Generation via Fragment Stitching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems for creating audio descriptions of products or processes require human input or scripted content, limiting their efficiency and scalability in generating natural-sounding voiceovers for various products or processes.
Innovation Solution
A method that uses a plurality of human voice recordings, automatically selects and stitches together audio fragments based on electronically obtained attribute values, eliminating the need for human intervention and script input, to create natural-sounding voiceover descriptions for products or processes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If human voice recordings are manually selected and stitched, then the audio description quality is high, but the time consumption and labor requirements increase significantly
Solution Approach 1:
The patent segments the voiceover creation process into discrete audio fragments corresponding to different product attributes. Each fragment is pre-recorded with proper prosody and can be independently selected and stitched together, transforming a manual process into an automated assembly process that reduces time consumption while maintaining quality
Solution Approach 2:
The system automatically selects and stitches audio fragments based on electronically obtained attribute values without requiring manual human intervention. The automated selection process uses rules to match audio fragments with product attributes, enabling the system to serve itself and eliminate time-consuming manual labor
2Productivity
If automated systems are used to generate audio descriptions, then efficiency and scalability improve, but the natural-sounding quality and context recognition may deteriorate
Solution Approach 1:
The patent performs preliminary action by having humans record audio fragments with proper prosody and context in advance. These pre-recorded fragments are then automatically assembled, combining the natural quality of human recording with the efficiency of automated processing. The preliminary recording phase captures authentic human speech patterns that automated synthesis cannot replicate
Solution Approach 2:
The system copies and reuses authentic human voice recordings rather than attempting to synthesize speech. By selecting and stitching existing human-recorded audio fragments, the system preserves the natural quality and context recognition of human speech while achieving automated production efficiency
3Adaptability or versatility
If human input and script input are required, then the audio description can be customized, but the scalability and ability to generate thousands of descriptions is limited
Solution Approach 1:
The patent creates a universal system where a single set of audio fragments can be used across multiple product descriptions. The fragments are organized by attribute categories and can be automatically combined in different configurations to generate descriptions for numerous products, enabling both customization and scalability simultaneously
Solution Approach 2:
The system dynamically assembles audio descriptions by automatically selecting and stitching audio fragments based on the specific product attributes. This dynamic assembly process allows the same foundational audio library to generate customized descriptions for thousands of different products without requiring manual recreation of each description
Data Source
AI summary
A method of building an audio description of a particular product of a class of products includes providing a plurality of human voice recordings, wherein each of the human voice recordings includes audio corresponding to an attribute value common to many of the products. The method also includes automatically obtaining attribute values of the particular product, wherein the attribute values reside electronically. The method also includes automatically applying a plurality of rules for selecting a subset of the human voice recordings that correspond to the obtained attribute values and automatically stitching the selected subset of human voice recordings together to provide a voiceover product description of the particular product. A similar method is used to build an audio description of a particular process.


