Location-Based Audio Story Generation with LLM Voice Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio tours are limited by pre-recorded messages and predetermined stops, making updates difficult and costly, and translation into new languages is time-consuming and expensive.
Innovation Solution
A computer-implemented method that dynamically generates story content based on user location and identified points of interest, using a large language model and audio synthesis engine to create personalized audio stories, allowing for on-demand updates and multilingual support.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If traditional pre-recorded audio tours are used, then content quality and consistency are maintained, but updating content becomes difficult and costly
Solution Approach 1:
The system uses text-to-speech synthesis to generate audio content from text scripts, creating synthetic voice copies that replace the need for actual voice actors. This allows content to be updated by simply modifying the text script rather than re-recording audio, dramatically reducing update complexity and cost while maintaining audio quality consistency
Solution Approach 2:
The patent replaces the mechanical process of human voice recording and manual audio production with an automated text-to-speech system. This substitution eliminates the need for coordinating with voice actors, scheduling recording sessions, and manual audio editing, thereby simplifying the content update process while maintaining professional audio quality
2Adaptability or versatility
If traditional audio tours with voice actors are used, then natural and engaging audio content is produced, but translation into new languages becomes time-consuming and expensive
Solution Approach 1:
The text-to-speech system serves multiple functions simultaneously: it generates audio content in the original language and can generate the same content in multiple different languages by simply changing the language parameter in the text script. This universal capability allows the same content to be adapted to numerous languages without requiring separate voice actors for each language, thereby dramatically improving language adaptability while reducing translation time
Solution Approach 2:
The system changes the language parameter of the text-to-speech synthesis engine to generate audio in different languages. By modifying this single parameter rather than rewriting entire scripts or coordinating with different voice actors, the system achieves rapid multilingual content generation, reducing translation time from weeks to minutes while maintaining natural-sounding audio quality
3Ease of operation
If customized audio stories are created for specific locations, then user engagement and personalization are improved, but content creation barriers increase
Solution Approach 1:
The system automatically generates personalized audio content by taking location data from the user's device, identifying relevant points of interest through API calls, and synthesizing customized audio stories without requiring manual content creation. This self-service approach allows any user to receive personalized audio guides for their location, dramatically improving ease of operation while eliminating content creation barriers through full automation
Solution Approach 2:
The system performs preliminary actions by pre-configuring the text-to-speech engine with templates and by establishing API connections to location services before user requests. When a user provides location data, the system can immediately generate personalized content without requiring real-time content creation, thereby improving user personalization while maintaining ease of content generation through pre-established infrastructure
Data Source
AI summary
A system for creating highly personalized stories can receive location information associated with a user, which can be used to help identify one or more points of interest (POIs). One of these POIs can be manually or automatically selected. A user can select or the system can auto-select a host, which can represent a personality (e.g., a true crime podcaster, a documentarian, a comedian, an art historian, and the like). The POI and host information can be used to generate a custom prompt that can be fed into a large language model (LLM) generative AI to generate an output used to create a story transcript. The story transcript and host information can then be fed into an audio synthesis engine to generate synthesized audio used to create story audio content. The story audio content and optionally the story transcript can then be presented to the user.


