Voice Imitation Cloud Conversion Reduces Storage
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current instant messaging applications have limited voice material options, leading to reduced enjoyment and inefficient use of storage space, as they require downloading multiple voice packages and can only produce one voice variation from recorded voice information.
Innovation Solution
A voice imitation method and apparatus that obtain training voices from multiple users, determine a conversion rule to transform a source user's voice into various imitation voices, and convert voice information using feature parameters like frequencies and formants, allowing for multiple voice variations from a single source.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple voice material packages are downloaded to increase voice variety, then the variety of available voices is improved, but the storage space occupation increases
Solution Approach 1:
The patent implements a cloud server that stores multiple voice material packages and provides them to multiple terminal devices. The cloud server acts as a universal resource that serves multiple users, allowing each user to access various voice materials without storing them locally. This enables one voice material package to serve multiple functions and multiple users simultaneously.
Solution Approach 2:
The patent uses cloud-based voice material packages that can be copied and accessed by multiple terminal devices simultaneously. Instead of each device storing its own copy, the voice materials are replicated across the cloud infrastructure and delivered on-demand to multiple users, reducing redundant storage while maintaining accessibility.
2Volume of stationary object
If only one voice material package is provided to reduce storage requirements, then the storage space occupation is reduced, but the voice variety decreases
Solution Approach 1:
The patent introduces a cloud server as an intermediary between voice material packages and terminal devices. The cloud server mediates the delivery of voice materials, allowing terminal devices to access a wide variety of voices without storing them locally. This intermediary enables users to have access to diverse voice materials while maintaining minimal local storage requirements.
Solution Approach 2:
The patent moves voice material storage from the local device dimension to the cloud dimension. Instead of expanding storage capacity on terminal devices, the system expands access to voice materials by adding a cloud-based dimension, where materials are stored centrally and distributed on-demand to multiple users.
3Speed
If voice material packages are stored locally on mobile devices, then the access speed is improved, but the storage space occupation increases
Solution Approach 1:
The patent implements a mechanism where voice material packages are pre-loaded or pre-fetched into a local buffer or cache on the terminal device before actual use. This preliminary action allows fast access during voice conversion operations while the bulk of the voice materials remain stored in the cloud, minimizing permanent local storage requirements.
4Device complexity
If a single source voice is converted to only one imitation voice to simplify the conversion process, then the conversion complexity is reduced, but the utilization rate of the voice material decreases
Solution Approach 1:
The patent enables a single voice material package to be used for generating multiple different imitation voices by applying different conversion rules or parameters. The same source voice material can be universally applied to create various voice styles, characters, or personas, increasing its utilization rate while maintaining manageable conversion complexity through standardized processing pipelines.
Data Source
AI summary
Method, apparatus, and storage medium for voice imitation are provided. The voice imitation method, includes: separately obtaining a training voice of a source user and training voices of a plurality of imitation users including a target user; determining, according to the training voice of the source user and a training voice of the target user, a conversion rule for converting the training voice of the source user into the training voice of the target user; collecting voice information of the source user; and converting the voice information of the source user into an imitation voice of the target user according to the conversion rule.


