A method and system for simulating personalized automated voice messages

TWI935823BActive Publication Date: 2026-08-11GAMANIA DIGITAL ENTERTAINMENT CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
TW114121010
Authority / Receiving Office
TW · TW
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2026-08-11
Estimated Expiration
2045-06-04

Smart Images

  • Figure TWG2TB001905762_001
    Figure TWG2TB001905762_001
  • Figure TWG2TB001905762_002
    Figure TWG2TB001905762_002
  • Figure TWG2TB001905762_003
    Figure TWG2TB001905762_003
Patent Text Reader

Abstract

An automated voice messaging system simulating personalized characteristics includes a server that performs natural language processing via a processor, analyzing multiple voice recordings of a specific person and assigning at least one persona tag. It utilizes automatic speech recognition and corpus annotation technologies to add contextual tags to the voice data and convert it into voice and text training data. AI is used to train and learn from the voice and text training data, generating multiple contextual voices corresponding to the specific person and text models incorporating their personality traits. Based on the contextual text models, the system synthesizes the voice and text, creating a corresponding contextual voice model, which is then stored as a unique voice model for that specific person. Based on the calendar, contextual voice emails matching the subject content are automatically sent to the member's mailing address.
Need to check novelty before this filing date? Find Prior Art

Claims

1. A system for simulating personalized automated voice messages, the system comprising: a server, the server including a processor performing natural language processing to analyze a plurality of voice data related to a specific person, and setting at least one persona tag for the specific person, the specific person including idols, virtual characters, anime characters, public figures, anchors, moderators, and supervisors; performing automatic speech recognition and corpus annotation technology to annotate a plurality of contextual tags in the voice data, the contextual tags including birthday context, festival context, anniversary context, event context, and conditional context, and storing the corresponding contextual tags as a plurality of voice training data, and further processing the voice training data... The training data is converted into multiple text training data; an AI training module learns from these voice training data, text training data, and the persona tag to generate multiple contextual voice models and multiple contextual text models corresponding to the specific person; a voice-text synthesis module synthesizes the contextual voice models based on the contextual text models and stores them as multiple specific person voice models corresponding to the contextual tags; and a mail sending module sends the specific person voice model corresponding to the contextual tag for a topic to a member's mailing address according to a calendar. The topic includes birthday themes, festival themes, anniversary themes, event themes, and conditional themes; a dynamic information collection module collects the specific person's posts on multiple platforms, analyzes the persona tags, and builds a user profile.

2. The system as described in claim 1, wherein the voice data is collected from a plurality of network platforms in connection with the voiceprint of a particular person or a particular person.

3. The system as described in claim 1, wherein the processor performs a voiceprint separation technique on the voice data to separate it into a background sound file and a voice file, and performs a data enhancement technique on the voice file.

4. The system as described in Request 1, wherein the speech-text synthesis module performs voiceprint optimization on the voice models of these specific individuals.

5. As described in Request 1, the AI ​​training module also includes a text optimization and filtering mechanism that uses a predefined blacklist and banned words to filter out sensitive or inappropriate text content that does not meet the requirements.

6. A method for simulating personalized automated voice messages, the method comprising: a processor of a server performing natural language processing to analyze a plurality of voice data related to a specific person, and setting at least one persona tag for the specific person, the specific person including idols, virtual characters, anime characters, public figures, anchors, moderators, and supervisors; the processor performing automatic speech recognition and corpus annotation technology to annotate a plurality of contextual tags in the voice data, the contextual tags including birthday context, festival context, anniversary context, event context, and conditional context, and storing the corresponding contextual tags as a plurality of voice training data; further, the processor converting the voice training data into a plurality of text training data; an AI training module converting the voice training data and the text training data into a plurality of text training data. This training material and the persona tag are used to learn and generate multiple contextual speech models and multiple contextual text models corresponding to the specific person; a speech-text synthesis module synthesizes the contextual speech models based on the contextual text models and stores them as multiple specific person speech models corresponding to the contextual tags; and a mail sending module sends the specific person speech model corresponding to the contextual tag to a member's mailing address based on a calendar, the contextual tag of the topic content, including birthday topics, festival topics, anniversary topics, event topics, and conditional topics; and a dynamic information collection collects the posting content of the specific person on multiple platforms, analyzes the persona tag, and builds a user profile.

7. The method described in claim 6, wherein the processor performs a voiceprint separation technique on the voice data to separate it into a background sound file and a voice file, and performs a data enhancement technique on the voice file.

8. The method described in claim 6, wherein the speech-text synthesis module performs voiceprint optimization on the voice models of these specific individuals.

9. As described in claim 6, the voice model of the specific person sent by the email sending module is further selected to match the persona tag of a member.

10. The method described in Request 6 further includes a text optimization and filtering mechanism that uses a predefined blacklist and banned words to filter out sensitive or inappropriate text content that does not meet the requirements.

Citation Information

Patent Citations

  • To-do list automatic prompt realizing method and device

    CN105608555A

  • In-vehicle voice privacy protection method, system and device and storage medium

    CN118197295A

  • Speech synthesis method, speech synthesis model training method, electronic device and computer program product

    CN119763547A

  • Speech synthesis dubbing system

    TW202230330A

  • Dual use of acoustic model in speech-to-text framework

    US20220028395A1