Method, system and computer program product for generating training data for speech emotion recognition models

TWI938917BActive Publication Date: 2026-09-11NATIONAL TSING HUA UNIVERSITY
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
TW114112339
Authority / Receiving Office
TW · TW
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2026-09-11
Estimated Expiration
2045-03-30

Smart Images

  • Figure TWG2TB001910485_001
    Figure TWG2TB001910485_001
  • Figure TWG2TB001910485_002
    Figure TWG2TB001910485_002
  • Figure TWG2TB001910485_003
    Figure TWG2TB001910485_003
Patent Text Reader

Abstract

This invention proposes a method for generating training data for a Speech Emotion Recognition (SER) model, comprising receiving cue data, including dialogue scene description information, dialogue participant information, and script content or instructions for automatically generating a script; generating text dialogue data based on the cue data using a large language model, the text dialogue data including at least one sentence and an emotion tag corresponding to each sentence; and converting the text dialogue data into emotionally charged speech dialogue data using at least one speech model. Furthermore, systems and computer program products using the above method are also proposed.
Need to check novelty before this filing date? Find Prior Art

Claims

1. A method for generating training data for a speech emotion recognition (SER) model, comprising: Receive prompt data, which includes dialogue scene description information and dialogue participant information, wherein the prompt data also includes at least one of script content and instructions for automatically generating script; Based on the prompt data as input to a large language model, use the large language model to generate text dialogue data, which includes at least one statement and an emotion tag corresponding to each statement. And using at least one speech model to convert the text dialogue data into speech dialogue data with emotion, including: using a first speech model to convert the text dialogue data into speech dialogue data without emotion. And using a second speech model, based on the emotion marker in the text dialogue data, the speech dialogue data without emotion is converted into speech dialogue data with emotion.

2. The method as described in request item 1, wherein the prompt data further includes dialog content distribution characteristics.

3. The method as described in claim 1, wherein the emotion tag includes: One of the following: happiness, neutrality, sadness, frustration, anger, surprise, fear, disgust, and contempt.

4. A system for generating training data for a speech emotion recognition (SER) model, the system comprising: Memory stores at least one instruction; A processor, coupled to the memory, wherein when the processor executes the at least one instruction, it is configured to: receive cue data, the cue data including dialogue scenario description information and dialogue participant information, wherein the cue data also includes at least one of script content and instructions for automatically generating a script; and, based on the cue data as input to a large language model, generate text dialogue data using the large language model, the text dialogue data including at least one statement and an emotion tag corresponding to each statement. And using at least one speech model to convert the text dialogue data into speech dialogue data with emotion, including: using a first speech model to convert the text dialogue data into speech dialogue data without emotion. And by using a second speech model, based on the emotion marker in the text dialogue data, the speech dialogue data without emotion is converted into speech dialogue data with emotion.

5. The system as described in claim 4, wherein the prompt information further includes dialog content distribution characteristics.

6. The system as described in claim 4, wherein the emotion tag includes: One of the following: happiness, neutrality, sadness, frustration, anger, surprise, fear, disgust, and contempt.

7. A computer program product comprising at least one instruction, which, when executed by a processor of an electronic device, enables the electronic device to perform the method as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Information generation method and device, electronic equipment and storage medium

    CN113051380A

  • Dialogue pre-labeling method and system, computer equipment and storage medium

    CN116860921A

  • Training method and device for voice emotion interaction model and electronic equipment

    CN118711572A

  • System, method, and article of manufacture for a voice recognition system for navigating on the internet utilizing audible information

    TW491991B

  • System and method for collaborative shopping, business and entertainment

    US20210224765A1