Information processing device, information processing method, and program

By employing multitask learning to estimate intonation, accent phrases, and nuclei in Japanese text, the method improves accent estimation accuracy and speech synthesis quality by reflecting the hierarchical linguistic structure.

JP7864585B2Active Publication Date: 2026-05-25LINE WORKS CORP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
LINE WORKS CORP
Filing Date
2022-07-28
Publication Date
2026-05-25

AI Technical Summary

Technical Problem

Existing accent estimation methods for Japanese text fail to adequately reflect the hierarchical language structure of accent phrases and nuclei, and do not account for the relationship with intonation phrases, leading to inaccuracies in accent estimation.

Method used

An information processing device and method that uses multitask learning to train a model for simultaneously estimating the delimiter positions of intonation phrases, accent phrases, and accent nuclei in Japanese text, employing a three-layer bidirectional long-short-term memory network model, conditional random field model, and autoregressive model to improve accuracy.

Benefits of technology

The proposed method enhances the accuracy of accent estimation in Japanese text by reflecting the hierarchical linguistic structure, resulting in improved speech synthesis quality and naturalness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007864585000014
    Figure 0007864585000014
  • Figure 0007864585000015
    Figure 0007864585000015
  • Figure 0007864585000016
    Figure 0007864585000016
Patent Text Reader

Abstract

To provide an information processing device, etc. that can improve accuracy of accent estimation for Japanese texts.SOLUTION: An information processing device includes: a data acquisition unit for acquiring input data including text data to be synthesized into speech; and an estimation unit that outputs from the input data a delimiter position of an intonation phrase in the text data, a delimiter position of an accent phrase in the text data, and results of estimation of a position of the accent nucleus in the text data, using a learned model learned based on multi-task learning including a first task for estimating from the input data a delimiter position of an intonation phrase in the text data, a second task for estimating from the input data a delimiter position of an accent phrase in the text data, and a third task of estimating from the input data a delimiter position of an accent nucleus in the text data.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art