Conversation transcription device, conversation transcription method, conversation transcription system, and conversation transcription program

JP2026121182APending Publication Date: 2026-07-23PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
Filing Date
2025-01-10
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Conventional audio recognition technologies struggle to accurately transcribe audio data containing simultaneous utterances from multiple speakers, resulting in low readability due to unclear correspondence between speakers' speech contents.

Method used

A conversation transcription device and method that separates audio data for each speaker, performs speech recognition, and generates a transcript that visualizes the response relationships between speakers' utterances using timestamp and contextual information.

Benefits of technology

Generates a transcript that clearly visualizes the response relationships between multiple speakers' utterances, improving readability and clarity in transcribed conversations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026121182000001_ABST
    Figure 2026121182000001_ABST
Patent Text Reader

Abstract

This invention provides a transcription device, method, and program that generate a transcription that visualizes the response relationships between the utterances of multiple users. [Solution] In a conversation content analysis system that performs transcription of conversation content, the processing device includes: a sound source separation unit that receives audio data in which the utterances of multiple speakers have been recorded, separates the audio data for each speaker and generates multiple separated data; a speech recognition unit that performs speech recognition on each of the multiple separated data and identifies the response relationships of the speech-recognized utterances of the multiple speakers; and an output result generation unit that generates and outputs a transcription that visualizes the utterances of the multiple speakers based on the identified response relationships of the utterances.
Need to check novelty before this filing date? Find Prior Art