Multimodal Meeting Summarization System for Information Loss Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing meeting minute generation systems that rely solely on spoken or written text summaries often result in information loss, as important points may not be fully captured or conveyed, leading to incomplete meeting records.
Innovation Solution
A system that integrates both voice and image data by performing speech recognition on voice recordings and character recognition on captured images, generating summary information that combines and summarizes both sources of data to create comprehensive meeting minutes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If only spoken text is summarized to generate meeting minutes, then the summarization process is simple and fast, but important information is lost when users do not speak what they wrote or write what they do not speak
Solution Approach 1:
The patent combines both spoken text summarization and written text summarization into a single meeting minutes generation system. The integration unit merges the output from the speech recognition unit and character recognition unit, ensuring that both auditory and visual information sources are captured and synthesized into comprehensive meeting minutes, thereby reducing information loss without requiring completely separate systems
Solution Approach 2:
The summarization apparatus is designed to perform multiple functions: it can process both spoken language (via speech recognition) and written text (via character recognition on images of writing media). This multi-functional capability allows a single system to handle diverse information sources present in meeting environments, making the system more versatile and reducing the need for multiple specialized tools
2Loss of information
If both spoken and written text are processed to generate meeting minutes, then information completeness is improved, but the processing time and computational resources increase
Solution Approach 1:
The system performs speech recognition and character recognition operations in parallel during the meeting, rather than sequentially after the meeting ends. By preliminarily processing both spoken and written content concurrently as they occur, the system reduces the total processing time required while ensuring both information sources are captured for comprehensive meeting minutes generation
Data Source
AI summary
A system that outputs information generated by summarizing contents of voices and images as texts. A CPU of the system performs, according to a program stored in a memory, recording voice data and captured image data, generating first text information by performing speech recognition on the acquired voice data, generating second text information by performing character recognition on the acquired image data, and generating summary text information smaller in the number of characters than the first text information and the second text information, based on the first text information and the second text information, according to a predetermined criterion.


