Generative Neural Network Meeting Storyboard Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for summarizing and retrieving key information from long conference calls are cumbersome and time-consuming, often requiring participants to manually transcribe and review extensive audio content.
Innovation Solution
A system and method for generating a storyboard from a media stream, utilizing generative neural network models to identify segments with related content, generate segment images, and create a visual summary of meetings, including dialog and participant contributions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If manual transcription and review of audio content is used, then key information can be captured, but the process becomes cumbersome and time-consuming
Solution Approach 1:
The patent replaces manual mechanical transcription processes with automated speech-to-text conversion and AI-powered semantic analysis. The system automatically converts audio content to text, identifies key segments, extracts important information, and generates summaries without human intervention, thereby eliminating the time-consuming manual transcription workflow while preserving all key information.
Solution Approach 2:
The system enables self-service by automatically processing meeting recordings end-to-end. The automated speech recognition, segment identification, key information extraction, and summary generation occur without requiring participant involvement. The meeting content is autonomously transformed into structured, searchable formats that participants can access independently.
2Loss of information
If extensive transcription is performed, then complete record of meeting content is available, but the complexity and time required increases significantly
Solution Approach 1:
The system extracts only the essential elements from complete meeting transcripts. Instead of processing and presenting entire transcripts, the AI identifies and extracts key segments, important statements, action items, and relevant information. This extraction approach maintains information completeness for critical content while eliminating unnecessary detail that会增加 complexity.
Solution Approach 2:
The patent segments meeting content into meaningful units such as individual speakers' contributions, topic-based sections, and chronologically ordered events. This segmentation allows the system to process and analyze manageable portions independently, reducing overall system complexity while maintaining complete coverage of meeting content through structured organization.
3Loss of information
If manual recording of key points is assigned to a participant, then meeting content can be memorialized, but the recorder cannot be an active participant
Solution Approach 1:
The system performs self-service by automatically documenting all meeting content without assigning this task to any participant. The automated speech recognition and analysis systems handle transcription, segmentation, and key point extraction independently, freeing all participants to remain fully engaged in the meeting discussion while ensuring complete documentation is generated afterward.
Solution Approach 2:
The patent replaces the mechanical role of human note-taking with automated computational systems. Instead of a participant manually recording key points, AI algorithms automatically capture, process, and organize meeting content, eliminating the conflict between documentation responsibilities and active participation while maintaining comprehensive meeting records.
Data Source
AI summary
A method for generating storyboards is described. An extraction prompt is provided to a first generative neural network model. The extraction prompt is a text-based prompt that instructs the first generative neural network model how to identify timestamps of segments having related content within transcripts according to dialog within the transcripts. A transcript of a meeting is provided as an input to the first generative neural network model. Segment timestamps for identified segments within the meeting are received from the first generative neural network model based on the extraction prompt and the transcript. Segment images for the identified segments are generated using a second generative neural network model, wherein each of the segment images represents segment content within a corresponding identified segment.


