Speech Data Document Generation with Automatic Editing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional image forming apparatuses lack the ability to automatically edit speech data converted from voice to text, resulting in reduced usability of the final document output.
Innovation Solution
A method and system for generating documents using speech data, which includes setting document editing information for features like speaker identification, sentence patterns, and formatting options, allowing for automatic editing and generation of documents based on these settings, enabling detailed editing and output options such as printing, storage, or transmission.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If speech data is uniformly converted into text file through VTT, then the conversion process is simple and fast, but the usability of the final document deteriorates due to lack of document editing
Solution Approach 1:
The system performs preliminary document editing actions automatically during the text generation process. It pre-processes the converted text by identifying speaker changes, extracting emphasized parts, and applying appropriate formatting before the document is finalized, thus eliminating the need for manual editing while maintaining high usability.
Solution Approach 2:
The document generation system performs self-editing by automatically detecting speech features and applying corresponding document formatting. The system identifies speaker changes, extracts key information, and formats the document autonomously without requiring user intervention, thereby maintaining productivity while improving usability.
2Ease of operation
If automatic document editing features are added, then document usability improves, but device complexity increases
Solution Approach 1:
The system integrates multiple document editing functions into a single unified platform. It combines speaker identification, emphasized part extraction, document formatting, and text generation capabilities in one system, allowing the device to perform various document processing tasks without requiring separate complex systems for each function.
Solution Approach 2:
The system manages complexity by dynamically adjusting processing parameters based on input characteristics. It adapts the level of automatic editing, the granularity of speaker separation, and the extent of format application based on the speech data properties, allowing flexible operation without requiring complex fixed configurations for all possible scenarios.
3Manufacturing precision
If detailed document editing options are provided, then document quality improves, but the processing time increases
Solution Approach 1:
The system applies partial automatic editing actions based on the detected speech characteristics. Instead of applying all possible editing options uniformly, it selectively applies only the relevant editing actions identified through speech feature analysis, such as speaker separation when multiple speakers are detected or emphasized part highlighting when specific speech patterns are identified, thus maintaining quality while minimizing processing time.
Data Source
AI summary
A document generation method and system using speech data, and an image forming apparatus including the document generation system. The method includes setting document editing information including at least one of document form information and sentence pattern information for editing a document when the speech data is generated as the document; converting the speech data into text; and generating the text as the document based on the document editing information.


