Knowledge Extraction from Collaborative Support Sessions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current tools are inefficient in extracting valuable knowledge from collaborative support sessions, which are lengthy, complex, and often contain background noise and multiple speakers, making it difficult to manually extract actionable information for future use.
Innovation Solution
A knowledge extraction system that uses machine learning and image processing to convert screen sharing content into text sequences with timestamps, synchronize commands and audio information, and generate a report with time-stamped entries of commands, parameters, and speech-based information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual extraction of information from troubleshooting sessions is performed, then information accuracy can be maintained, but productivity is significantly reduced due to the lengthy and complex nature of sessions
Solution Approach 1:
The patent introduces an automated information extraction system that acts as an intermediary between the troubleshooting session data and the final knowledge base. This system uses machine learning models to automatically transcribe audio, convert screen sharing to text, extract commands, and synchronize multi-modal data, thereby maintaining information accuracy while dramatically improving extraction productivity
Solution Approach 2:
The patent replaces the manual mechanical process of information extraction with an automated computational system. Machine learning models substitute for human analysts in transcribing audio, identifying commands, extracting parameters, and synthesizing knowledge, thereby eliminating the productivity bottleneck while preserving accuracy through algorithmic consistency
2Productivity
If automated tools are used to extract information from troubleshooting sessions, then productivity is improved, but measurement precision deteriorates due to background noise and multiple speakers
Solution Approach 1:
The patent segments the troubleshooting session into distinct components: audio streams are separated by speaker, screen sharing content is divided into discrete frames, and commands are extracted as individual entities. This segmentation allows the system to process each component separately with specialized algorithms, improving overall accuracy despite the presence of background noise and multiple speakers
Solution Approach 2:
The patent introduces multiple intermediary processing layers including audio transcription models, screen-to-text conversion systems, and command extraction algorithms. These intermediaries act as filters and translators that progressively refine the raw data from each modality, maintaining information accuracy while enabling automated high-volume processing
3Loss of information
If comprehensive information is captured from all session modalities, then knowledge completeness is improved, but device complexity increases due to multiple data sources requiring synchronization
Solution Approach 1:
The patent merges multiple data sources (audio transcripts, screen sharing text, identified commands, and parameters) into a unified knowledge representation. By combining these diverse modalities into a single integrated output format, the system achieves comprehensive knowledge capture while managing complexity through unified data structures and synchronization mechanisms
4Loss of information
If lengthy troubleshooting sessions are processed in detail, then knowledge completeness is improved, but loss of time increases due to the extended duration of sessions
Solution Approach 1:
The patent replaces time-consuming manual processing of lengthy sessions with automated machine learning systems that can rapidly transcribe, analyze, and extract knowledge from extended troubleshooting recordings. This substitution enables complete processing of long sessions without proportional increases in human time investment
Data Source
AI summary
At a communication server, a first computer device and a second computer device are connected to a collaborative support session configured to support audio communications, screen sharing, and control of the first computer device by the second computer device. Screen sharing video image content is converted to a text sequence with timestamps. A text log with timestamps is generated from the text sequence. Using a command-based machine learning model, a sequence of commands and associated parameters, with timestamps, are determined from the text log. Audio is analyzed to produce speech-based information with timestamps. The command sequence is time-synchronized with the speech-based information based on the timestamps of the command sequence and the timestamps of the speech-based information. A knowledge report for the collaborative support session is generated. The knowledge report includes entries each including a timestamp, commands and associated parameters, and speech-based information that are time-synchronized to the timestamp.


