Correlating Written Feedback to Telemetry via NLP Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Manufacturers and developers face challenges in efficiently identifying and addressing performance and reliability issues in computer products due to the vast amount of written feedback and telemetry data, with manual parsing being time-consuming and many telemetry events being irrelevant.
Innovation Solution
A system that applies natural language processing to correlate semantically similar instances of written feedback with relevant telemetry data, extracting timestamps and device identifications to retrieve relevant telemetry event logs, and analyzing these to identify common events, thereby linking customer feedback context with actionable telemetry insights.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If manufacturers manually parse through written feedback to identify performance and reliability issues, then they can understand user experiences and context, but the process becomes time-consuming and inefficient when dealing with millions of feedback instances
Solution Approach 1:
The patent replaces manual mechanical parsing of feedback text with automated natural language processing systems. The NLP model automatically processes, clusters, and analyzes feedback instances, substituting human manual analysis with computational mechanisms that can handle millions of feedback records efficiently while preserving contextual understanding.
Solution Approach 2:
The system enables self-service analysis where the NLP model autonomously performs feedback processing, clustering, and correlation with telemetry data without requiring manual intervention. The system serves itself by automatically identifying patterns, grouping similar feedback, and linking to relevant telemetry events, eliminating the need for continuous human oversight in the analysis process.
2Reliability
If manufacturers collect and analyze telemetry data to identify performance and reliability issues, then they can diagnose technical problems, but many telemetry events are irrelevant and the analysis becomes inefficient
Solution Approach 1:
The patent extracts and isolates relevant telemetry events from the broader dataset by clustering feedback instances and selecting only those telemetry events that are temporally and contextually relevant to the clustered feedback. This extraction process filters out irrelevant telemetry data, retaining only the portions that directly relate to identified performance and reliability issues.
Solution Approach 2:
The system segments the telemetry data analysis process into distinct phases: clustering feedback instances by semantic similarity, extracting temporal and contextual information, retrieving relevant telemetry events based on timestamps and device identifiers, and analyzing only the segmented relevant portions. This segmentation allows efficient processing by focusing computational resources on specific subsets of data rather than analyzing all telemetry events uniformly.
3Loss of information
If manufacturers correlate written feedback with telemetry data, then they can link user experiences with technical events, but the volume of data makes manual correlation impractical
Solution Approach 1:
The patent introduces temporal information (timestamps) and device identifiers as intermediary elements that bridge feedback data and telemetry data. These intermediaries serve as connecting keys that allow the system to correlate feedback instances with relevant telemetry events without requiring complex direct mapping relationships. The intermediaries simplify the correlation process by providing clear temporal and contextual anchors for matching data from different sources.
Data Source
AI summary
Disclosed herein is a system that automatically correlates related instances of written feedback associated with a computer product to relevant telemetry data so that a performance and/or reliability issue associated with the computer product can be identified and fixed in a more efficient manner. The system applies a natural language processing model to instances of written feedback that have been received for the computer product to identify a cluster of instances of written feedback which describe the same issue. For each instance of written feedback in the cluster, the system uses an identification and a timestamp to retrieve a telemetry event log. The system analyzes the telemetry event logs to identify a common telemetry event. This effectively correlates meaningful written descriptions with common telemetry behavior(s). Information regarding this correlation is provided to a computing device of a user (e.g., a developer) tasked with examining the performance and/or reliability issue.


