Stream Summarization System for Information Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face difficulty in identifying and understanding the major points from vast amounts of information on social networking sites, real-time messaging services, and blogs due to the presence of irrelevant content, making it hard to discern what is relevant without reading through entire threads or posts.

Innovation Solution

A system that analyzes streams of data from various sources, groups related content into clusters, summarizes the data, and presents it in a timeline format, allowing users to easily understand topics without needing to read through all the original content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If users read through entire threads, blogs, walls, or other postings to understand major points, then comprehension accuracy is improved, but time consumption increases

Engineering Contradiction:
Improvecomprehension accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system extracts and identifies only the relevant information from vast streams of data by analyzing content against user-defined criteria. It pulls out key entities, events, and relationships while filtering out irrelevant noise, enabling users to access only the essential information needed for comprehension without reading entire threads or posts.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system introduces an intermediary processing layer between raw data streams and user consumption. This intermediary component analyzes, clusters, and summarizes information automatically, acting as a filter that translates complex data streams into digestible summaries, thereby reducing the time users need to spend while maintaining comprehension accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If users read through all content to identify relevant information, then information completeness is improved, but productivity decreases

Engineering Contradiction:
Improveinformation completenessVSAvoidprocessing efficiency
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The system performs preliminary analysis and organization of information before it reaches the user. By pre-processing data streams to identify, cluster, and tag relevant content in advance, the system ensures that when users access the information, it is already organized and ready for quick consumption, maintaining completeness while boosting productivity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system replaces the mechanical process of manual reading and filtering with automated computational analysis. Instead of users mechanically reading through all content, the system uses algorithms to automatically identify, cluster, and present relevant information, maintaining information completeness while dramatically improving processing efficiency.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Quantity of substance

If the system processes vast amounts of data from multiple sources, then information coverage is improved, but system complexity increases

Engineering Contradiction:
Improveinformation coverageVSAvoidsystem complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The system segments the complex task of processing vast data streams into distinct manageable modules: data collection, data analysis, clustering, summarization, and presentation. Each module handles a specific aspect of processing, making the overall system more manageable despite the volume of information it processes from multiple sources.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system employs universal processing mechanisms that can handle diverse data types and sources through a common framework. The same core algorithms and processing logic can analyze different formats (text, video, audio) and sources (blogs, social media, news sites), allowing the system to maintain information coverage across multiple sources without proportionally increasing complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS8326880B2Summarizing streams of information
Publication Date: 2012.12.04 MICROSOFT TECHNOLOGY LICENSING LLC
  • US8326880B2 patent drawing
  • US8326880B2 patent drawing
  • US8326880B2 patent drawing

AI summary

Concepts and technologies are described herein for summarizing streams of information. A stream of information is obtained and analyzed. One or more entities are identified in the stream. The data in the stream is grouped into one or more clusters corresponding to the identified entities. The data in the clusters is summarized, and a timeline corresponding to the data in the cluster is determined. In some embodiments, a format can be selected for presentation of the summarized stream data. The data in the stream can be formatted in the selected format, and the summarized data can be presented in the selected format. In some embodiments, an update feature can be used to update the data in the summarized stream. The data in the stream can be updated, and the updated summarized stream can be formatted and presented.