Adaptive Data Aggregation via Machine Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data aggregation services rely on APIs or pre-programmed scripts to access data, which are unreliable and costly to maintain, especially when websites change their layouts or formats, and lack coverage in industries like utility, healthcare, and insurance due to high development and maintenance costs.

Innovation Solution

The use of natural language processing (NLP) and machine learning (ML) techniques to automate data aggregation from websites without pre-programmed scripts, allowing for adaptive operation across multiple websites with different schemas and layouts, reducing maintenance costs and improving reliability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional data aggregation services use APIs or pre-programmed scripts to access data, then they can obtain structured data access, but they become unreliable and costly to maintain when websites change their layouts or formats

Engineering Contradiction:
Improvedata aggregation reliabilityVSAvoidmaintenance complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system enables self-service by allowing the data aggregation system to automatically adapt to website changes without requiring manual intervention. The machine learning models continuously learn from website structures and automatically adjust to layout changes, eliminating the need for developers to maintain and update scripts when websites change.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system changes from using fixed, pre-programmed parameters (hardcoded scripts) to dynamic, learned parameters (machine learning models that adapt). The models learn website structures and navigate based on learned patterns rather than fixed parameters, allowing automatic adaptation when website layouts change.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If data aggregation services develop custom scripts for each website, then they can access multiple data sources, but the development and maintenance costs increase significantly

Engineering Contradiction:
Improvewebsite coverageVSAvoiddevelopment cost
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

The system implements universality by creating a single, general-purpose data aggregation platform that can access multiple different websites without requiring custom development for each site. The machine learning models learn the structure of each website automatically and can navigate any website that follows similar patterns, eliminating the need for separate scripts for each data source.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system uses copying by training machine learning models on example website structures and then applying these learned patterns to navigate and extract data from similar websites. Instead of creating unique scripts for each website, the system copies and adapts navigation patterns from learned examples to new websites.

Inventive Principle:
Principle #26Copying

3Quantity of substance

If conventional systems use pre-programmed scripts for data aggregation, then they can access data from known sources, but they lack coverage in industries like utility, healthcare, and insurance

Engineering Contradiction:
Improvedata source coverageVSAvoidscript development effort
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The system transitions from static, pre-programmed scripts to dynamic, adaptive machine learning models. The models continuously learn from website structures and can automatically adapt to new website formats in industries like utility, healthcare, and insurance without requiring upfront script development for each industry.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11526579B1System and methods for performing automatic data aggregation
Publication Date: 2022.12.13 SOPHTRON INC
  • US11526579B1 patent drawing
  • US11526579B1 patent drawing
  • US11526579B1 patent drawing

AI summary

Systems, apparatuses, and methods for automated data aggregation, automated webpage navigation, or automatically performing a task by entering data into multiple webpages. In some embodiments, this is achieved by use of techniques such as natural language processing (NLP) and machine learning to enable the automation of data aggregation and other tasks involving websites without the use of pre-programmed scripts.