Automated Data Science Service Pipeline for Predictive Modeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for providing data science and artificial intelligence as-a-service struggle to efficiently process and analyze large volumes of diverse data, requiring professional data scientists and complex infrastructure.

Innovation Solution

An automated method that cleans and enriches raw data using algorithms to make fields consistent, removes uncontributory fields, and adds new fields for predictive model building, ultimately rendering predictive models in deliverable markup language documents.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If automated data cleaning and enrichment algorithms are implemented, then predictive model accuracy is improved, but processing time and computational complexity increase

Engineering Contradiction:
Improvepredictive model accuracyVSAvoiddata processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by performing data cleaning and enrichment operations before the main predictive modeling process. The system pre-processes raw data by removing duplicates, handling missing values, standardizing formats, and enriching with external data sources in advance, so that the actual model training receives already-prepared high-quality data, improving accuracy while managing processing time through staged execution.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the data processing workflow into distinct modular stages: data ingestion, cleaning operations, enrichment processes, validation, and model training. Each stage can be independently optimized and executed, allowing parallel processing where possible and enabling the system to manage computational complexity by breaking down the overall time-consuming process into manageable segments that can be scheduled efficiently.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If comprehensive data enrichment is performed to add new fields for predictive modeling, then model predictive accuracy is enhanced, but device complexity and computational resources required increase

Engineering Contradiction:
Improvepredictive accuracyVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements universality by designing a multi-functional data enrichment engine that can perform multiple types of enrichment operations through a single unified system. The engine can internally generate derived fields, externally fetch data from multiple sources, validate against schemas, and transform data formats all within one platform. This consolidates what would otherwise require multiple separate tools and reduces overall system complexity while comprehensively enriching data for predictive modeling.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces an intermediary enrichment layer between raw data ingestion and predictive model training. This intermediary component acts as a mediator that systematically adds, validates, and transforms data fields before they reach the modeling stage. The enrichment engine serves as a buffer that manages the complexity of data preparation independently, allowing the core predictive modeling system to focus on its primary function while the intermediary handles the complexity of comprehensive data enrichment.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If professional data scientists are replaced with automated algorithms, then service accessibility and ease of operation improve, but the sophistication and expertise previously provided by human specialists may be reduced

Engineering Contradiction:
Improveservice accessibilityVSAvoidexpert-level analysis quality
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent applies self-service by enabling automated predictive modeling systems to perform data preparation, feature engineering, model selection, and validation independently without requiring human data scientists. The system automatically ingests raw data, cleans and enriches it using configured algorithms, selects appropriate predictive models based on data characteristics, trains and validates models, and deploys them for inference. This makes the service accessible to users without specialized expertise while maintaining reliability through automated best practices and continuous validation.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent utilizes parameter changes by allowing the automated system to dynamically adjust processing parameters, algorithm selections, and model configurations based on the characteristics of the input data. The system can automatically detect data types, distributions, and quality metrics, then adaptively select appropriate cleaning strategies, enrichment methods, and modeling approaches. This flexibility enables the system to maintain expert-level analysis quality across diverse datasets while remaining fully automated and accessible to non-expert users.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12282831B2Method of automating data science services
Publication Date: 2025.04.22 BRIGHTERION INC
  • US12282831B2 patent drawing
  • US12282831B2 patent drawing
  • US12282831B2 patent drawing

AI summary

An automated method of predictive model development first cleans up raw supervised and unsupervised training data with a step that uses an algorithm to make every field of every record consistent, cohesive, and productive. Then the resulting flat data is given texture in a next step by a data enrichment algorithm that culls fields that do not contribute to predictive model building and that adds new fields computed from data combinations that are tested to add value to later steps that build different types of predictive models. Another late step for building smart-agents and their entity profiles uses another algorithm that benefits greatly from the cleaned and highly enriched training data. The predictive models and smart-agents and their entity profiles are then rendered as deliverable predictive model markup language documents in a final step executed by a specialized algorithm.