Descriptive Data Pipeline Platform for Automated Quality Enforcement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data engineering approaches require significant effort and specialized skills to develop customized pipelines, making them inefficient and non-scalable for handling large volumes of user requests, and often lack centralized data governance.
Innovation Solution
A computer-implemented method using a platform-as-a-service solution that allows users to define data pipelines through a descriptive language, enabling self-service creation of data pipelines, automating quality controls, and implementing data governance within a centralized platform, utilizing microservices for simplified maintenance and execution across multiple computing instances.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If customized data pipelines are developed manually for each user request, then data transformation accuracy is improved, but development time and effort increase significantly
Solution Approach 1:
The patent uses template-based pipeline configurations that can be copied and reused across multiple user requests. Pre-defined pipeline templates capture proven data transformation patterns, allowing rapid deployment of accurate pipelines without manual development for each request.
Solution Approach 2:
The system allows customization of pipeline parameters (such as data sources, transformation rules, and output destinations) while maintaining the core pipeline structure. This enables adaptation to specific user needs through parameter modification rather than complete pipeline redesign.
2Reliability
If specialized data engineering skills are used for pipeline development, then pipeline quality is improved, but accessibility to broader users deteriorates
Solution Approach 1:
The patent implements self-service capabilities where users can independently configure and deploy data pipelines using standardized templates and interfaces. This eliminates the need for specialized data engineering expertise while maintaining pipeline quality through enforced best practices embedded in the template system.
Solution Approach 2:
The platform provides a universal interface and template system that serves both technical and non-technical users. The same standardized templates that ensure quality for expert users also enable self-service for business users without specialized skills.
3Manufacturing precision
If centralized data governance is implemented, then data quality control is improved, but system complexity increases
Solution Approach 1:
The patent integrates data governance controls directly into the pipeline template definitions and execution framework. Quality checks, validation rules, and governance policies are merged with the pipeline configuration itself, eliminating the need for separate governance systems while maintaining centralized control.
4Adaptability or versatility
If custom pipeline development is performed for each request, then adaptability to specific needs is improved, but scalability to handle large volumes of requests deteriorates
Solution Approach 1:
The patent segments pipeline configurations into reusable template components that can be independently selected and combined. This modular approach allows rapid assembly of customized pipelines from pre-validate d building blocks, significantly increasing the rate at which adaptive pipelines can be deployed.
Data Source
AI summary
A computer-implemented method executed using a first networked computer and comprising receiving a digitally stored workflow pattern that specifies at least an input data source, a data transformation process, an output data destination, a data quality assertion and a data quality source; the workflow pattern comprising a structured plurality of name declarations and value specifications that are human readable and machine readable; the data transformation process specified in the workflow pattern including one or more references to processing logic, a processing logic source outside the workflow pattern at which the processing logic is stored, and one or more available process engines that are capable of processing the processing logic; machine parsing the workflow pattern and dividing the workflow pattern into a plurality of execution units, each execution unit being associated with a particular process engine among the one or more available process engines; accessing the input data source specified in the workflow pattern and loading at least a portion of data from the input data source into main memory; accessing the processing logic source at a second networked computer and loading a copy of the processing logic specified in the workflow pattern from the second networked computer; for each of the execution units, selecting a particular process engine among the plurality of available process engines, calling the particular process engine, programmatically providing access to the portion of data and the copy of the processing logic, and receiving output data that has been created by the particular process engine after transforming the portion of data; translating the data quality assertion into a data quality request and automatically forwarding the data quality request to the data quality source at a third computer, the data quality request comprising the data quality assertion, and receiving a response to the request that specifies whether the output data conforms to the data quality assertion.


