Schema Views for Data Transformation Jobs
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current tools are inadequate for managing and manipulating large-scale XML schemas, which often consist of thousands of nodes arranged in complex hierarchies, necessitating improved methods for data transformation and parsing across different formats.
Innovation Solution
A computer program product and method that generates views of subsets of nodes from a schema, allowing for the creation of data transformation jobs to convert input data from one format to another, utilizing a client GUI, application server, engine server, and repository to process and transform data through a sequence of steps, including parsing and composing operations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If complete XML schemas with thousands of nodes are used to define data models, then data transformation accuracy and completeness are improved, but system complexity and difficulty of manipulation increase significantly
Solution Approach 1:
The patent divides the complete XML schema into multiple manageable portions by creating views that select and present only specific nodes and elements relevant to particular data transformation tasks. This segmentation allows users to work with smaller, more manageable schema portions rather than the entire complex schema at once, while still maintaining transformation accuracy for the targeted data elements.
Solution Approach 2:
The patent extracts relevant portions of the schema by using view definitions that selectively retrieve only the necessary nodes, elements, and relationships from the complete schema. This extraction mechanism enables data transformation operations to focus on specific data models without needing to process the entire complex schema, thereby reducing operational complexity while preserving transformation precision.
2Adaptability or versatility
If large-scale schemas with hundreds of XSD files are used to define complex data models, then data model comprehensiveness is improved, but ease of operation and data manipulation decrease
Solution Approach 1:
The patent creates a universal view management system that can handle multiple data models and transformation tasks through a single interface. Views can be defined once and reused across different data transformation jobs, allowing the same view infrastructure to serve multiple purposes including parsing, composing, and transforming data between various formats and schemas.
Solution Approach 2:
The patent uses view definitions as templates or copies that can be instantiated multiple times for different data transformation tasks. Instead of working with the original complex schema directly, users work with view copies that are pre-configured with specific node selections, making each operation simpler while maintaining access to the comprehensive data model through the view hierarchy.
3Adaptability or versatility
If manual configuration of data transformation jobs is used, then transformation customization flexibility is improved, but productivity and time required for transformation increase
Solution Approach 1:
The patent performs preliminary configuration by allowing users to define views and transformation templates in advance before actual data transformation tasks. These pre-defined views with their associated node selections and transformation rules can be quickly instantiated and applied to different data files without requiring manual reconfiguration, thereby maintaining customization flexibility while significantly improving transformation productivity.
Data Source
AI summary
Provided is a method for processing input data in a storage system and in communication with a repository. Views are generated that comprise a tree of nodes selected from a subset of nodes in a hierarchical representation of a schema. The views are saved to the repository. At least one of the views are used to create a job comprising a sequence of data transformation steps to transform the input data described by input schemas to the output data described by output schemas.


