Machine Learning Code Mapping for Schema Incompatibility

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data processing systems face challenges in handling and classifying diverse and inconsistent codes from different organizations and software platforms, making it difficult to import and analyze data across disparate sources due to incompatible data schemas.

Innovation Solution

The implementation of machine learning techniques to classify codes by determining similarities between recognized and unrecognized codes, allowing for accurate mapping and classification regardless of the source schema, using a network environment with a verifier computing environment, code classifier service, and machine learning model.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional data processing systems are used to handle codes from different organizations, then data can be processed using existing schemas, but the system cannot correctly handle new or unrecognized codes and fails when source schemas are incompatible

Engineering Contradiction:
Improveability to handle diverse code schemasVSAvoidcorrect classification accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent introduces an intermediary system consisting of a code classifier service and machine learning model that acts as a mediator between diverse source schemas and the target data warehouse schema. This intermediary automatically classifies and maps codes from different organizations into a unified schema, enabling the system to handle diverse code schemas while maintaining reliable classification through trained machine learning algorithms.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If manual classification methods are used for code mapping, then accurate classification can be achieved, but the process becomes time-consuming and cannot scale to handle large volumes of data from multiple sources

Engineering Contradiction:
Improvecode classification accuracyVSAvoiddata processing throughput
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent replaces manual mechanical classification processes with an automated machine learning-based code classifier service. The machine learning model automatically performs code classification and mapping tasks that would otherwise require manual intervention, thereby maintaining high classification accuracy while dramatically increasing data processing throughput and scalability.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Stability of the object's composition

If a unified schema is imposed on all data sources, then data consistency can be achieved, but the system loses the ability to accommodate organization-specific code structures and granularities

Engineering Contradiction:
Improvedata schema consistencyVSAvoidaccommodation of source-specific schemas
Core Design Contradiction:
Stability of the object's compositionVSAdaptability or versatility

Solution Approach 1:

The patent segments the code mapping process into distinct components: a configurable mapping layer that handles organization-specific code structures, a machine learning classification layer that adapts to different code granularities, and a unified output schema. This segmentation allows the system to maintain data consistency in the target schema while preserving the ability to accommodate various source-specific code structures through configurable mapping rules and adaptive classification.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20240160911A1Machine learning based code mapping
Publication Date: 2024.05.16 ADP INC
  • US20240160911A1 patent drawing
  • US20240160911A1 patent drawing
  • US20240160911A1 patent drawing

AI summary

Disclosed are various embodiments for machine learning based code mapping. A computing device can obtain a set of code identifiers. Then, the computing device can provide a first one of the set of code identifiers to a machine learning model and receive a potential classification from the machine learning model in response. Then, the computing device can determine that the confidence score for the potential classification meets or exceeds a predefined threshold value. In response, the computing device can then create a first mapping pair that links the first one of the set of code identifiers as being associated with the bucket.