Natural Language Code Generation for Automated Data Transformation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning platforms require technical expertise for configuration and data preparation, making them inaccessible to non-experts and inefficient for automated model generation and execution.

Innovation Solution

A system that automatically generates and executes machine learning models using natural language descriptions of data manipulations, analyzing user data and tasks to create models with minimal user input, incorporating encoders and large language models for real-time model generation and execution.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional machine learning platforms are used, then model accuracy can be achieved, but technical expertise is required for configuration and data preparation

Engineering Contradiction:
Improvemodel accuracyVSAvoiduser accessibility
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent introduces natural language processing as an intermediary layer between the user and the machine learning platform. Users can describe data manipulation tasks in natural language, and the system automatically translates these descriptions into executable code and model configurations, eliminating the need for users to have technical expertise in platform configuration

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system automatically generates executable code, configures models, and prepares data without requiring user intervention in technical aspects. The platform self-adjusts parameters and executes tasks based on high-level user instructions, making the complex machine learning process self-serving and accessible to non-experts

Inventive Principle:
Principle #25Self-service

2Reliability

If detailed technical configuration is performed manually, then model performance can be optimized, but time consumption increases significantly

Engineering Contradiction:
Improvemodel performanceVSAvoidmodel generation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system pre-configures model architectures, data processing pipelines, and execution environments based on the user's natural language description before actual model training begins. This preliminary automated setup eliminates the time-consuming manual configuration phase while ensuring proper model performance through pre-validated configurations

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces manual mechanical configuration processes with automated code generation and execution systems. The system automatically translates natural language requirements into optimized model configurations and executes them, substituting the time-intensive manual tuning process with rapid automated system operations

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS12498908B2Methods and systems for automatically generating and executing computer code using a natural language description of a data manipulation to be performed on a data set
Publication Date: 2025.12.16 AKKIO INC
  • US12498908B2 patent drawing
  • US12498908B2 patent drawing
  • US12498908B2 patent drawing

AI summary

A method for automatically generating and executing computer code includes receiving, by a machine learning engine, a user-specified data set and a user-specified task. The machine learning engine analyzes at least one characteristic of the user-specified data set and at least one characteristic of the user-specified task and generates at least one machine learning model for processing the user-specified data set. The machine learning model generates a first output by processing the user-specified data set. The machine learning engine receives a natural language description of a user-requested data transformation task for execution with a subset of the first output and directs a large language model to identify an archetype of the user-requested data transformation task. The large language model applies the user-requested data transformation task to the subset of the first output using the archetype to generate a second output.