Amplifying Source Code Signals for Machine Learning Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Machine learning models trained for source code tasks face inefficiencies in identifying and learning source code signals, requiring large amounts of data and time, and may not be effective without reliable signal identification.

Innovation Solution

A method that involves identifying and amplifying source code signals, generating amplified code that is functionally equivalent to the original, and providing it for machine learning models to improve training efficiency, using a signal amplifier system that includes a code analyzer and re-writer to enhance signal visibility without altering the model's architecture or objective.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a statistical approach is used to train ML models to identify source code signals, then the model can learn source code tasks, but it requires relatively large amounts of training data and time

Engineering Contradiction:
Improvemodel effectivenessVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-processing source code data to identify and amplify source code signals before training the ML model. The signal amplifier analyzes source code and generates amplified versions with enhanced signals, preparing the data in advance so the model can learn more efficiently during training without requiring as much raw data or time

Inventive Principle:
Principle #10Preliminary action

2Reliability

If a statistical approach is used to train ML models to identify source code signals, then the model can learn source code tasks, but it requires relatively large amounts of training data

Engineering Contradiction:
Improvemodel effectivenessVSAvoidtraining data volume
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent applies parameter changes by transforming the source code data parameters through signal amplification. The signal amplifier modifies the source code to enhance the visibility and prominence of source code signals, changing the data characteristics so that less training data is needed to achieve the same learning effectiveness

Inventive Principle:
Principle #35Parameter changes

3Reliability

If traditional training methods are used, then ML models can be trained for source code tasks, but the training efficiency is low

Engineering Contradiction:
Improveprediction accuracyVSAvoidtraining efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent introduces an intermediary component - the signal amplifier - that sits between the raw source code data and the ML model training process. This intermediary analyzes and amplifies source code signals in the data before presenting it to the model, improving training efficiency while maintaining prediction accuracy

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20220253723A1Amplifying source code signals for machine learning
Publication Date: 2022.08.11 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20220253723A1 patent drawing
  • US20220253723A1 patent drawing
  • US20220253723A1 patent drawing

AI summary

Embodiments are disclosed for a method. The method includes identifying one or more source code signals in a source code. The method also include generating an amplified code based on the identified signals and the source code. The amplified code is functionally equivalent to the source code. Further, the amplified code includes one or more amplified signals. The method additionally includes providing the amplified code for a machine learning model that is trained to perform a source code relevant task.