Fine-Tuned LLM Testing for Domain-Specific Code Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing large language models (LLMs) trained on general datasets are ineffective for domain-specific code generation due to the lack of tailored training data, leading to inaccurate and inefficient code generation.
Innovation Solution
A method and system for testing a fine-tuned LLM using a domain-specific test dataset, which includes determining LLM-generated problem statements and codes, and evaluating accuracy based on a percentage match with test codes to ensure precision.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If LLMs are trained on general datasets, then training data availability is improved, but code generation accuracy for domain-specific tasks deteriorates
Solution Approach 1:
The patent applies local quality by creating domain-specific test datasets that are tailored to evaluate LLM performance in particular domains (e.g., healthcare, finance, engineering). Instead of using generic code datasets, the system generates test cases specific to the target domain, allowing evaluation of domain-specific code generation capabilities while maintaining the ability to use general training data for model training.
2Manufacturing precision
If LLMs are fine-tuned on domain-specific training data, then code generation accuracy for the domain is improved, but the complexity of data preparation and evaluation deteriorates
Solution Approach 1:
The patent implements self-service by enabling the system to automatically generate domain-specific test datasets and evaluation metrics without requiring manual intervention. The system can autonomously extract domain concepts, generate appropriate test cases, and evaluate model performance, thereby reducing the complexity burden on data preparation and evaluation processes while maintaining high domain-specific accuracy.
3Ease of operation
If conventional general code datasets are used for evaluation, then evaluation process simplicity is improved, but evaluation relevance to domain-specific tasks deteriorates
Solution Approach 1:
The patent applies dynamics by creating a flexible evaluation system that can adapt to different domains. Instead of using static general datasets, the system dynamically generates domain-specific test datasets based on the target domain characteristics. This allows the evaluation process to maintain simplicity while becoming relevant to specific domain tasks, as the test dataset can be automatically adjusted to match the domain being evaluated.
Data Source
AI summary
A method and system of testing a fine-tuned LLM for domain specific code generation is disclosed. Further, a processor receives a test dataset corresponding to a domain from a code repository. Further, the processor determines an LLM generated problem statement corresponding to the test code using the fine-tuned LLM. The fine-tuned LLM is fine-tuned based on a training dataset. Further, the fine-tuned LLM is prompted based on the LLM generated problem statement to determine an LLM generated code for a corresponding test function. The accuracy level of the fine-tuned LLM is determined based on a percentage match between the LLM generated code with the test code for each of the set of test functions.


