Lingual Matching Tuning for Hard-Pair Dataset Refinement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing datasets for finetuning classification models often include 'easy' examples with very similar or very different representations, which do not contribute value during subsequent training processes, leading to suboptimal model performance.

Innovation Solution

A system and method that iteratively processes a dataset to reduce the frequency of highly similar and dissimilar data pairs, using modules to evaluate and refine the training dataset through resampling, false sample pair generation, and metric calculation to enhance model performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If datasets include easy examples with very similar or very different representations, then the dataset size increases, but the training effectiveness decreases

Engineering Contradiction:
Improvedataset sizeVSAvoidtraining effectiveness
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent applies parameter changes by modifying the similarity threshold parameter to dynamically filter training examples. By adjusting this threshold, the system selectively includes or excludes examples based on their similarity metrics, thereby changing the composition of the training dataset to optimize both size and effectiveness simultaneously

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If similar records are matched via finetuning to receive similar vector embeddings, then classification accuracy improves, but computational resources increase

Engineering Contradiction:
Improveclassification accuracyVSAvoidcomputational resources
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent applies preliminary action by pre-computing and storing vector embeddings for all training records before the actual finetuning process. This preliminary embedding generation allows the model to efficiently compare and match similar records during training without performing computationally intensive calculations in real-time, thereby reducing overall computational resource requirements

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20260080303A1System and method for refining lingual matching tuning
Publication Date: 2026.03.19 INTUIT INC
  • US20260080303A1 patent drawing
  • US20260080303A1 patent drawing
  • US20260080303A1 patent drawing

AI summary

A system and method are provided for refining lingual matching tuning.