Semantic Similarity Model Training via Dataset Correlation Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The integration of multiple types of datasets in the fine-tuning phase for training semantic similarity models results in extreme and undesirable accuracy, as current methods fail to effectively utilize high-quality annotation data and correlate datasets with target fields.

Innovation Solution

A method is introduced to calculate correlations between a target field and application fields for training datasets, allowing the semantic similarity model to be trained with these datasets in a more targeted manner, dividing datasets into high-correlation and low-correlation sets for sequential training to enhance learning capability and accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If multiple types of datasets are integrated together to train the semantic similarity model in the fine-tuning phase, then the training effect is improved, but the accuracy of the trained semantic similarity model becomes extreme and undesirable

Engineering Contradiction:
Improvetraining effectVSAvoidaccuracy of semantic similarity model
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent segments the training process by dividing multiple datasets into different groups based on their correlation with the target field. Instead of integrating all datasets together, the method separates them into high-correlation and low-correlation groups, training the model with each group separately in sequence. This segmentation resolves the contradiction by preventing the negative interaction between diverse datasets while still utilizing their respective training benefits.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by treating different datasets differently based on their specific characteristics and correlation with the target field. High-correlation datasets are used for primary fine-tuning to maximize accuracy, while low-correlation datasets are used for supplementary training. This differentiated approach ensures that each dataset contributes optimally to the model's performance without causing accuracy degradation.

Inventive Principle:
Principle #3Local quality

2Reliability

If high-quality annotation data is used for fine-tuning the pre-trained semantic similarity model, then the model potential is mined and model effect is improved, but the accuracy becomes extreme when multiple datasets are integrated

Engineering Contradiction:
Improvemodel effectVSAvoidaccuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent performs preliminary action by calculating the correlation between the target field and each dataset's application field before training. This pre-analysis allows the method to determine the optimal training sequence and select which datasets to use, preventing accuracy degradation before it occurs. The correlation calculation is performed in advance to guide the subsequent training process.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces dynamics by making the training process adaptive to the target field characteristics. The training sequence and dataset selection are not fixed but are determined dynamically based on the calculated correlations between the target field and dataset fields. This dynamic approach allows the model to benefit from high-quality annotation data while avoiding accuracy extremes.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12118063B2Method, apparatus, electronic device and storage medium for training semantic similarity model
Publication Date: 2024.10.15 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US12118063B2 patent drawing
  • US12118063B2 patent drawing
  • US12118063B2 patent drawing

AI summary

The present disclosure provides a method, apparatus, electronic device and storage medium for training a semantic similarity model, which relates to the field of artificial intelligence. A specific implementation solution is as follows: obtaining a target field to be used by a semantic similarity model to be trained; calculating respective correlations between the target field and application fields corresponding to each of training datasets in known multiple training datasets; training the semantic similarity model with the training datasets in turn, according to the respective correlations between the target field and the application fields corresponding to each of the training datasets. According to the technical solution of the present disclosure, it is possible to, in the fine-tuning phase, more purposefully train the semantic similarity model with the training datasets with reference to the correlations between the target field and the application fields corresponding to the training datasets, thereby effectively improving the learning capability of the sematic similarity model and effectively improving the accuracy of the trained semantic similarity model.