Machine Reading Comprehension Model Noise Resistance Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The complexity and difficulty of training machine reading comprehension models are increased by adding more data from other sources, which complicates the training process and requires manual intervention to enhance noise resistance.
Innovation Solution
A method is proposed where an initial model is trained to generate an intermediate model, noise text is automatically generated and added to samples to create noise samples, and correction training is performed on the intermediate model to improve its noise resistance without modifying the model or requiring manual participation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If more data from other sources is added to the model as input to enhance noise resistance, then the anti-noise capability is improved, but the model complexity and training difficulty increase
Solution Approach 1:
The patent applies preliminary action by pre-processing the input text to identify and remove noise words before the model processes the data. This preliminary cleaning step prevents noise from interfering with the model's learning process, thereby improving anti-noise capability without increasing model complexity or training difficulty
Solution Approach 2:
The patent extracts and removes harmful noise elements from the input data using a predefined noise word list. By taking out the noise components before processing, the model receives cleaner input data, which improves reliability without requiring the model itself to become more complex
2Reliability
If more data from other sources is added to the model as input to enhance noise resistance, then the anti-noise capability is improved, but the training complexity increases
Solution Approach 1:
The patent performs preliminary action by pre-processing the input text to identify and remove noise words before the model processes the data. This preliminary cleaning step prevents noise from interfering with the model's learning process, thereby improving anti-noise capability without increasing model complexity or training difficulty
Solution Approach 2:
The patent extracts and removes harmful noise elements from the input data using a predefined noise word list. By taking out the noise components before processing, the model receives cleaner input data, which improves reliability without requiring the model itself to become more complex
3Reliability
If manual intervention is used to enhance noise resistance, then the anti-noise capability is improved, but the cost and complexity increase
Solution Approach 1:
The patent implements self-service by automatically identifying and removing noise words using a predefined noise word list and automated processing rules. This eliminates the need for manual intervention in noise filtering, thereby improving anti-noise capability while reducing training cost and complexity
Solution Approach 2:
The patent extracts and removes harmful noise elements from the input data using a predefined noise word list. By taking out the noise components before processing, the model receives cleaner input data, which improves reliability without requiring the model itself to become more complex
Data Source
AI summary
Embodiments of the present disclosure provide a method and an apparatus for training a machine reading comprehension model, and a storage medium. The method includes: training an initial model to generate an intermediate model based on sample data; extracting samples to be processed from the sample data according to a first preset rule; generating a noise text according to a preset noise generation method; adding the noise text to each of the samples to be processed respectively to generate noise samples; and performing correction training on the intermediate model based on the noise samples to generate the machine reading comprehension model.


