LLM Speech Modality Training via SQA Data Composition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing large language models (LLMs) struggle to effectively integrate speech modality, leading to limitations in performing tasks such as speech-to-text and speech-to-text translations, due to overfitting and lack of in-context learning capabilities.
Innovation Solution
The method involves training LLMs on a combination of automatic speech recognition (ASR) data and speech comprehension test question-answer (SQA) data, with a greater proportion of SQA data, to enhance speech modality and improve in-context learning.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If LLMs are trained on ASR data to integrate speech modality, then speech processing capability is improved, but overfitting occurs and general contextual abilities degrade
Solution Approach 1:
The patent changes the training data composition parameter by using a greater proportion of SQA data (e.g., 70-90%) compared to ASR data (e.g., 10-30%). This parameter adjustment prevents overfitting to ASR tasks while maintaining speech processing capability, thereby resolving the contradiction between speech capability improvement and general ability preservation
2Measurement precision
If speech models are trained on task-specific data, then speech recognition accuracy is improved, but in-context learning capability is lost
Solution Approach 1:
The patent performs preliminary action by pre-training the LLM on extensive SQA datasets before fine-tuning on ASR data. This preliminary SQA training establishes strong in-context learning capabilities that enable the model to adapt to new speech tasks without losing general learning ability, thus resolving the contradiction between recognition accuracy and adaptability
3Ease of operation
If cascading systems are built to link ASR with LLMs, then voice-to-text transformation is enabled, but actual speech processing functionality is limited
Solution Approach 1:
The patent merges ASR data and SQA data into a unified training framework where the LLM is jointly trained on both types of data, with SQA data comprising the majority. This merging enables the LLM to directly process speech inputs and perform various speech-related tasks (translation, analysis, etc.) rather than being limited to simple voice-to-text transformation, thus resolving the contradiction between operational ease and functional versatility
Data Source
AI summary
Systems and methods are provided for enhancing the speech modality in a large language model (LLM) and for retaining in-context learning capabilities without overfitting to trained tasks. Systems obtain a first set of training data comprising tuples of a sample of speech combined with synthetically generated pairings of speech comprehension test questions and answers that correspond to the sample of speech and obtain a second set of training data comprising pairings of automatic speech recognition data. Systems generate and align a first set of encodings of the first set of training data and a second set of encodings of the second set of training data. Systems train the LLM on a greater amount of the first set of training data than the second set of training data and use the trained LLM to perform a natural language processing task.


