LLM Speech Modality Training via SQA Data Composition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing large language models (LLMs) struggle to effectively integrate speech modality, leading to limitations in performing tasks such as speech-to-text and speech-to-text translations, due to overfitting and lack of in-context learning capabilities.

Innovation Solution

The method involves training LLMs on a combination of automatic speech recognition (ASR) data and speech comprehension test question-answer (SQA) data, with a greater proportion of SQA data, to enhance speech modality and improve in-context learning.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If LLMs are trained on ASR data to integrate speech modality, then speech processing capability is improved, but overfitting occurs and general contextual abilities degrade

Engineering Contradiction:
Improvespeech processing capabilityVSAvoidgeneral contextual abilities
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent changes the training data composition parameter by using a greater proportion of SQA data (e.g., 70-90%) compared to ASR data (e.g., 10-30%). This parameter adjustment prevents overfitting to ASR tasks while maintaining speech processing capability, thereby resolving the contradiction between speech capability improvement and general ability preservation

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If speech models are trained on task-specific data, then speech recognition accuracy is improved, but in-context learning capability is lost

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidin-context learning capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent performs preliminary action by pre-training the LLM on extensive SQA datasets before fine-tuning on ASR data. This preliminary SQA training establishes strong in-context learning capabilities that enable the model to adapt to new speech tasks without losing general learning ability, thus resolving the contradiction between recognition accuracy and adaptability

Inventive Principle:
Principle #10Preliminary action

3Ease of operation

If cascading systems are built to link ASR with LLMs, then voice-to-text transformation is enabled, but actual speech processing functionality is limited

Engineering Contradiction:
Improvevoice-to-text transformationVSAvoidspeech processing functionality
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The patent merges ASR data and SQA data into a unified training framework where the LLM is jointly trained on both types of data, with SQA data comprising the majority. This merging enables the LLM to directly process speech inputs and perform various speech-related tasks (translation, analysis, etc.) rather than being limited to simple voice-to-text transformation, thus resolving the contradiction between operational ease and functional versatility

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20250140238A1Methods and systems for enhancing multimodal capabilities in large language models
Publication Date: 2025.05.01 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20250140238A1 patent drawing
  • US20250140238A1 patent drawing
  • US20250140238A1 patent drawing

AI summary

Systems and methods are provided for enhancing the speech modality in a large language model (LLM) and for retaining in-context learning capabilities without overfitting to trained tasks. Systems obtain a first set of training data comprising tuples of a sample of speech combined with synthetically generated pairings of speech comprehension test questions and answers that correspond to the sample of speech and obtain a second set of training data comprising pairings of automatic speech recognition data. Systems generate and align a first set of encodings of the first set of training data and a second set of encodings of the second set of training data. Systems train the LLM on a greater amount of the first set of training data than the second set of training data and use the trained LLM to perform a natural language processing task.