A low-resource language-oriented speech large model adaptation method and device
By employing a dynamic query transformer and lightweight adaptive techniques, the query vector convergence problem of large speech models in low-resource language scenarios is solved, achieving efficient cross-modal speech signal and text alignment, and improving the model's generalization ability and computational efficiency.
CN122116885APending Publication Date: 2026-05-29MINZU UNIVERSITY OF CHINA +1
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MINZU UNIVERSITY OF CHINA
- Filing Date
- 2026-04-28
- Publication Date
- 2026-05-29
Smart Images

Figure CN122116885A_ABST
Abstract
The application discloses a speech large model adaptation method and device for a low-resource language, and relates to the technical field of natural language processing. The method comprises the following steps: first, frame a speech signal, input an encoder with frozen parameters, and extract a time sequence acoustic feature sequence; then, pass through a feature enhancement module composed of linear projection, two-layer Transform and a feedforward network, capture global acoustic context, fuse fine-grained acoustic features through a target modal encoder and linear mapping, construct an adaptation structure based on a dynamic query transformer, utilize an explicit inductive bias connected with time sequence classification to guide the dynamic generation of a query vector, and input the query vector into a large language model trained through a frozen training strategy to generate text. The application realizes efficient feature compression and cross-modal alignment of a variable-length speech signal in a low-resource scene, and realizes the best convergence of model performance under very few samples.
Need to check novelty before this filing date? Find Prior Art