This invention discloses a lightweight intent-assisted input scheme based on speech prosodic features, belonging to the field of
artificial intelligence and voice
interaction technology. The scheme utilizes a local
speech synthesis or recognition front-end module on the
client side to extract prosodic features such as tone, intonation, pauses, and
rhythm from the user's speech, quantifying them into lightweight parameter labels with extremely small volumes, such as question level, exclamation level, pause length, and tone intensity. These labels are then encapsulated with the speech-to-text and sent to the
server. The
server-side
intent recognition model uses these labels as
prior information to dynamically adjust the
inference path,
pruning invalid intent branches, reducing computational consumption, and improving recognition efficiency and accuracy. This invention does not change the core model structure, does not significantly increase the transmission burden, and can reuse existing speech module logic, achieving front-end computation and
system cost reduction and efficiency improvement. This invention is applicable to various voice assistants and intelligent dialogue systems, and has broad application prospects.