This invention relates to the field of voice interaction in smart devices, and discloses a cascaded voice
interaction method for
moxibustion devices. The method collects user voice, preprocesses it to obtain effective voice frames and acoustic feature maps; inputs the feature maps into a first-level voice
processing unit, runs a low-power wake-word detection and short command recognition model, and directly outputs control commands when short commands are recognized; if no short command is recognized, it calculates a complex intent tendency
score, and if the
score exceeds a threshold, it wakes up a second-level voice
processing unit; the second-level unit runs a locally deployed lightweight large
language model to perform end-to-end voice recognition and semantic understanding, and generates interactive commands based on a
moxibustion knowledge base; finally, the commands are executed and the results are broadcast. This invention, through a two-level cascaded architecture, balances low
power consumption and high intelligence, achieves offline complex semantic understanding, and improves the voice interaction experience in
moxibustion scenarios.