The invention relates to a toy
control system based on multi-
modal interaction, which is integrated in a toy main body and comprises an instruction receiving unit for receiving a voice instruction through a
microphone array; the
instruction processing unit is used for carrying out localized recognition and semantic understanding on the voice instruction and generating a corresponding control instruction based on an action
label output by a local semantic understanding model; and the
action control unit is used for driving a motor to execute actions according to the control instruction and comprises advancing, retreating, head rotating or arm swinging of the toy main body. According to the
control system, three interaction
modes of voice, vision and action are fused, a localized
edge computing framework is adopted, voice recognition
response delay is controlled within 300 milliseconds, action and vision feedback are synchronously coordinated,
time sequence dislocation is avoided, instant and accurate interaction is ensured, collaborative response of speaking, watching and moving at the same time is achieved, and the
control system has the advantages of being simple in structure, convenient to operate and high in practicability. And the natural fluency of interaction and the use immersion of children are obviously improved.