The application discloses an end-to-end voice dialogue
system and method based on cross-
modal retrieval enhancement, comprising a voice input and
feature extraction module, a cross-
modal retrieval enhancement module, an end-to-end voice generation module and a voice
memory module. Through the
system and method process setting, direct mapping and generation from voice to voice can be realized, error accumulation problems caused by
modular architecture can be eliminated, and the overall robustness of the
system is improved. Through the end-to-end model, the training and deployment process is simplified, the
scalability and real-time interaction capability of the system are improved, and the system can be better applied to intelligent customer service and other application scenarios with extremely high requirements for accuracy and efficiency. Through the efficient cross-
modal retrieval enhancement generation mode, relevant field professional knowledge can be retrieved from a text
knowledge base in real time according to received audio, the accuracy and timeliness of the response information are ensured, direct, efficient and accurate retrieval from voice to text knowledge is realized, and the cognitive ability of the system in the
vertical field is enhanced.