The invention relates to the technical field of voice
processing, can be applied to business scenes of financial science and technology,
medical health and the like, and discloses a voice text bidirectional
conversion method, device, equipment and medium, and the method comprises the steps: respectively executing voice recognition or voice synthesis operation according to the type of input information; for the voice information,
noise suppression parameters are generated in combination with the lip movement video data,
noise reduction
processing is executed, and the recognition accuracy is improved; for text information, a pre-generated speaker style vector is obtained, the vector is cited in the
speech synthesis process to generate natural personalized speech, and lip movement information and tactile feedback which are synchronous with
speech output are generated. According to the method, complex
noise is suppressed by fusing lip movement data, personalized voice is generated by using the style vector, and lip movement and touch information is output, so that bidirectional real-time conversion of voice and text in a complex environment is realized, and recognition accuracy, voice naturalness and interaction
synchronism are effectively improved.