The invention provides a natural speech emotion synthesis and recognition method and
system fused with
deep learning, and relates to the technical field of
speech processing, and the method comprises the steps: obtaining to-be-processed speech data and corresponding text content, extracting acoustic feature representation and
semantic feature representation, building edge connection between a
time sequence frame node of an acoustic feature and a semantic unit node of a
semantic feature, constructing a bidirectional emotion association graph, calculating an edge weight, and propagating and updating based on graph
convolution operation to obtain fusion feature representation; inputting the fusion features into an emotion classifier to obtain an emotion state identifier, calculating an emotion target area
mask according to edge connection
weight distribution, and generating an emotion regulation and control parameter; and performing
speech synthesis based on the emotion regulation and
control parameters and performing consistency
verification to obtain synthesized speech and emotion deviation feedback information. According to the method, deep fusion of
acoustics and
semantics is realized through the bidirectional emotion association graph, and the
emotion recognition accuracy and the emotion expressive force of
speech synthesis are improved.