The invention discloses an intelligent
processing method and
system for multi-
modal data of an oral broadcast
video based on
deep learning, and relates to the technical field of oral broadcast video
data processing, and the method comprises the following steps: 1, collecting multi-
modal oral
broadcast data in real time, recognizing emotion types, and carrying out the data preprocessing; step 2, constructing an oral playing
animation generation model, judging the synchronization degree of the
mouth shape, the head action and the audio in the oral playing
animation, and taking coordinated regulation measures; 3, judging the consistency of the
mouth shape and the emotion intensity in the oral playing
animation, and implementing an anomaly repair measure; 4, constructing a depth
time sequence prediction model, evaluating the difference degree between the actual oral playing animation and the prediction result of the depth
time sequence prediction model, and carrying out
video optimization and
model tuning; the problems that when
mouth shape supplement and animation filling are carried out by AI, precise synchronization of mouth shapes, expressions, head actions and emotional fluctuations is difficult to achieve, so that visual abnormity is caused, and the reality sense and the
professional degree of videos are affected are solved.