The invention relates to the technical field of speech
semantics, and discloses a context-based multi-
modal data generation method, device and equipment and a medium, and the method comprises the steps: carrying out the semantic segmentation of context text information through a semantic recognition model, and determining a target semantic tag;
multimedia materials are obtained, association is carried out according to the
multimedia materials and the context text information, and initial multi-
modal data are generated; and generating target multi-
modal data according to the initial multi-
modal data and the multi-
modal data template. According to the mode, the target semantic
label of the context text is extracted through the preset semantic recognition model, the text subjected to voice transcription is dynamically associated with the
multimedia material according to the semantic
label, and the multi-
modal data integrating the image-text, the audio and the video is generated, so that a user only needs natural dialogue and does not need typewriting or complex operation, and the user experience is improved. In the business fields of financial science and technology,
medical health, old-age care and the like, the interaction efficiency between the intelligent voice interaction
system and the user is improved.