This invention discloses a method for collaborative multimodal
content generation in the field of
artificial intelligence, specifically a method for multimodal content collaborative generation, comprising the following steps: S1: Receiving multimodal generation requests input by the user; S2: Based on the
scenario to which the request belongs; S3: Storing data through three-dimensional classification, intelligent retrieval based on
semantic similarity, and
adaptive optimization of historical resources; S4: First, based on a pre-trained cross-
modal semantic model, combined with constraint reports and reused resource packages; S5: Real-time monitoring of the
system's computing power status, combined with the differences in resource requirements of each modality generator and the urgency of the user's request; S6: Scheduling each modality generator according to semantic tags, association relationships, constraint
adaptation reports, and scheduling strategies. This invention constructs a unified multimodal
semantic space, maps the initial features of each modality to this space, and establishes
modal semantic tag association relationships, achieving strong semantic associations in multimodal content, avoiding semantic conflicts between different modalities, and improving content consistency.