Multi-modal task scheduling method, apparatus, device, and medium

By analyzing the task description information and computing node characteristic data of multimodal tasks, a task scheduling strategy is generated, which solves the problem of low resource utilization in multimodal task scheduling and realizes efficient and stable multimodal task execution.

CN122431813APending Publication Date: 2026-07-21广东南方新媒体股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
广东南方新媒体股份有限公司
Filing Date
2026-03-18
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Current multimodal task scheduling methods lack the ability to automatically identify and adaptively schedule the characteristics of multimodal tasks, resulting in problems such as low computing power utilization, extended task queuing time, uneven load on computing nodes, or idle resources.

Method used

By using task description information based on multimodal tasks and node feature data of computing nodes, a task scheduling strategy is generated, including the determination of task feature data, the acquisition of node feature data, the generation of task scheduling strategy, and the distribution of execution tasks, thereby realizing automated parsing and dynamic optimization scheduling of tasks.

Benefits of technology

It significantly improves the resource utilization and load balancing of heterogeneous computing power clusters, enhances the adaptability and system robustness of multimodal generation tasks, shortens execution time, and improves computing power utilization and task execution efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122431813A_ABST
    Figure CN122431813A_ABST
Patent Text Reader

Abstract

The application discloses a multi-modal task scheduling method and device, equipment and medium. Through the automatic analysis and structured decomposition of complex multi-modal creation intention, the natural language description of the user is intelligently converted into multiple execution tasks with clear dependency relationship. Then, through the accurate matching between task characteristics and node characteristics, dynamic optimization scheduling is carried out according to the real-time resource state and hardware capability to generate a scheduling strategy. The overload or resource idling of the computing node is effectively avoided, the overall resource utilization and load balancing level of the heterogeneous computing power cluster are significantly improved, the adaptability and system robustness of the multi-modal generation task in the dynamic environment are enhanced, thereby an efficient, stable and scalable automatic scheduling strategy is provided for the multi-modal AI creation task, the execution time of the complex multi-modal generation task is greatly shortened, and the computing power utilization and task execution efficiency are significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a multimodal task scheduling method, apparatus, device, and medium. Background Technology

[0002] Current AI (Artificial Intelligence) creation generally involves multimodal tasks, which typically require calling different large-scale models and consuming heterogeneous computing resources. However, most current multimodal task scheduling methods are geared towards single-modal or fixed-model scenarios, lacking the ability to automatically identify and adaptively schedule the characteristics of multimodal tasks. This can easily lead to problems such as low system computing power utilization, extended task queuing times, uneven load on computing nodes, or idle resources. Summary of the Invention

[0003] In view of this, the present invention provides a multimodal task scheduling method, apparatus, electronic device and medium to solve the technical problems of insufficient multimodal task scheduling capabilities, resulting in inefficient computing power utilization and uneven resource allocation.

[0004] Firstly, a multimodal task scheduling method is provided, which includes: Based on the task description information of multimodal tasks, the task feature data of multimodal tasks are determined. The task feature data includes multiple execution tasks corresponding to the multimodal task, the task modality of each execution task, the task requirements of each execution task, the task dependencies between multiple execution tasks, and the resource requirements for executing each execution task. Obtain node characteristic data of multiple computing power nodes, including resource status data and capability attribute information of each computing power node; Based on task feature data and node feature data, a task scheduling strategy for multimodal tasks is generated. The task scheduling strategy includes the target computing power node corresponding to each task and the execution time window of each task. Based on the task scheduling strategy, each execution task is distributed to its corresponding target computing power node to generate multimodal works corresponding to multimodal tasks.

[0005] Secondly, a multimodal task scheduling device is provided, the device comprising: The determination module is used to determine the task feature data of the multimodal task based on the task description information of the multimodal task. The task feature data includes multiple execution tasks corresponding to the multimodal task, the task mode of each execution task, the task requirements of each execution task, the task dependencies between multiple execution tasks, and the resource requirements for executing each execution task. The acquisition module is used to acquire node characteristic data of multiple computing power nodes, including resource status data and capability attribute information of each computing power node. The first generation module is used to generate a task scheduling strategy for multimodal tasks based on task feature data and node feature data. The task scheduling strategy includes the target computing power node corresponding to each task and the execution time window of each task. The second generation module is used to distribute each execution task to its corresponding target computing power node based on the task scheduling strategy, and generate multimodal works corresponding to multimodal tasks.

[0006] Thirdly, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described multimodal task scheduling method.

[0007] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described multimodal task scheduling method.

[0008] The aforementioned multimodal task scheduling method, device, electronic equipment, and storage medium firstly, through automated parsing and structured decomposition of complex multimodal creative intentions, intelligently transforms the user's natural language description into multiple execution tasks with clear dependencies. Subsequently, by constructing a bidirectional precise matching mechanism between task features and node features, dynamic optimization scheduling is performed based on real-time resource status and hardware capabilities to generate a global scheduling strategy. Each execution task is then distributed to its corresponding target computing node for triggering execution, ultimately synthesizing a complete cross-modal work. This effectively avoids computing node overload or resource idleness, significantly improves the overall resource utilization and load balancing level of heterogeneous computing power clusters, enhances the adaptability and system robustness of multimodal generation tasks in dynamic environments, and thus provides an efficient, stable, and scalable automated scheduling strategy for multimodal AI creation tasks. This significantly shortens the execution time of complex multimodal generation tasks and significantly improves computing power utilization and task execution efficiency. Attached Figure Description

[0009] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 This is a flowchart illustrating a multimodal task scheduling method according to an embodiment of the present invention; Figure 2This is a schematic diagram of the structure of a multimodal task scheduling device in one embodiment of the present invention. Detailed Implementation

[0010] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in the present invention are only for illustrative and descriptive purposes and are not intended to limit the scope of protection of the present invention.

[0011] Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this invention illustrate operations implemented according to some embodiments of the invention. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or performed simultaneously. Moreover, those skilled in the art, guided by the content of this invention, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.

[0012] Furthermore, the embodiments described herein are merely some, not all, of the embodiments of the invention. The components of the embodiments of the invention described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0013] It should be noted that the term "comprising" will be used in the embodiments of the present invention to indicate the presence of a feature subsequently declared, but does not exclude the addition of other features. It should also be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0014] The following is a detailed description of this case, in conjunction with the relevant accompanying drawings in the instruction manual.

[0015] Please see Figure 1 This specification provides a multimodal task scheduling method, which specifically includes the following steps: S10: Based on the task description information of the multimodal task, determine the task feature data of the multimodal task; The task feature data includes multiple execution tasks corresponding to the multimodal task, the target task modality of each execution task, the task requirements of each execution task, the task dependencies between multiple execution tasks, and the resource requirements for executing each execution task.

[0016] It is understood that the executing entity of this invention can be a multimodal task scheduling device, a terminal, or a server; no specific limitation is made here. This embodiment of the invention will be described using a server as an example.

[0017] In this step, the task description information of the multimodal task submitted by the user is received. The multimodal task is a multimodal AI creation task, and the task description information typically carries the user's creative intent in the form of natural language or structured instructions, such as "generate a short video with a spring theme, including background music and subtitles." Upon receiving the task description information, it is intelligently parsed and decomposed into a workflow to determine the task feature data of the multimodal task. This includes the multiple atomic single-modal execution tasks corresponding to the multimodal task. For example, the multimodal task "generate a short video with a spring theme, including background music and subtitles" includes single-mode execution tasks such as "generate descriptive text," "generate background image," "synthesize background music," "generate subtitles," and "synthesize video." The task feature data also includes the target task modality of each execution task, such as text generation, image generation, speech synthesis, or music composition; the user's task requirements for each execution task; the task dependencies between multiple execution tasks established by logic or data flow, such as text generation must precede image and subtitle generation, and final synthesis can only be performed after all media materials are ready; and the resource requirements for executing each task, including estimated video memory usage, computation time, and storage I / O requirements.

[0018] The above methods can automatically parse complex creative intentions, decompose multimodal tasks into independently schedulable sub-tasks, and accurately determine the task characteristics of each task, thereby providing a structured data foundation for flexible and intelligent workflow orchestration.

[0019] In one embodiment of this application, a specific feature extraction scheme is provided. In S10, that is, based on the task description information of the multimodal task, the task feature data of the multimodal task is determined, which specifically includes the following steps S11-S15: S11: In response to the multimodal generation task request, obtain the task description information included in the multimodal generation task request.

[0020] In this step, when a multimodal generation task request is received from the front-end workflow or user interface, the task description information submitted by the user is extracted from the request. This information usually carries the user's comprehensive creative intent in the form of natural language or structured instructions, such as "create a digital music video that blends Chinese-style melodies with the artistic conception of landscape paintings".

[0021] S12: Based on the task description information, perform workflow decomposition on the multimodal task to determine multiple execution tasks, the task requirements of each execution task, and the task dependencies between multiple execution tasks.

[0022] In this step, based on the acquired task description information, the multimodal task is intelligently decomposed into a workflow. Through semantic understanding and intent recognition technology, the complex task is broken down into multiple atomic execution tasks that can be executed and scheduled independently. Subsequently, the specific task requirements for each execution task are determined, such as output format, style constraints, quality level, etc., and the task dependencies between these execution tasks are parsed.

[0023] In one embodiment of this application, a specific task decomposition scheme is provided. In S12, based on task description information, a workflow decomposition is performed on the multimodal task to determine multiple execution tasks, the task requirements of each execution task, and the task dependencies between the multiple execution tasks. Specifically, this includes the following steps S121-S122: S121: Based on the large language model, perform semantic parsing on the task description information to determine the multiple execution tasks included in the multimodal task, as well as the task requirements of each execution task.

[0024] In this step, the task description information submitted by the user typically contains multiple implicit, logically separable creative intentions at the linguistic level, each representing a task to be performed. When a composite description such as "create a short film that blends ink-wash animation with guqin melodies" is received, the description, along with the pre-defined system capability framework, is submitted to the large language model. Based on its deep understanding of semantics, common-sense knowledge, and task planning capabilities, the large language model outputs a structured parsing result. This result clearly indicates several specific tasks to be performed, such as: "generate a sequence of ink-wash style animations," "create a melody with guqin timbre," and "synchronize the animation and audio." Simultaneously, the large language model also extracts and associates the specific requirements of each task from the original description. For example, it specifies "black, white, and gray main color scheme and smooth brushstrokes" for the ink-wash animation task, "pentatonic scale and soothing rhythm" for the guqin melody task, and "audio-visual synchronization and addition of fade-in and fade-out transitions" for the synthesis task.

[0025] The above method enables the automated conversion from user natural language commands to a refined, operable task list and its specifications, providing accurate data for subsequent model matching, resource scheduling, and task execution.

[0026] S122: Based on preset workflow construction rules, determine the task dependencies between multiple execution tasks.

[0027] In this step, the preset workflow construction rules are essentially a formalized dependency determination logic library, defining the possible dependency patterns between different types of tasks. For example, the rules include: the video rendering task depends on the completion of all visual material generation tasks; the audio mixing task depends on the output of both voice synthesis and background music generation tasks; the subtitle embedding task must be executed after video rendering is completed but can be performed in parallel with audio synthesis, etc. By matching and reasoning with these rules to the identified set of execution tasks, the dependencies between tasks are automatically determined. These dependencies clarify which tasks can be executed in parallel, which tasks must be executed sequentially, and the data transfer paths between tasks, providing a process topology structure for subsequent intelligent scheduling and resource allocation.

[0028] S13: Based on the preset task modality set, determine the target task modality corresponding to each execution task.

[0029] In this step, a task modality classification system is pre-constructed, forming a preset task modality set. This set covers common content generation types in the current AI creation field, including but not limited to text generation, image generation, speech synthesis, music composition, video rendering, and 3D model generation. Each modality clearly defines its core functional boundaries, typical input / output formats, and related processing characteristics. After decomposing the multimodal task into multiple execution tasks, a pre-trained deep learning model is used to map each execution task to the most matching modality category in the preset task modality classification set, based on the content nature, processing objective, and required result format. This clarifies the technical field to which each execution task belongs and the model type to be invoked, providing crucial type basis for subsequent model selection, resource estimation, and scheduling strategy formulation.

[0030] For example, the task of “generating story narration” is classified as the “speech synthesis” modality; the task of “generating illustrations for the story” is classified as the “image generation” modality.

[0031] S14: Based on the target task modality and task requirements, determine the target task model corresponding to each execution task.

[0032] In this step, after determining the target task modality and its specific task requirements for each execution task, a pre-defined model resource knowledge base is accessed. This knowledge base registers various available AI models and records information such as the modality types supported by each model, its preferred style categories, supported output specifications, required operating environment, and key performance indicators. First, all candidate models that match the target task modality category are selected. Then, based on constraints such as style, specifications, and quality in the task requirements, matching and optimization are performed to select the model with the best overall performance as the target task model for that execution task.

[0033] By employing the above methods, we can ensure that each task is executed by an AI model with highly matched capabilities and optimal output.

[0034] In one embodiment of this application, a specific scheme for determining the task model is provided. In S14, that is, based on the target task mode and task requirements, the target task model corresponding to each execution task is determined, specifically including the following steps S141-S143: S141: Based on the target task modality, determine at least one candidate task model corresponding to each execution task.

[0035] In this step, after determining the target task modality for each execution task, a pre-built model registry is accessed. Each model in this registry is labeled with one or more task modalities it supports. Using the target task modality as the query condition, all labeled model entries supporting that modality are quickly retrieved and returned, forming a candidate task model set for that task. For example, for the image super-resolution modality, the candidate set includes ESRGAN, Real-ESRGAN, and SwinIR models.

[0036] S142: Obtain the attribute information of each candidate task model.

[0037] S143: Based on task requirements and attribute information, determine the target task model for each task to be performed from at least one candidate task model.

[0038] For steps S142-S143, after obtaining the candidate task model set for each execution task, the complete attribute information of each candidate model is extracted from the model resource library. This attribute information includes, but is not limited to: model architecture characteristics, training dataset type, supported specific style variants, output resolution range, inference speed metrics, memory usage curve, quantization support level, and metadata such as licensing agreements and invocation costs. Subsequently, the specific task requirements of each execution task, such as generating a 4K resolution landscape painting style image with a rendering time of less than 5 seconds, are compared with the attribute information of each candidate model. The candidate models are comprehensively evaluated across multiple dimensions using preset scoring rules or optimization algorithms. Finally, the model with the highest comprehensive score or the one that best meets the specific optimization objective (such as lowest cost or fastest speed) is selected as the target task model for that execution task.

[0039] By employing the above methods, we can ensure that each task is executed by the AI ​​model with the best matching attributes and optimal performance, effectively improving the quality and efficiency of multimodal creation tasks.

[0040] S15: Based on the target task model and task requirements, determine the resource requirements data needed to execute each task.

[0041] In this step, after selecting the target task model for each execution task, and considering the specific requirements of the task, such as "generating a 1024x1024 pixel image, setting the iteration step count to 50, and using a specific sampler," refined resource calculations are performed. First, the model's baseline resource profile is queried to obtain its typical resource consumption under standard configuration, including basic GPU memory usage and typical computational cost per inference. Then, based on the specific parameters and requirements of the current task, the baseline resource consumption is dynamically adjusted. For example, higher output resolution, more iteration steps, or more complex sampling methods will linearly or non-linearly increase GPU memory requirements and computation time. Furthermore, if the task requirements include constraints such as real-time generation or low latency, the estimated level of computing resources (such as GPU computing power) will be increased accordingly. Finally, the resource requirements for the task execution are comprehensively calculated, specifically including the estimated peak GPU memory usage, expected CPU core usage, memory usage, estimated execution time, and potential disk I / O or network bandwidth requirements. These quantified resource demand data provide a reliable scientific basis for subsequent precise scheduling and resource allocation.

[0042] In one embodiment of this application, a specific resource requirement data determination scheme is provided. In S15, that is, based on the target task model and task requirements, the resource requirement data required to execute each task is determined, specifically including the following steps S151-S152: S151: Obtain the baseline resource requirements data for each target task model.

[0043] S152: Based on the task requirements of each execution task, the baseline resource requirement data is corrected to determine the resource requirement data for each execution task.

[0044] For steps S151-S152, after each execution task selects a target task model, the baseline resource requirement data for that model is retrieved from the model resource archive. This data records the model's baseline resource requirements under preset standard configurations, including but not limited to GPU memory usage, CPU computational load, memory usage, and single inference time. Subsequently, combined with the specific task requirements of each execution task, such as parameterized constraints like increasing the output resolution to 2048×2048, increasing the sampling steps to 100, and enabling high-definition restoration, the baseline resource requirement data is refined and corrected. The correction process quantifies the impact coefficients of specific parameter changes on various resource requirements based on preset resource consumption adjustment algorithms or empirical models. Through the comprehensive calculation of the baseline value and correction coefficients, the system finally generates accurate resource requirement data for that task execution. This data provides crucial quantitative input for the subsequent scheduling system to perform resource matching, node selection, and load prediction, and is the foundation for achieving efficient and accurate resource scheduling.

[0045] S20: Obtain node characteristic data of multiple computing nodes; Among them, node characteristic data includes resource status data and capability attribute information of each computing node.

[0046] In this step, lightweight monitoring agents deployed on each computing node continuously collect node characteristic data from multiple computing nodes in the cluster, including resource status data and capability attribute information. Resource status data reflects the real-time load of the nodes, including but not limited to the current utilization rate of GPU / CPU, available video memory and system memory capacity, network bandwidth utilization, and task queue length. Capability attribute information describes the inherent configuration and capabilities of the nodes, such as the GPU model and computing power level, CUDA driver version, deployed deep learning frameworks and model libraries, and storage media type and capacity.

[0047] S30: Generate a task scheduling strategy for multimodal tasks based on task feature data and node feature data; The task scheduling strategy includes the target computing power nodes corresponding to each task and the execution time window for each task.

[0048] In this step, based on the aforementioned task and node characteristic data, intelligent scheduling calculations are performed. This involves simultaneously balancing task dependencies, resource requirements, real-time node load, hardware capability matching, and global efficiency goals to generate an optimal task scheduling strategy for the multimodal task. The generated task scheduling strategy not only assigns a target computing node for each task but also plans the execution time window for each task, thereby ensuring that the entire workflow functions like a highly efficient and collaborative intelligent production system.

[0049] By accurately matching and dynamically scheduling task requirements with node capabilities, the problem of node overload or resource idleness is effectively avoided, and the cluster's computing power resources are balanced.

[0050] In one embodiment of this application, a specific task scheduling strategy generation scheme is provided. In S30, that is, based on task feature data and node feature data, a task scheduling strategy for multimodal tasks is generated, which specifically includes the following steps S31-S33: S31: Match and verify the resource requirements data of each task with the capability attribute information of each computing node to determine the candidate node set for each task.

[0051] In this step, after obtaining the precise resource requirements for each task, these requirements are systematically matched and verified against the capability attribute information registered by each computing node in the cluster. First, it checks whether the node hardware meets the basic requirements, such as whether the GPU model supports the CUDA version and whether the total video memory capacity exceeds the peak requirements of the task. Second, it verifies software environment compatibility, such as whether the node has installed the specified version of the deep learning framework or specific dependency libraries. Finally, it confirms runtime availability, such as whether the node has deployed the target model required for the task or a compatible runtime. Only nodes that pass all levels of verification are included in the candidate node set for that task.

[0052] For example, for an image generation task that requires an Ampere architecture GPU, more than 24GB of video memory, a PyTorch 2.0+ environment, and a pre-loaded Stable Diffusion XL model, only nodes that simultaneously meet all these capability attributes will enter the candidate set.

[0053] S32: Based on the resource status data and task dependencies of each computing power node in the candidate node set, allocate target computing power nodes to each execution task and determine the execution time window for each execution task.

[0054] In this step, real-time resource status data for each node in the candidate node set is first acquired, including dynamic indicators such as currently available GPU memory, CPU utilization, remaining memory, run queue length, and network bandwidth usage. Simultaneously, a directed acyclic graph (DAG) is constructed based on task dependencies to clarify the execution order and data dependencies between tasks. Based on this, a comprehensive decision is made using scheduling optimization algorithms, such as list-based heuristic scheduling, constraint programming, or reinforcement learning models, under the following multiple constraints: a task can only begin after all its predecessor tasks have been completed; a task must be assigned to a candidate node with currently and predicted future resource availability; tasks executed on the same node cannot overlap in time and must satisfy resource continuity. Finally, the algorithm outputs a target computing power node and an execution time window for each task, i.e., the planned start time and the expected end time, thus forming a complete and executable scheduling schedule.

[0055] The above methods enable global efficiency optimization and resource contention coordination for multi-task, multi-node systems in a dynamic resource environment.

[0056] S33: Based on each execution task and its corresponding target computing node, the execution time window of each execution task, and the task dependency relationship, generate a task scheduling strategy for multimodal tasks.

[0057] In this step, after determining the target computing node and specific execution time window for each task, these discrete allocation decisions need to be integrated into a unified and coordinated global plan. Specifically, the task identifier of each task, its assigned target computing node identifier, planned start time, expected end time, and dependencies between tasks are structured into a machine-readable and scheduler-executable task scheduling strategy. This strategy includes at what time, on which computing node, which task, which prerequisite tasks the task depends on, and which subsequent tasks its output will be passed to. It may also include high-level control information such as cross-node data transfer path planning, backup node arrangements for failover, and critical path identification. This complete scheduling strategy, as the overall control blueprint for the entire multimodal task execution, is distributed to the task dispatcher and node executors, thereby systematically driving all subtasks to execute sequentially, on time, and collaboratively on a node basis in the distributed computing environment, ultimately achieving the complex multimodal creation goal.

[0058] S40: Based on the task scheduling strategy, each execution task is distributed to its corresponding target computing power node to generate multimodal works corresponding to multimodal tasks.

[0059] In this step, based on the generated task scheduling strategy, each execution task is assigned to the corresponding target computing power node. Each node loads the required task model and input data, independently executes the assigned task, and relies on the system's internal mechanisms to manage data dependencies and synchronize states between tasks. After completing the execution of all atomic single-modal subtasks, the intermediate results generated on different nodes, belonging to different modalities, are aligned temporally, spatially merged, and format-encapsulated according to the logical dependencies and creative intent defined in the initial task definition, ultimately generating a unified, directly usable multimodal work.

[0060] In a practical application scenario, the system receives a user's request to "generate a short video about spring, including background music and subtitles." The user's requirement is broken down into five tasks: generating descriptive text (T1), generating a series of scene images (T2), synthesizing spoken language (T3), generating synchronized subtitles (T4), and creating background music (T5). The task dependencies are determined: T2, T3, and T4 depend on the text output of T1. Subsequently, based on the memory, computing power, and real-time node status required for each task, T1 is assigned to the text generation node, T2 to the image generation node, T3 and T5 to the audio synthesis node, and T4 to the subtitle processing node. Each computing node executes the tasks in parallel or sequentially, outputting the text, image sequence, audio, subtitle files, and background music respectively. After all subtasks are completed, the audio and background music are mixed into an audio track based on timestamps, the image sequence is synthesized into a video stream, and the subtitles are overlaid onto the video screen. Finally, the video is encoded and packaged into a unified MP4 format short video. This video is a delivered multimodal work that achieves the synchronous integration of text, images, voice, music and subtitles, fully presenting audiovisual content with a spring theme.

[0061] In one embodiment of this application, a specific task scheduling strategy update scheme is provided. That is, after distributing each execution task to its corresponding target computing power node based on the task scheduling strategy, the scheme further includes the following steps: Obtain task execution status data; Based on task execution status data, resource consumption feedback data for each target computing power node is generated; Obtain real-time resource status data for each computing node; The task scheduling strategy is updated based on task execution status data, resource consumption feedback data, and real-time resource status data.

[0062] In this embodiment, task execution status data is continuously acquired during the execution phase, including real-time progress percentages, completed / incomplete status, and error codes for each task. Based on this real-time status data, resource consumption feedback data for each target computing node is generated, including quantitative indicators such as the actual observed memory usage curve, CPU utilization fluctuations, and deviations between actual and estimated task execution times. Simultaneously, the system continuously acquires real-time resource status data for each computing node, reflecting the latest load distribution and resource availability of the cluster at the current moment. Finally, a comprehensive analysis is performed based on task execution status data, resource consumption feedback data, and real-time resource status data. Delayed or failed tasks are identified through execution status, and changes in cluster load are perceived through resource consumption feedback data and real-time resource status. Based on the aforementioned multi-source dynamic information, task scheduling strategies are updated in real-time for tasks that have not yet started execution or are queued, including reallocating computing nodes, adjusting execution sequences, triggering failover, and dynamically balancing node loads, thereby achieving dynamic optimization and adaptive adjustment of the multimodal task execution process.

[0063] As can be seen, in the above scheme, firstly, the user's natural language description is intelligently transformed into multiple execution tasks with clear dependencies through automated parsing and structured decomposition of complex multimodal creative intentions. Then, by constructing a bidirectional precise matching mechanism between task features and node features, dynamic optimization scheduling is performed based on real-time resource status and hardware capabilities to generate a global scheduling strategy. Each execution task is then distributed to its corresponding target computing node to trigger execution, ultimately synthesizing a complete cross-modal work. This effectively avoids computing node overload or resource idleness, significantly improves the overall resource utilization and load balancing level of the heterogeneous computing cluster, enhances the adaptability and system robustness of multimodal generation tasks in dynamic environments, and thus provides an efficient, stable, and scalable automated scheduling strategy for multimodal AI creation tasks. This significantly shortens the execution time of complex multimodal generation tasks and significantly improves computing power utilization and task execution efficiency.

[0064] In one embodiment of this application, a multimodal task scheduling and computing power allocation system for AI creation workflows is provided to implement the aforementioned multimodal task scheduling method. Specifically, the system includes a task access module for receiving multimodal generation task requests from the front-end workflow or user interface; a task classification module for identifying multiple execution tasks and extracting key parameters, including task modality, model name, memory requirements, and execution priority; a computing power resource monitoring module for real-time monitoring of GPU memory usage, task queue length, and temperature load information of each computing power node; a scheduling decision module for executing a multi-level scheduling algorithm based on task feature data and node status data to generate a task scheduling strategy; a task allocation module for distributing each execution task to the target computing power node and managing its execution; and an execution feedback module for collecting execution status and result feedback and updating the task scheduling strategy. The system automatically analyzes task characteristics and dynamically selects the optimal computing power node based on the modality and model load information of the input task. Based on a multi-level scheduling algorithm using weights and priorities, it achieves parallel allocation and load balancing of tasks across multiple GPUs and servers. It also supports task dependency resolution and cross-modal execution pipeline processing. It can be widely used in multimodal creation scenarios such as AI music creation, picture book generation, automated film and television post-production, and digital human synthesis, significantly improving computing power utilization and task execution efficiency.

[0065] Optionally, the scheduling decision module adopts a hierarchical scheduling algorithm, including a global load balancing layer and a local computing power allocation layer. The global layer allocates task weights based on node performance, while the local layer dynamically adjusts the task allocation ratio based on the real-time GPU utilization rate.

[0066] Optionally, the task classification module automatically identifies task modalities based on a deep learning model, and can distinguish different modal types such as text generation, image generation, speech generation, and video rendering.

[0067] Optionally, the computing resource monitoring module periodically collects information on node memory, graphics card model, bandwidth, and temperature, and reports it to the central scheduler via a lightweight agent.

[0068] Optionally, the system supports task dependency resolution, enabling pipelined scheduling of multimodal chained tasks, such as text→image→audio synthesis, to achieve automatic task connection and resource reuse across modalities.

[0069] In one embodiment, a multimodal task scheduling device is provided, which corresponds one-to-one with the multimodal task scheduling method described in the above embodiments. For example... Figure 2 As shown, the multimodal task scheduling device 100 includes: a determination module 101, an acquisition module 102, a first generation module 103, and a second generation module 104. Detailed descriptions of each functional module are as follows: The determination module 101 is used to determine the task feature data of the multimodal task based on the task description information of the multimodal task. The task feature data includes multiple execution tasks corresponding to the multimodal task, the task mode of each execution task, the task requirements of each execution task, the task dependencies between multiple execution tasks, and the resource requirements for executing each execution task. The acquisition module 102 is used to acquire node characteristic data of multiple computing power nodes, wherein the node characteristic data includes resource status data and capability attribute information of each computing power node; The first generation module 103 is used to generate a task scheduling strategy for multimodal tasks based on task feature data and node feature data. The task scheduling strategy includes the target computing power node corresponding to each task and the execution time window of each task. The second generation module 104 is used to distribute each execution task to its corresponding target computing power node based on the task scheduling strategy, and generate multimodal works corresponding to multimodal tasks.

[0070] In one embodiment, the determining module 101 is specifically used for: In response to a multimodal generation task request, obtain the task description information included in the multimodal generation task request; Based on task description information, the workflow of multimodal tasks is decomposed to determine multiple execution tasks, the task requirements of each execution task, and the task dependencies between multiple execution tasks. Based on a pre-defined set of task modes, the target task mode corresponding to each task to be executed is determined. Based on the target task modality and task requirements, determine the target task model corresponding to each execution task; Based on the target task model and task requirements, determine the resource requirements data needed to execute each task.

[0071] In one embodiment, the determining module 101 is further configured to: Based on the large language model, semantic parsing of task description information is performed to determine the multiple execution tasks contained in the multimodal task, as well as the task requirements of each execution task. Based on preset workflow construction rules, the task dependencies between multiple execution tasks are determined.

[0072] In one embodiment, the determining module 101 is further configured to: Based on the target task modality, at least one candidate task model is determined for each execution task. Obtain the attribute information of each candidate task model; Based on task requirements and attribute information, determine the target task model for each task to be performed from at least one candidate task model.

[0073] In one embodiment, the determining module 101 is further configured to: Obtain baseline resource requirements data for each target task model; The baseline resource requirement data is revised based on the task requirements of each task to determine the resource requirement data for each task.

[0074] In one embodiment, the first generation module 103 is specifically used for: The resource requirements data of each task are matched and verified with the capability attribute information of each computing node to determine the candidate node set for each task. Based on the resource status data and task dependencies of each computing node in the candidate node set, target computing nodes are allocated to each execution task, and the execution time window of each execution task is determined. Based on each execution task and its corresponding target computing node, the execution time window of each execution task, and the task dependency relationship, a task scheduling strategy for multimodal tasks is generated.

[0075] In one embodiment, the acquisition module 102 is further configured to acquire task execution status data.

[0076] In one embodiment, the device further includes: The third generation module is used to generate resource consumption feedback data for each target computing power node based on task execution status data.

[0077] In one embodiment, the acquisition module 102 is further configured to acquire real-time resource status data of each computing node; In one embodiment, the device further includes: The update module is used to update the task scheduling strategy based on task execution status data, resource consumption feedback data, and real-time resource status data.

[0078] This invention provides a multimodal task scheduling device 100. First, through automated parsing and structured decomposition of complex multimodal creative intentions, the user's natural language description is intelligently transformed into multiple execution tasks with clear dependencies. Then, by constructing a bidirectional precise matching mechanism between task features and node features, dynamic optimization scheduling is performed based on real-time resource status and hardware capabilities to generate a global scheduling strategy. Each execution task is then distributed to its corresponding target computing node for triggering execution, ultimately synthesizing a complete cross-modal work. This effectively avoids computing node overload or resource idleness, significantly improves the overall resource utilization and load balancing level of heterogeneous computing power clusters, enhances the adaptability and system robustness of multimodal generation tasks in dynamic environments, and thus provides an efficient, stable, and scalable automated scheduling strategy for multimodal AI creation tasks. This significantly shortens the execution time of complex multimodal generation tasks and significantly improves computing power utilization and task execution efficiency.

[0079] Specific limitations regarding the multimodal task scheduling device can be found in the limitations of the multimodal task scheduling method described above, and will not be repeated here. Each module in the aforementioned multimodal task scheduling device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independently of the processor in the electronic device, or stored in software in the memory of the electronic device, so that the processor can call and execute the operations corresponding to each module.

[0080] In one embodiment, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the multimodal task scheduling method described above.

[0081] In one embodiment, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the above-described multimodal task scheduling method.

[0082] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or electronic device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0083] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0084] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0085] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A multimodal task scheduling method, characterized in that, include: Based on the task description information of the multimodal task, the task feature data of the multimodal task is determined. The task feature data includes multiple execution tasks corresponding to the multimodal task, the task modality of each execution task, the task requirements of each execution task, the task dependencies between the multiple execution tasks, and the resource requirements for executing each execution task. Obtain node characteristic data of multiple computing power nodes, wherein the node characteristic data includes resource status data and capability attribute information of each computing power node; Based on the task feature data and the node feature data, a task scheduling strategy for the multimodal task is generated, wherein the task scheduling strategy includes the target computing power node corresponding to each task and the execution time window of each task. Based on the task scheduling strategy, each execution task is distributed to its corresponding target computing power node to generate the multimodal work corresponding to the multimodal task.

2. The multimodal task scheduling method according to claim 1, characterized in that, The step of determining the task feature data of the multimodal task based on the task description information of the multimodal task specifically includes: In response to a multimodal generation task request, the task description information included in the multimodal generation task request is obtained; Based on the task description information, the multimodal task is decomposed into a workflow to determine multiple execution tasks, the task requirements of each execution task, and the task dependencies between the multiple execution tasks. Based on a pre-defined set of task modes, the target task mode corresponding to each task to be executed is determined. Based on the target task modality and the task requirements, determine the target task model corresponding to each execution task; Based on the target task model and the task requirements, the resource requirements data needed to execute each task are determined.

3. The multimodal task scheduling method according to claim 2, characterized in that, The step of decomposing the multimodal task based on the task description information to determine multiple execution tasks, the task requirements of each execution task, and the task dependencies between the multiple execution tasks specifically includes: Based on a large language model, semantic parsing is performed on the task description information to determine the multiple execution tasks included in the multimodal task, as well as the task requirements of each execution task. Based on preset workflow construction rules, the task dependencies between the multiple execution tasks are determined.

4. The multimodal task scheduling method according to claim 2, characterized in that, The step of determining the target task model corresponding to each execution task based on the target task modality and the task requirements specifically includes: Based on the target task modality, at least one candidate task model is determined for each execution task; Obtain the attribute information of each candidate task model; Based on the task requirements and the attribute information, the target task model for each task is determined from the at least one candidate task model.

5. The multimodal task scheduling method according to claim 2, characterized in that, The step of determining the resource requirements data needed to execute each task based on the target task model and the task requirements specifically includes: Obtain baseline resource requirements data for each target task model; The baseline resource requirement data is corrected based on the task requirements of each execution task to determine the resource requirement data for each execution task.

6. The multimodal task scheduling method according to claim 1, characterized in that, The step of generating the task scheduling strategy for the multimodal task based on the task feature data and the node feature data specifically includes: The resource requirement data of each execution task is matched and verified with the capability attribute information of each computing node to determine the candidate node set for each execution task. Based on the resource status data and task dependencies of each computing node in the candidate node set, target computing nodes are allocated to each execution task, and the execution time window for each execution task is determined. Based on each execution task and its corresponding target computing node, the execution time window of each execution task, and the task dependency relationship, the task scheduling strategy for multimodal tasks is generated.

7. The multimodal task scheduling method according to claim 1, characterized in that, After distributing each execution task to its corresponding target computing power node based on the task scheduling strategy, the process further includes: Obtain task execution status data; Based on the task execution status data, resource consumption feedback data for each target computing power node is generated; Obtain real-time resource status data for each computing node; The task scheduling strategy is updated based on the task execution status data, the resource consumption feedback data, and the real-time resource status data.

8. A multimodal task scheduling device, characterized in that, include: The determination module is used to determine the task feature data of the multimodal task based on the task description information of the multimodal task. The task feature data includes multiple execution tasks corresponding to the multimodal task, the task mode of each execution task, the task requirements of each execution task, the task dependencies between the multiple execution tasks, and the resource requirements for executing each execution task. The acquisition module is used to acquire node characteristic data of multiple computing power nodes, wherein the node characteristic data includes resource status data and capability attribute information of each computing power node; The first generation module is used to generate a task scheduling strategy for the multimodal task based on the task feature data and the node feature data, wherein the task scheduling strategy includes the target computing power node corresponding to each execution task and the execution time window of each execution task. The second generation module is used to distribute each execution task to its corresponding target computing power node based on the task scheduling strategy, and generate the multimodal works corresponding to the multimodal tasks.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the multimodal task scheduling method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the multimodal task scheduling method as described in any one of claims 1 to 7.