Media content processing method and apparatus, device, and storage medium

By dynamically scheduling the content generation pipeline and task services through a unified scheduling platform, the problem of redundant development caused by the access of new services is solved, and an efficient and flexible material production process is achieved.

WO2026102663A1PCT designated stage Publication Date: 2026-05-21BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2024-11-14
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

During the material creation and production process, the integration of new services leads to a large amount of repetitive development work, and lacks flexibility and adaptability, making it difficult to efficiently schedule and monitor external services.

Method used

This paper provides a unified scheduling platform that dynamically schedules the services required by content generation pipelines and tasks, enabling different content generation pipelines and tasks to schedule and parse external services, thereby reducing redundant development costs and improving adaptability.

Benefits of technology

By using a unified scheduling platform, the cost of developing repetitive modules is reduced, flexible and efficient adaptation to new services is achieved, and the task scheduling and result parsing processes are simplified.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024132094_21052026_PF_FP_ABST
    Figure CN2024132094_21052026_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a media content processing method and apparatus, a device, and a storage medium. The method comprises: in response to a first task request from a first content generation pipeline, creating a first task of a first type on the basis of the first task request, wherein the first content generation pipeline is configured to generate media content by calling one or more services, the first task request indicates at least the first type and a first call parameter, and the first task is associated with a callable first service (610); in response to the first task being scheduled for execution, on the basis of the first call parameter, calling the first service to execute the first task (620); and in response to obtaining a processing result of the first task from the first service, sending the processing result to the first content generation pipeline for generating media content (630). Therefore, by providing a unified scheduling platform between content generation pipelines and services to be called, the content generation pipelines and the services required by various tasks can be dynamically scheduled.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, apparatus, devices and storage media for media content processing Technical Field

[0001] The exemplary embodiments disclosed herein relate generally to the field of computers, and more particularly to methods, apparatus, devices, and computer-readable storage media for media content processing. Background Technology

[0002] The internet provides access to a wide variety of resources. For example, it allows access to various applications, products, audio and video content, and more. Furthermore, the internet has become a widely used new form of information dissemination for content creation and service promotion. High-quality and efficient content creation processes are becoming increasingly important.

[0003] For example, in terms of content production, high-quality content can increase view rates and click-through rates for content (e.g., advertisements). Furthermore, with the rapid development of automated content production technologies, more and more sophisticated media processing and content generation capabilities are emerging, bringing even more possibilities for high-quality content production.

[0004] Summary of the Invention

[0005] In a first aspect of this disclosure, a method for processing media content is provided. The method includes: in response to a first task request from a first content generation pipeline, creating a first task of a first type based on the first task request, wherein the first content generation pipeline is configured to generate media content by invoking one or more services, and the first task request at least indicates a first type and first invoking parameters, the first task being associated with an invokable first service; in response to the first task being scheduled for execution, invoking the first service based on the first invoking parameters to execute the first task; and in response to receiving a processing result of the first task from the first service, sending the processing result to the content generation pipeline for use in generating media content.

[0006] In a second aspect of this disclosure, an apparatus for media content processing is provided. The apparatus includes: a task creation module configured to create a first task of a first type based on a first task request from a first content generation pipeline, wherein the first content generation pipeline is configured to generate media content by invoking one or more services, and the first task request at least indicates a first type and first invoking parameters, and the first task is associated with an invokable first service; a parameter invoking module configured to invoke a first service to execute the first task based on the first invoking parameters in response to the first task being scheduled for execution; and a result sending module configured to send a processing result of the first task to the content generation pipeline for generating media content in response to receiving a processing result of the first task from the first service.

[0007] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.

[0008] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.

[0009] In a fifth aspect of this disclosure, a computer program product is provided. The computer program product is tangibly stored in a computer storage medium and includes computer-executable instructions that, when executed by a device, cause the device to perform the method of the first aspect.

[0010] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0011] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0012] Figure 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;

[0013] Figure 2 illustrates a schematic diagram of an example architecture for media content processing according to some embodiments of the present disclosure;

[0014] Figure 3 shows a schematic diagram of an example structure of a predetermined parameter structure according to some embodiments of the present disclosure;

[0015] Figure 4 illustrates a schematic diagram of an example architecture for the state of media content processing according to some embodiments of the present disclosure;

[0016] Figure 5 shows a flowchart of an example process for scheduling and executing a task according to some embodiments of the present disclosure;

[0017] Figure 6 shows a flowchart of a process for media content processing according to some embodiments of the present disclosure;

[0018] Figure 7 shows a block diagram of an apparatus for media content processing according to some embodiments of the present disclosure; and

[0019] Figure 8 shows a block diagram of an electronic device capable of implementing several embodiments of the present disclosure. Detailed Implementation

[0020] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0021] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.

[0022] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0023] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.

[0024] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.

[0025] As used in this paper, the term "model" refers to a model that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs using multiple layers of processing units. A neural network model is an example of a deep learning-based model. In this paper, "model" may also be referred to as a "machine learning model," "learning model," "machine learning network," or "learning network," and these terms are used interchangeably.

[0026] Machine learning typically comprises three phases: training, testing, and application (also known as inference). In the training phase, a given model is trained using a large amount of training data, iteratively updating parameter values ​​until the model can consistently generate inferences that meet the expected goals from the training data. Through training, the model can be considered to have learned the relationship between inputs and outputs (also known as an input-output mapping) from the training data. The parameter values ​​of the trained model are determined. In the testing phase, test inputs are applied to the trained model to test whether it can provide the correct output, thus determining the model's performance. The testing phase can sometimes be integrated into the training phase. In the application or inference phase, the trained model can be used to process actual model inputs based on the trained parameter values ​​to determine the corresponding model output.

[0027] As briefly mentioned earlier, with the rapid development of automated content generation technology, more and more mature media processing and content generation capabilities are emerging, bringing more possibilities to the production of high-quality materials. For example, in the process of producing materials, external services can be called to perform various processing on the content, thereby combining them to achieve the desired final effect. However, in the process of material production, the integration of new processing capabilities and services has become one of the most manpower-intensive parts of the R&D process. In addition to connecting and debugging the input parameters of new processing capabilities, it is also necessary to spend a considerable amount of time and effort developing the scheduling logic and result receiving logic for these new capabilities. Furthermore, the capacity of the external services to be scheduled must be considered, and corresponding rate limiting strategies must be developed. When multiple new processing capabilities need to be integrated, the above repetitive work will occupy a significant amount of the R&D personnel's energy.

[0028] For example, the process of generating content for a certain application (e.g., a novel application) is as follows: Based on the text information input by the user (e.g., novel excerpt), the terminal device first needs to call a first service (e.g., a storyboard model capable of performing storyboarding tasks on the excerpt) to obtain the storyboard excerpt. Next, the terminal device continues to receive the input storyboard excerpt and calls a second task (e.g., a text-to-image model capable of performing text-to-image tasks) to generate the images required for each storyboard. Then, the terminal device continues to receive the input images and calls a third task (e.g., a video compositing model capable of performing video generation tasks) to generate the final video.

[0029] In the above process, each task needs to interface with downstream processing capabilities, and the service call resource addresses, capacities, return formats (synchronous, asynchronous), and output result formats for each task are different. Therefore, if each content production pipeline and each task within that pipeline is developed separately, it will lead to repetitive work and insufficient flexibility and adaptability. Consequently, without a universal task interface adaptation layer, it is also not conducive to monitoring and maintaining downstream tasks.

[0030] This disclosure presents a scheme for media content processing. According to various embodiments of this disclosure, if a first task request is received from a content generation pipeline, a first task is created based on the first task request. The content generation pipeline is configured to generate target media content, and the first task request at least indicates a first task type and first invocation parameters. The first task is associated with a callable first service. If it is determined that the first task is scheduled for execution, the first service is invoked according to the first invocation parameters to execute the first task. Then, the processing result of the first task is received from the first service and sent to the content generation pipeline for generating media content. By providing a unified scheduling platform between the content generation pipeline and the services to be invoked, the services required by various content generation pipelines and tasks are dynamically scheduled. This enables the reuse of processes such as scheduling and parsing of external services by different content generation pipelines and tasks, reducing the cost of developing a large number of repetitive modules, and also allows for flexible and efficient adaptation to new services. For example, if a new content generation pipeline is developed or new tasks need to be introduced, the service scheduling capabilities of the scheduling platform can be directly reused without repeated development.

[0031] The following section provides a detailed description of various example implementations of this scheme, with reference to the accompanying drawings.

[0032] Example Environment

[0033] Figure 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In this example environment 100, an electronic device 120 deploys a content generation pipeline 140 and a scheduling framework 110. The content generation pipeline 140 includes multiple tasks 142-1, 142-2, ..., 142-N, which can indicate the processing power required in the process of generating media content. For ease of discussion, tasks 142-1, 142-2, ..., 142-M can be collectively referred to as task 142. Depending on the specific requirements for content generation, the content generation pipeline 140 can be configured to include different tasks. Although multiple tasks are shown sequentially in the figure, two or more tasks can also be executed in parallel, depending on the specific application requirements.

[0034] In environment 100, multiple tasks 142-1, 142-2, ..., 142-N have corresponding services 130-1, 130-2, ..., 130-N. For ease of discussion, services 130-1, 130-2, ..., 130-N can be collectively referred to as service 130. Electronic device 120 can invoke scheduling framework 110 and / or attached devices of scheduling framework 110 to generate media content. That is, when scheduling framework 110 deployed in electronic device 120 receives a task request from content generation pipeline 140, it can invoke the service corresponding to that task to execute the task. In some examples, different tasks can also be implemented by the same service. In some scenarios, user 110 can also publish the generated media content through electronic device 120. "Media content" can include one or more types of content, such as video, images, GIFs, image sets, audio, text, etc.

[0035] Although shown as integrated within the same device, the content generation pipeline 140 and scheduling framework 110 can also be implemented in separate devices. In some embodiments, the scheduling framework 110 deployed in electronic device 120 can communicate with service 130 to generate media content based on content generation pipeline 140. Electronic device 120 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, electronic device 120 can also support any type of user-facing interface (such as "wearable" circuitry). Service 130 can be various types of computing systems / servers capable of providing computing power, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, and so on.

[0036] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure. For example, embodiments of this disclosure can be applied to any suitable one or more applications, and are not limited to office suites.

[0037] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.

[0038] Example Interaction

[0039] The process of media content processing according to this disclosure is described below with reference to FIG2. FIG2 shows a schematic diagram of an example frame 200 for media content processing according to some embodiments of this disclosure. Hereinafter, the example embodiments will be described primarily with respect to scheduling frame 110. It should be understood that the actions described with respect to scheduling frame 110 can be performed by electronic device 120, or can be performed by electronic device 120 in conjunction with its server (e.g., server).

[0040] As shown in the example architecture 200 in Figure 2, in cases where media content needs to be produced based on multiple tasks, the electronic device 120 can invoke the scheduling framework 110 to produce the media content. That is, the scheduling framework 110 receives a task request from task 142-1 among multiple tasks in the first content generation pipeline. Subsequently, based on the task request of task 142-1, the scheduling framework 110 invokes the application programming interface (API) 212 to create task 142-1.

[0041] In some examples, a task request can indicate the type of task. For instance, for a workflow in a novel application that generates video based on text, task 142-1 could be a storyboard task, task 142-2 could be a text-to-image task, and task 142-3 could be a video compositing task. A task request can also indicate the parameters for calling the task. For example, for generating video based on text, the parameters could indicate the text.

[0042] In some embodiments, for each type, the following can be configured: a limit on the number of parallel processing tasks corresponding to each type, a parameter structure configuration for each type of task, and / or a result parsing configuration for each type of task. In some examples, for each type, a flow control configuration 216, a parameter structure configuration 217, and a result parsing configuration 218 can be provided in advance.

[0043] Flow control configuration 216 instructs users (who can be referred to as developers) to pre-set the upper limit of tasks of a certain type that can be processed in parallel. Through flow control configuration, the amount of downstream calls can be controlled, preventing a large number of tasks from failing or timed out when the capacity of downstream services is insufficient.

[0044] Parameter structure configuration 217 instructs users (who can be referred to as developers) to pre-configure parameter structures for a specific task type, so that users only need to construct the parameters during subsequent use. The user's parameter construction will be described in detail below with reference to Figure 3. In this way, users only need to configure the parameter construction logic for different types of tasks, without needing to worry about the task scheduling logic.

[0045] Result parsing configuration 218 indicates that users (who can be referred to as developers) can pre-configure result parsing functionality for a specific task type. By configuring result parsing for a particular task type, the scheduling framework can parse the results scheduled from the service. This eliminates the need to configure result receiving and parsing logic in the content generation pipeline.

[0046] In this way, by registering the configurations corresponding to different types of tasks in the scheduling framework, redundant development can be avoided. Different content generation pipelines or different stages in the same content generation pipeline can complete the scheduling of the required service capabilities by registering the task configuration information and taking advantage of the scheduling mechanism and flow control mechanism of the scheduling framework.

[0047] Figure 3 illustrates a schematic diagram of an example structure 300 of a predetermined parameter structure according to some embodiments of the present disclosure. As shown in the example structure 300 in Figure 3, the parameter structure configuration corresponding to each type of task may include: Configuration 311: primarily storing the concurrency configurations for different types of tasks (I Task implementation class), including global concurrency and single-instance concurrency. This configuration can be updated in real time through a dynamic configuration center. Task Scheduler 312: responsible for task scheduling, containing the concurrency configuration and a coroutine pool for batch retrieval and processing of task data to be scheduled.

[0048] Core Interface (I Task) 313: The interface for the scheduling framework (i.e., the concrete application interface of the strategy pattern). Understandably, when new processing capabilities need to be integrated, users (e.g., developers) only need to implement the methods in this interface, without needing to concern themselves with scheduling, processes, etc. The main functions of the methods in this interface include: Parameter Check: Parameter check for each task. Submit: Responsible for constructing task parameters and invoking the service. Process Callback: Responsible for parsing the service's callback structure. Polling Result: Responsible for querying the execution status of downstream services when they do not support callbacks. Get Task Info. Task Type: Returns the current task type.

[0049] Implementation class 315: Indicates the parameters used to provide the corresponding construction tasks. For example, if generating a video based on text requires a storyboard task, a text-to-image task, and a video compositing task, the user can add a class 317 for the storyboard task, a class 318 for the text-to-image task, and a class 319 for the video compositing task in implementation class 315. In some examples, all three classes implement the `ITask` interface and implement their specific logic in their methods. Task information (`Task Info`) 316: Indicates the task entity, mainly recording the task type, parameters, results, and task status.

[0050] If the scheduling framework 110 receives a first task request from the first content generation pipeline, it creates a first task of a first type based on the first task request. Here, the first content generation pipeline is configured to generate media content by invoking one or more services. The first task request from the content generation pipeline at least indicates a first type and first invocation parameters, and the requested first task is associated with a callable first service.

[0051] After a task is created, the scheduling framework 110 can schedule the execution of each task according to the scheduling policy. If it is determined that the first task is scheduled for execution, the scheduling framework 110 calls the first service to execute the first task based on the first invocation parameter. As shown in Figure 2, if the scheduling framework 110 determines that task 142-1 is scheduled for execution, it calls service 130-1 to execute task 142-1 based on the first invocation parameter (e.g., the novel text used to generate the video).

[0052] In some embodiments, if the scheduling framework 110 determines that the first invocation parameter indicated by the first task request conforms to the parameter structure requirements corresponding to the first type, it stores the task information of the first task in the task queue. Then, according to the task scheduling policy, the scheduling framework 110 determines the first task to be scheduled for execution from the task queue, and the first task is in a pending execution state.

[0053] As shown in Figure 2, after creating task 142-1, the scheduling framework 110 uses the parameter verification module 213 to determine whether the calling parameters corresponding to task 142-1 meet the parameter structure requirements of the first type, and then stores the task information of task 142-1 in the task queue 223. For example, a storyboard task needs to verify the input keywords, a video compositing task needs to verify the input video resources, and so on. Subsequently, the scheduling framework 110 sequentially determines the task 142-1 in the pending execution state from the task queue 223, and schedules the service corresponding to task 142-1 to execute task 142-1.

[0054] In some embodiments, the task information of the first task may include: a field indicating the first type of the first task, a field indicating the first call parameter of the first task, a field indicating the processing result of the first task, and / or a field indicating the status of the first task. As shown in the example architecture 300 of Figure 3, the task information (Task Info) of task 142-1 records the type field (Task Type), the parameter field (Task Param), the result field (Result), the status field (Task Status), and the identifier (Downstream ID) corresponding to task 142-1. The identifier (Downstream ID) corresponding to task 142-1 is used to record the unique identifier returned by the downstream when task 142-1 calls the downstream.

[0055] The following description continues with reference to Figures 2 and 5. After storing the task information of task 142-1, the scheduling framework 110 uses its included flow control verification module 219 to determine whether to schedule and execute task 142-1 based on the pre-configured upper limit of parallel processing tasks.

[0056] In some embodiments, scheduling framework 110 determines whether the number of tasks belonging to the first type and in a running state exceeds the upper limit of parallel processing tasks corresponding to the first type. If the scheduling framework determines that the number of tasks exceeds the upper limit of parallel processing tasks, it determines that the first task should be scheduled for execution. In some examples, scheduling framework 110 includes a coroutine pool for recording the number of tasks of the same type running at a target time. If scheduling framework 110 determines that the number of tasks of the first type in the coroutine pool does not exceed the upper limit of parallel processing tasks, it can schedule task 142-1 for execution.

[0057] In some embodiments, if the scheduling framework 110 determines that a first task has been created, it locks the first task using a distributed lock corresponding to the first type. Subsequently, if the scheduling framework 110 determines that the number of tasks does not exceed the limit for parallel processing tasks and the first task is in a running state, it releases the distributed lock on the first task to determine that the first task should be scheduled for execution. In some examples, the scheduling framework 110 may use a distributed lock to lock task 142-1 among at least one task in the task queue that is in a pending state. Then, when the scheduling framework 110 determines that the number of tasks does not exceed the limit for parallel processing tasks, it can set the state of task 142-1 to a running state. Subsequently, the scheduling framework 110 releases the distributed lock on task 142-1 to invoke service 130-1 to execute task 142-1.

[0058] The process of scheduling and executing the first task will be described below with reference to Figure 5. Figure 5 shows a flowchart of an example process 500 for scheduling and executing a task according to some embodiments of this disclosure.

[0059] In process 500, the Task Scheduler creates a pool of coroutines at runtime, where the execution dimension of each coroutine is a type. In block 511, the scheduling framework 110 queries the number of tasks currently in the running state for the current type. In block 512, the scheduling framework 110 calculates the number of runnable tasks based on the concurrency configuration. In block 513, the scheduling framework 110 extracts task 142-1, which is in the pending running state, from task queue 223, and then locks task 142-1 using a distributed lock (e.g., which can be implemented based on Redis). In some embodiments, the lock granularity is type. This approach prevents multiple coroutines of the same type from being scheduled simultaneously, thus preventing the rate limiting configuration from failing.

[0060] In box 514, if scheduling framework 110 determines that the number of tasks in the running state of the current type does not exceed the limit of parallel processing tasks, it changes the state of task 142-1, which is in the waiting state, to the running state and releases the lock. In box 515, scheduling framework 110 calls service 130-1 to run task 142-1. In box 516, scheduling framework 110 receives the result of service 130-1 running task 142-1 and updates the state of task 142-1 in task queue 223.

[0061] The following description, referring to Figures 2 and 3, describes how the scheduling framework 110 determines the scheduled execution task 142-1, and then uses its included task scheduling module 220 to construct the calling parameters for calling service 130-1 based on the pre-configured parameter structure, and to call service 130-1 based on service addressing.

[0062] In some embodiments, the scheduling framework 110 can construct service invocation parameters for calling the first service based on the parameter structure configuration corresponding to the first type. Then, the scheduling framework 110 uses the service invocation parameters to call the first service via the address information of the first service to execute the first task. As shown in Figure 3, for the need to access a new processing capability (for example, the processing capability belongs to type A), a class for implementing the processing capability can be added in the implementation class 314, and the configuration of type A can be configured in the configuration (Config) 311. Subsequently, the scheduling framework 110 calls the service to execute the task based on the address information of the service.

[0063] The following description, referring to Figures 2 and 4, continues to describe how the scheduling framework 110 determines, after the service 130-1 completes the execution of task 142-1, to receive the processing result of task 142-1 from service 130-1 using the result receiving module 221. Figure 4 shows a schematic diagram of an example architecture 400 for the state during media content processing according to some embodiments of this disclosure.

[0064] In some embodiments, the scheduling framework 110 receives the processing result of the first task from the first service and sends the processing result to the first content generation pipeline for generating media content. In some examples, service 130-1 can send a callback of the result after executing task 142-1 to the scheduling framework 110. In some embodiments, the scheduling framework 110 can also poll the result corresponding to task 142-1. Subsequently, the scheduling framework 110 returns the result corresponding to task 142-1 to the first content generation pipeline to generate media content. In some embodiments, the content pipeline can also query the result corresponding to task 142-1.

[0065] In some embodiments, after obtaining the processing result of the first task from the first service, the scheduling framework 110 can parse the processing result based on the result parsing configuration corresponding to the first type. Subsequently, the scheduling framework 110 sends the successfully parsed processing result to the first content generation pipeline for generating media content. In some examples, the scheduling framework 110 obtains the processing result of task 142-1 called back by service 130-1 via the result receiving module 221. The scheduling framework 110 parses the processing result of task 142-1 called back by service 130-1 via the result parsing function in the result receiving module 221. Further, the scheduling framework 110 sends the successfully parsed processing result to the content generation pipeline for generating media content. Accordingly, the scheduling framework 110 updates the status information of task 142-1 in the task queue 223 in real time.

[0066] In some embodiments, if the scheduling framework 110 determines that the parsing result of the processing result indicates parsing failure or processing failure of the first service, it updates the status of the first task to a first status. In some embodiments, the first status indicates that the processing result corresponding to the first task failed to be generated. In some embodiments, if the scheduling framework 110 determines that the status of the first task is the first status, it re-invokes the first service based on the first invocation parameters to execute the first task.

[0067] In some examples, if the scheduling framework 110 determines, via the result parsing function in the result receiving module 221, that the parsing result of the processing result indicates parsing failure or processing failure of the first service, then it updates the status information of task 142-1 to a first status indicating that the processing result corresponding to task 142-1 has failed to be generated. Then, the scheduling framework 110 uses the first invocation parameters to re-invoke the first service to execute the first task. Accordingly, the scheduling framework 110 updates task 142-1, which is in a failed state, to a pending execution state.

[0068] As shown in the example architecture 400 in Figure 4, after task 142-1 is created, its status is set to pending execution status 411. After task 142-1 is submitted and passes verification, its status is set to running status 412. After task 142-1 is scheduled for execution, it is set to success status 413. In some embodiments, if the scheduling framework 110 determines, via the result parsing function in the result receiving module 221, that the parsing result indicates parsing failure or processing failure of the first service, then the status of task 142-1 is updated to failure status 414. A task in failure status 414 can be invoked by a predetermined mechanism, and its status is set back to pending execution status 411 for rescheduling and execution.

[0069] In some embodiments, if the scheduling framework 110 determines that the parsing result of the processing result indicates parsing failure or that the number of processing failures of the first service exceeds a predetermined number, it updates the first task in the first state to the third state. In some embodiments, the third state indicates that the media content generation corresponding to the first task has ended. It is understood that if the number of times the task in the failure state 414 is invoked by the predetermined mechanism exceeds a predetermined number, the process of generating the media content corresponding to that task is directly terminated.

[0070] In some embodiments, the scheduling framework can interface with multiple content generation pipelines or schedule multiple tasks within the same content generation pipeline. For example, if a second task request is received from a first or second content generation pipeline, the scheduling framework 110 can create a second task of a second type based on the second task request. The second task request at least indicates a second type and second invocation parameters, and the second task is associated with an invoked second service. If the scheduling framework 110 determines that the second task is scheduled for execution, it invokes the second service based on the second invocation parameters to execute the second task. Then, if the scheduling framework 110 obtains the processing result of the second task from the second service, it sends the processing result to the first or second content generation pipeline for generating media content.

[0071] As shown in the example architecture 200 in Figure 2, the scheduling framework 110 receives a task request for task 142-2 from among multiple tasks in the first content generation pipeline. Subsequently, the scheduling framework 110, based on the task request for task 142-2, invokes the application programming interface (API) 212 to create task 142-2. In some embodiments, if the scheduling framework 110 determines that task 142-2 is scheduled for execution, it invokes service 130-2 to execute task 142-2 based on a second invocation parameter (e.g., a novel's text for generating a video).

[0072] Then, the scheduling framework 110 obtains the processing result of task 142-2 called back by service 130-2 via the result receiving module 221. The scheduling framework 110 parses the processing result of task 142-2 called back by service 130-2 via the result parsing function in the result receiving module 221. Further, the scheduling framework 110 sends the successfully parsed processing result to the content generation pipeline for generating media content.

[0073] In summary, the embodiments disclosed herein enable the abstraction and extraction of common processes from different services, as well as the establishment of an asynchronous scheduling mechanism. This allows upstream services to call the same interface, introducing different task capabilities through parameter control. Consequently, when accessing new processing capabilities, users only need to focus on parameter concatenation and result parsing when calling the service. They do not need to concern themselves with task scheduling and downstream capacity, thus saving tedious repetitive work. Furthermore, based on the common flow control mechanism, different flow control measures can be implemented for each task type simply by modifying the configuration.

[0074] Example process

[0075] Figure 6 shows a flowchart of an example process 600 for media content processing according to some embodiments of the present disclosure. Process 600 can be implemented at the scheduling framework 110. Process 600 will now be described with reference to Figure 1.

[0076] As shown in Figure 6, in box 610, scheduling framework 110 responds to a first task request from a first content generation pipeline, and creates a first task of a first type based on the first task request, wherein the first content generation pipeline is configured to generate media content by invoking one or more services, and the first task request at least indicates a first type and a first invocation parameter, and the first task is associated with a callable first service.

[0077] In box 620, the scheduling framework 110 responds to the first task being scheduled for execution by invoking the first service based on the first invocation parameters to execute the first task.

[0078] In box 630, in response to receiving the processing result of the first task from the first service, the scheduling framework 110 sends the processing result to the content generation pipeline for generating media content.

[0079] In some embodiments, process 600 further includes: in response to the first call parameter indicated by the first task request conforming to the parameter structure requirements corresponding to the first type, storing the task information of the first task in a task queue; and determining from the task queue, based on a task scheduling strategy, that the first task should be scheduled for execution.

[0080] In some embodiments, the task information of the first task includes at least one of the following: a field indicating a first type of the first task; a field indicating a first invocation parameter of the first task; a field indicating the processing result of the first task; and / or a field indicating the status of the first task.

[0081] In some embodiments, determining the first task to be scheduled for execution from the task queue includes: determining whether the number of tasks belonging to the first type and in a running state exceeds the upper limit of parallel processing tasks corresponding to the first type; and determining the first task to be scheduled for execution in response to the number of tasks not exceeding the upper limit of parallel processing tasks.

[0082] In some embodiments, determining from the task queue that the first task is to be scheduled for execution further includes: locking the first task using a distributed lock corresponding to the first type in response to the creation of the first task; and releasing the distributed lock on the first task in response to the number of tasks not exceeding the upper limit of parallel processing tasks and the state of the first task being running, so as to determine that the first task is to be scheduled for execution.

[0083] In some embodiments, invoking a first service to perform a first task based on a first invocation parameter includes: constructing service invocation parameters for invoking the first service based on a parameter structure configuration corresponding to a first type; and invoking the first service to perform the first task using the service invocation parameters via the address information of the first service.

[0084] In some embodiments, sending the processing result to the second content generation pipeline for generating media content includes: in response to obtaining the processing result of the first task from the first service, parsing the processing result based on the result parsing configuration corresponding to the first type; and sending the successfully parsed processing result to the first content generation pipeline for generating media content.

[0085] In some embodiments, process 600 further includes: in response to the parsing result of the processing result indicating parsing failure or the processing of the first service failing, updating the state of the first task to a first state, the first state indicating that the processing result corresponding to the first task failed to be generated; in response to determining that the state of the first task is the first state, re-invoking the first service based on the first invocation parameter to execute the first task.

[0086] In some embodiments, process 600 further includes: in response to a second task request from a first content generation pipeline or a second content generation pipeline, creating a second task of a second type based on the second task request, the second task request indicating at least a second type and second invocation parameters, the second task being associated with an invokeable second service; in response to the second task being scheduled for execution, invoking the second service based on the second invocation parameters to execute the second task; and in response to obtaining a processing result of the second task from the second service, sending the processing result to the first content generation pipeline or the second content generation pipeline for generating media content.

[0087] Example devices and equipment

[0088] Embodiments of this disclosure also provide corresponding apparatus for implementing the methods or processes described above. Figure 7 shows a schematic structural block diagram of an example apparatus 700 for media content processing according to certain embodiments of this disclosure. Apparatus 700 may be implemented as or included in scheduling framework 110. The various modules / components in apparatus 700 may be implemented by hardware, software, firmware, or any combination thereof.

[0089] As shown in Figure 7, the apparatus 700 includes a task creation module 710 configured to create a first task of a first type based on a first task request from a first content generation pipeline, wherein the first content generation pipeline is configured to generate media content by invoking one or more services, and the first task request at least indicates a first type and first invocation parameters, and the first task is associated with an invokable first service. The apparatus 700 also includes a parameter invocation module 720 configured to invoke a first service to execute the first task based on the first invocation parameters in response to the first task being scheduled for execution. The apparatus 700 further includes a result sending module 730 configured to send the processing result of the first task to the content generation pipeline for generating media content in response to receiving the processing result from the first service.

[0090] In some embodiments, the apparatus 700 further includes a scheduling execution first task determination module, configured to store the task information of the first task in a task queue in response to the first call parameter indicated by the first task request conforming to the parameter structure requirements corresponding to the first type; and to determine from the task queue that the first task should be scheduled for execution based on a task scheduling strategy.

[0091] In some embodiments, the task information of the first task includes at least one of the following: a field indicating a first type of the first task; a field indicating a first invocation parameter of the first task; a field indicating the processing result of the first task; and / or a field indicating the status of the first task.

[0092] In some embodiments, the scheduling execution first task determination module is further configured to determine whether the number of tasks belonging to the first type and in the running state exceeds the upper limit of parallel processing tasks corresponding to the first type; and in response to the number of tasks not exceeding the upper limit of parallel processing tasks, determine that the first task should be scheduled for execution.

[0093] In some embodiments, the scheduling execution first task determination module is further configured to lock the first task using a distributed lock corresponding to the first type in response to the creation of the first task; and to release the distributed lock on the first task in response to the number of tasks not exceeding the upper limit of parallel processing tasks and the state of the first task being running, so as to determine that the first task is to be scheduled for execution.

[0094] In some embodiments, the parameter invocation module 720 is further configured to construct service invocation parameters for invoking the first service based on the parameter structure configuration corresponding to the first type; and to invoke the first service via the address information of the first service using the service invocation parameters to perform the first task.

[0095] In some embodiments, the result sending module 730 is further configured to, in response to obtaining the processing result of the first task from the first service, parse the processing result based on the result parsing configuration corresponding to the first type; and send the successfully parsed processing result to the first content generation pipeline for generating media content.

[0096] In some embodiments, the parameter calling module 720 is further configured to update the status of the first task to a first status in response to the parsing result of the processing result indicating parsing failure or the processing of the first service failing, wherein the first status indicates that the processing result corresponding to the first task has failed to be generated; and in response to determining that the status of the first task is the first status, to re-call the first service to execute the first task based on the first calling parameters.

[0097] In some embodiments, the result sending module 730 is further configured to: in response to a second task request from a first content generation pipeline or a second content generation pipeline, create a second task of a second type based on the second task request, wherein the second task request at least indicates a second type and second invocation parameters, and the second task is associated with an invokeable second service; in response to the second task being scheduled for execution, invoke the second service based on the second invocation parameters to execute the second task; and in response to obtaining the processing result of the second task from the second service, send the processing result to the first content generation pipeline or the second content generation pipeline for generating media content.

[0098] Figure 8 shows a block diagram of an electronic device 800 capable of implementing various embodiments of the present disclosure. The electronic device 800 may be implemented as or include the electronic device 120 of FIG1 or the device 700 of FIG7.

[0099] As shown in Figure 8, the electronic device 800 is in the form of a general-purpose computing device. Components of the electronic device 800 may include, but are not limited to, one or more processors or processing units 810, memory 820, storage devices 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860. The processing unit 810 may be a physical or virtual processor and is capable of performing various processes according to programs stored in the memory 820. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 800.

[0100] Electronic device 800 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 800, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 820 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 830 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 800.

[0101] Electronic device 800 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG8, disk drives for reading or writing from removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading or writing from removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. Memory 820 may include computer program product 825 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.

[0102] The communication unit 840 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 800 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 800 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0103] Input device 850 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 860 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 800 can also communicate with one or more external devices (not shown) via communication unit 840 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 800, or with any device that enables electronic device 800 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0104] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0105] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0106] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0107] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0108] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0109] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A media content processing method, comprising: In response to a first task request from a first content generation pipeline, a first task of a first type is created based on the first task request, wherein the first content generation pipeline is configured to generate media content by invoking one or more services, and the first task request at least indicates the first type and a first invoking parameter, and the first task is associated with a callable first service; In response to the first task being scheduled for execution, the first service is invoked based on the first invocation parameters to execute the first task; as well as In response to receiving the processing result of the first task from the first service, the processing result is sent to the content generation pipeline for generating the media content.

2. The method according to claim 1, further comprising: In response to the first call parameter indicated by the first task request conforming to the parameter structure requirements corresponding to the first type, the task information of the first task is stored in the task queue; as well as Based on the task scheduling strategy, the first task to be scheduled for execution is determined from the task queue.

3. The method according to claim 2, wherein the task information of the first task includes at least one of the following: A field used to indicate the first type of the first task; A field used to indicate the first invocation parameter of the first task; Fields used to indicate the processing result of the first task; and / or A field used to indicate the status of the first task.

4. The method of claim 2, wherein determining the first task to be scheduled for execution from the task queue comprises: Determine whether the number of tasks belonging to the first type and in a running state exceeds the upper limit of parallel processing tasks corresponding to the first type; as well as In response to the fact that the number of tasks does not exceed the upper limit of the parallel processing tasks, it is determined that the first task should be scheduled for execution.

5. The method of claim 4, wherein determining from the task queue that the first task is to be scheduled for execution further comprises: In response to the creation of the first task, the first task is locked using the distributed lock corresponding to the first type; as well as In response to the fact that the number of tasks does not exceed the upper limit of the parallel processing tasks and the first task is in a running state, the distributed lock on the first task is released to determine that the first task is to be scheduled for execution.

6. The method of claim 1, wherein invoking the first service based on the first invocation parameter to perform the first task comprises: Based on the parameter structure configuration corresponding to the first type, construct service call parameters for calling the first service; as well as Using the service call parameters, the first service is invoked via the address information of the first service to execute the first task.

7. The method of claim 1, wherein sending the processing result to the second content generation pipeline for generating the media content comprises: In response to obtaining the processing result of the first task from the first service, the processing result is parsed based on the result parsing configuration corresponding to the first type; as well as The processing result after successful parsing is sent to the first content generation pipeline to generate the media content.

8. The method according to claim 7, further comprising: In response to the parsing result of the processing result indicating parsing failure or the processing of the first service failing, the status of the first task is updated to a first status, the first status indicating that the processing result corresponding to the first task failed to be generated; as well as In response to determining that the state of the first task is the first state, the first service is re-invoked based on the first invocation parameters to execute the first task.

9. The method according to claim 1, further comprising: In response to a second task request from the first content generation pipeline or the second content generation pipeline, a second task of a second type is created based on the second task request, the second task request indicating at least the second type and second invocation parameters, and the second task is associated with a callable second service; In response to the second task being scheduled for execution, the second service is invoked based on the second invocation parameters to execute the second task; as well as In response to obtaining the processing result of the second task from the second service, the processing result is sent to the first content generation pipeline or the second content generation pipeline for generating the media content.

10. An apparatus for media content processing, comprising: A task creation module is configured to respond to a first task request from a first content generation pipeline, and to create a first task of a first type based on the first task request, wherein the first content generation pipeline is configured to generate media content by invoking one or more services, and the first task request at least indicates a first type and a first invocation parameter, and the first task is associated with a callable first service; The parameter invocation module is configured to invoke the first service to execute the first task based on the first invocation parameters in response to the first task being scheduled for execution. as well as The result sending module is configured to, in response to receiving the processing result of the first task from the first service, send the processing result to the content generation pipeline for generating the media content.

11. An electronic device, comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 9 when executed by the at least one processing unit.

12. A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the method according to any one of claims 1 to 9.

13. A computer program product tangibly stored in a computer storage medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method according to any one of claims 1 to 9.