Media data processing method and system, and cluster, storage medium and program product
By allowing users to configure AI models for media task processing, and combining cloud platform presets and user-defined AI models, the problem of insufficient applicability of AI technology in media transcoding services is solved, thereby improving the reliability and efficiency of media processing.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
- Filing Date
- 2025-04-30
- Publication Date
- 2026-05-07
AI Technical Summary
In cloud-based media transcoding services, existing AI technologies have limited applicability and do not support modification, resulting in low reliability of media transcoding results.
Users can configure AI models for media task processing, supporting the flexibility and applicability of media processing templates. By using cloud platform presets and user-defined AI models, the adaptability is improved and the media processing process is simplified.
It improves the reliability and efficiency of media processing results, reduces the need to build end-to-end business systems, and enhances the flexibility and applicability of media processing templates.
Smart Images

Figure CN2025092518_07052026_PF_FP_ABST
Abstract
Description
Media data processing methods, systems, clusters, storage media, and program products
[0001] This application claims priority to Chinese patent application filed on October 30, 2024, with application number 202411533550.X and entitled "Method, System, Cluster, Storage Medium and Program Product for Processing Media Data", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of cloud computing technology, and more specifically to a method, system, cluster, storage medium, and program product for processing media data. Background Technology
[0003] Media transcoding refers to converting media data from one encoding format to another. With the development of cloud computing technology, media transcoding services can be built on cloud services, saving users the high costs of purchasing, building, and managing transcoding hardware and software, and avoiding issues such as configuration optimization and transcoding parameter adaptation. At the same time, media transcoding services can leverage the elastic scalability of cloud services to meet the needs of actual transcoding business.
[0004] In cloud-based media transcoding services, artificial intelligence (AI) technology can be used to process media data. However, the AI technologies provided by different cloud services are applicable to limited scenarios, and the AI technologies integrated into media transcoding services cannot be modified. Therefore, when the media processing scenario is incompatible with the AI technologies integrated into the media transcoding service, the reliability of the media transcoding results is low. Summary of the Invention
[0005] This application provides a method, system, cluster, storage medium, and program product for processing media data, in order to improve the flexibility of AI technology integrated in media transcoding services and enhance the adaptability between AI technology integrated in media processing scenarios and media transcoding services.
[0006] Firstly, this application provides a media data processing method applied to a cloud management platform. During media data processing, the client sends an algorithm configuration request to the cloud management platform. This algorithm configuration request includes transcoding parameters and one or more first identifiers used to indicate an AI model. Based on the algorithm configuration request sent by the client, the cloud management platform obtains a media processing template. The client then sends a first request to the cloud management platform for processing first media data. Upon receiving the first request from the client, the cloud management platform processes the first request according to the media processing template to obtain second media data.
[0007] Compared to media transcoding services integrated with AI technology provided by cloud service providers, the method presented in the first aspect allows users to configure the AI models during media task processing themselves. The media processing templates are modifiable, thus offering greater flexibility and applicability. Furthermore, the AI models available to users include pre-set AI models from the cloud platform and user-defined AI models. This increased availability of AI models enhances the compatibility between media processing templates and media processing scenarios, thereby ensuring the reliability of the media processing results.
[0008] Furthermore, users can configure the AI model during media task processing themselves, eliminating the need to build an end-to-end business system for media processing, thereby improving the efficiency of applying media processing algorithms to actual business operations.
[0009] In one optional implementation, obtaining the algorithm configuration request is specifically implemented as follows: the client sends a second request to the cloud management platform to instruct the cloud management platform to configure the media processing template. In response to the client's second request, the cloud management platform provides a first interface to the client. This first interface includes an algorithm input box. In response to a user's first set of operations based on the first interface, the cloud management platform provides a second interface. In this second interface, the algorithm input box includes one or more first identifiers. In response to a user's second operation on the second interface, the cloud management platform obtains the algorithm configuration request.
[0010] Based on this optional implementation method, users send algorithm configuration requests to the cloud management platform through interface interaction, ensuring that users configure the media processing template themselves.
[0011] In one optional implementation, the first set of operations includes a first sub-operation and a second sub-operation. In this implementation, in response to the user's first set of operations on the first interface, the cloud management platform provides a first sub-interface in response to the user's first sub-operation on the algorithm input box. This first sub-interface includes identifiers for multiple AI models. In response to the user's second sub-operation on the identifiers of the multiple AI models in the first interface, the cloud management platform provides a second interface.
[0012] Optionally, multiple AI models include: AI models preset in the cloud management platform, and user-defined AI models uploaded by different users.
[0013] Based on this optional implementation, users select the corresponding first identifier from multiple AI model identifiers provided by the cloud management platform, ensuring that users can configure the AI model used in the media processing process themselves. Furthermore, the multiple AI models provided by the cloud management platform include user-defined AI models uploaded by different users, thus enabling the reuse of AI models in different media processing templates.
[0014] In one optional implementation, the first interface provided by the cloud management platform also includes a parameter input box. The second interface provided by the cloud management platform includes a transcoding parameter or a second identifier indicating the transcoding parameter in the parameter input box.
[0015] In this way, users can set transcoding parameters through the interface, improving the interactivity between the cloud management platform and the user.
[0016] In one optional implementation, the specific implementation is as follows: the first set of operations includes a third sub-operation and a fourth sub-operation. In responding to the user's first set of operations on the first interface, the cloud management platform, in response to the user's third sub-operation on the algorithm input box in the first interface, provides a second sub-interface. This second sub-interface includes a parameter input box and an algorithm input box, and the algorithm input box includes one or more first identifiers. The cloud management platform, in response to the user's fourth sub-operation on the parameter input box in the second sub-interface, provides the second interface.
[0017] In this way, users can configure the AI model and transcoding parameters in the media processing process through the third and fourth sub-operations, ensuring that the media processing templates are configured by the users themselves. This improves the flexibility and applicability of the media processing templates.
[0018] In one optional implementation, the first set of operations includes a fifth sub-operation and a sixth sub-operation. In responding to the user's first set of operations on the first interface, the cloud management platform, in response to the user's fifth sub-operation on the parameter input box in the first interface, provides a third sub-interface. This third sub-interface includes a parameter input box and an algorithm input box, and the parameter input box includes transcoding parameters or a second identifier. In response to the user's sixth sub-operation on the algorithm input box in the third sub-interface, the cloud management platform provides a second interface.
[0019] This ensures that users can configure media processing templates themselves, thereby improving the flexibility and applicability of media processing templates.
[0020] In one optional implementation, the user-defined AI model is the AI model uploaded by the client. The AI model preset by the cloud platform is the AI model pre-configured by the cloud management platform.
[0021] Optionally, the pre-configured AI models include one or two of the following: AI models provided by the cloud management platform through local files, and AI models provided by the cloud management platform through model call interfaces.
[0022] In this way, by increasing the number of available AI models, the adaptability between media processing templates and media processing scenarios can be improved, thereby ensuring the reliability of media processing results.
[0023] In one optional implementation, the client sends a model upload request to the cloud management platform. This request includes the resource specifications required for the candidate AI model and the file access path for executing the candidate AI model. The cloud management platform receives and responds to the client's model upload request, storing the candidate AI model as a user-defined AI model.
[0024] In this way, by having users upload their own AI models, the number of usable AI models can be increased, and the applicability of the AI models available on the cloud management platform can also be enhanced.
[0025] In one optional implementation, obtaining the model upload request from the client is specifically implemented as follows: The client sends a third request to the cloud management platform to request the upload of a user-defined AI model. The cloud management platform responds to the client's third request by providing a third interface. This third interface includes a model file path input box and a resource specification input box. In response to the user's second set of operations on the third interface, the cloud management platform provides a fourth interface. In this fourth interface, the model file path input box includes the file access path of the candidate AI model, and the resource specification input box includes the resource specifications required by the candidate AI model. In response to the user's third operation on the fourth interface, the cloud management platform obtains the model upload request.
[0026] Based on this possible implementation method, users can upload their own AI models, which can increase the number of usable AI models and enhance the applicability of the AI models available on the cloud management platform.
[0027] In one optional implementation, the cloud management platform also stores test cases or media data test sets, which are uploaded by the client. In storing the candidate AI model as a user-defined AI model, the cloud management platform retrieves the candidate AI model based on the file access path. The cloud management platform then verifies the candidate AI model based on the test cases or media data test sets. If the candidate AI model passes verification, the cloud management platform stores it as a user-defined AI model.
[0028] Optionally, if the candidate AI model fails the verification, the cloud management platform sends an upload failure message to the client, prompting the client to re-upload the candidate AI model.
[0029] Based on this optional implementation method, the AI model uploaded by the user can be verified using test cases or media data test sets, which can ensure the availability of the AI model and thus improve its reliability.
[0030] In one alternative implementation, obtaining the media processing template based on the algorithm configuration request is specifically achieved by the cloud management platform associating the transcoding parameters with one or more AI models indicated by the first identifier to obtain the media processing template.
[0031] This simplifies the media processing template configuration process, thereby improving media data processing efficiency.
[0032] In one optional implementation, the cloud management platform includes multiple transcoding nodes. In the process of processing the first media data according to the media processing template to obtain the second media data, the cloud management platform selects the target transcoding node to execute the media task from the multiple transcoding nodes based on the resource specifications of the media processing template, schedules the media task to the target transcoding node, and controls the target transcoding node to process the first media data according to the media processing template to obtain the second media data.
[0033] Optionally, the specifications of the target transcoding node are greater than or equal to the resource specifications of the media processing template; the resource specifications of the media processing template are related to the resource specifications of the AI model included in the media processing template.
[0034] Based on this optional implementation method, the cloud management platform selects the corresponding target transcoding node according to the resource specifications of the media processing template, ensuring the compatibility between the target transcoding node for processing media processing tasks and the resource specifications of the media processing template, thereby ensuring the normal operation of media processing tasks and thus guaranteeing the reliability of media processing.
[0035] In one optional implementation, controlling the target transcoding node to process the first media data according to the media processing template is specifically implemented as follows: The cloud management platform controls the target transcoding node to decode the first media data to obtain first intermediate media data. The cloud management platform controls the target transcoding node to call the AI model included in the media processing template to process the first intermediate media data to obtain second intermediate media data. The cloud management platform controls the target transcoding node to encode the second intermediate media data according to the transcoding parameters to obtain second media data.
[0036] Based on this optional implementation, the target transcoding node processes media by calling the AI model included in the media processing template, eliminating the need to build an end-to-end media processing business system. This improves the efficiency of applying the AI model to actual business operations, thereby enhancing the efficiency of the media processing method.
[0037] In one optional implementation, the cloud management platform further includes multiple candidate media processing templates. These templates are created based on the algorithm configuration request from the second client. The client sends a fourth request to the cloud management platform. This fourth request instructs the cloud management platform to process the third media data. The cloud management platform receives the fourth request, determines the target media processing template from the multiple candidate templates, and processes the third media data according to the target template to obtain the fourth media data.
[0038] Based on this optional implementation, in the process of a user requesting media processing services, the cloud management platform provides the user with media processing templates created by other users, allowing the user to select the target media processing template from multiple candidate templates. This enables the reuse of multiple media processing templates, thereby improving media processing efficiency.
[0039] In one optional implementation, the fourth request includes a third identifier. Based on this third identifier, the cloud management platform selects the target media processing template that matches the third identifier from multiple candidate media processing templates.
[0040] In this way, the target media processing template can be quickly obtained through a third identifier. Furthermore, allowing the user to select the target media processing template ensures compatibility between the target media processing template and the third-party request.
[0041] In one optional implementation, the fourth request also carries scene information of the third media data. The cloud management platform selects a target media processing template from multiple candidate media processing templates that matches the scene information of the third media data based on this scene information.
[0042] In this way, the target media processing template can be quickly obtained through scene information. Furthermore, it can improve the compatibility between the target media processing template and third-party media data, thereby ensuring the reliability of the processing results.
[0043] Secondly, this application provides a media data processing system deployed on a cloud management platform, which includes a communication module, a processing module, and a storage module.
[0044] The communication module acquires an algorithm configuration request and receives a first request from the client. The first request instructs the processing of the first media data. The algorithm configuration request includes one or more first identifiers indicating the AI model and transcoding parameters. These transcoding parameters include one or more of the following: encoding style, bitrate, and resolution. The AI models corresponding to the multiple first identifiers include: preset AI models in the cloud platform and user-defined AI models.
[0045] The processing module is used to obtain a media processing template based on the algorithm configuration request, and to process the first media data according to the media processing template to obtain the second media data.
[0046] The storage module is used to store executable program code, data from the media data processing system during media processing, such as AI models preset in the cloud platform, user-defined AI models, transcoding parameters, first media data, and second transcoding data.
[0047] Thirdly, this application provides a computing device cluster including at least one computing device. Each computing device includes a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the computing device cluster to perform the method as described in the first aspect or any optional implementation thereof.
[0048] Fourthly, embodiments of this application provide a computer program product containing instructions that, when executed by a computing device cluster, cause the computing device cluster to perform the method described in the first aspect or any optional implementation thereof.
[0049] Fifthly, embodiments of this application provide a computer-readable storage medium including computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the instructions in the computing program stored in the computer-readable storage medium to perform the method in the first aspect or any optional implementation of the first aspect.
[0050] The technical effects of any of the implementations in aspects two through five can be found in the first aspect or any optional implementation of the first aspect. Further details are omitted here. Based on the implementations provided in the above aspects, this application can be further combined to provide more implementations. Attached Figure Description
[0051] Figure 1 is a flowchart illustrating the media conversion process;
[0052] Figure 2 is a schematic diagram of the structure of a computer system provided in this application;
[0053] Figure 3 is a flowchart illustrating the media data processing method provided in this application.
[0054] Figure 4 is a schematic diagram of the media processing template provided in this application;
[0055] Figure 5 is a schematic diagram of the upload process for the user-defined AI model provided in this application;
[0056] Figure 6 is a schematic diagram of the process for selecting the target transcoding node provided in this application;
[0057] Figure 7 is a flowchart illustrating the media processing tasks performed by the target transcoding node provided in this application;
[0058] Figure 8 is a flowchart illustrating the media data processing method provided in this application (II).
[0059] Figure 9 is a flowchart illustrating the media data processing method provided in this application.
[0060] Figure 10 is a schematic diagram of the interface for obtaining the algorithm configuration request provided in this application;
[0061] Figure 11 is a schematic diagram of the interface for selecting an AI model provided in this application;
[0062] Figure 12 is a schematic diagram of the interface for obtaining the algorithm configuration request provided in this application;
[0063] Figure 13A is a schematic diagram of the interface for obtaining transcoding parameters provided in this application;
[0064] Figure 13B is a schematic diagram of the interface for obtaining transcoding parameters provided in this application (II).
[0065] Figure 13C is a schematic diagram of the interface for obtaining the algorithm configuration request provided in this application;
[0066] Figure 14 is a schematic diagram of the interface for obtaining the algorithm configuration request provided in this application;
[0067] Figure 15A is a schematic diagram of the interface for uploading candidate AI models provided in this application;
[0068] Figure 15B is a schematic diagram of the interface for uploading candidate AI models provided in this application (II).
[0069] Figure 16A is a schematic diagram of the interface for the resource specifications of the input candidate AI model provided in this application;
[0070] Figure 16B is a schematic diagram of the interface for the resource specifications of the input candidate AI model provided in this application.
[0071] Figure 17 is a flowchart illustrating the media data processing method provided in this application (Part 4);
[0072] Figure 18 is a schematic diagram of the media data processing system provided in this application;
[0073] Figure 19 is a schematic diagram of the structure of the computing device provided in this application;
[0074] Figure 20 is a schematic diagram of the structure of the computing device cluster provided in this application;
[0075] Figure 21 is a schematic diagram of the network connection between computing devices in the computing device cluster provided in this application. Detailed Implementation
[0076] The AI technology integrated into cloud-based media transcoding services has limited applicability and does not support modification. When the media processing scenario is not compatible with the AI technology integrated into the media transcoding service, the reliability of the media transcoding results is low.
[0077] Based on this, to improve the flexibility of AI technology integrated into media transcoding services and enhance the adaptability between AI technology integrated into media processing scenarios and media transcoding services, thereby ensuring the reliability of media transcoding results, this application provides a media data processing method. During media data processing, the user configures the AI model for the media task processing process, and the media processing template supports modification, thus increasing the flexibility and applicability of the media processing template. Furthermore, the AI models available to the user include preset AI models in the cloud platform and user-defined AI models. By increasing the number of usable AI models, the adaptability between the media processing template and the media processing scenario can be improved, thereby ensuring the reliability of the media processing results. In addition, since the user configures the AI model for the media task processing process, there is no need to build an end-to-end media processing business system, thereby improving the efficiency of applying media processing algorithms to actual business operations.
[0078] Specifically, based on the algorithm configuration request sent by the client, the cloud platform configures a media processing template. This media processing template includes transcoding parameters and one or more AI models corresponding to first identifiers. The AI models corresponding to the multiple first identifiers include preset AI models in the cloud platform and user-defined AI models. Upon receiving the first request from the client, the cloud platform processes the first media data indicated by the first request according to the media processing template to obtain the second media data.
[0079] It should be noted that the media data processing in the embodiments of this application includes, but is not limited to, media transcoding, media style conversion, and content generation. This application does not limit these aspects.
[0080] The terminology used in the implementation section of this application is only for explaining specific embodiments of this application and is not intended to limit this application. A brief introduction to some concepts that may be involved in this application is given below.
[0081] Media transcoding refers to converting media data from one encoding format to another. Taking images as an example, this includes converting an image from low resolution to high resolution, or vice versa, or converting an image from lossy compression (Joint Photographic Experts Group, JPG) format to portable network graphics (PNG) format, etc.
[0082] As shown in Figure 1(a), the media conversion process includes the following two stages: Stage ① decodes the input media data to obtain decoded media data; Stage ② encodes the decoded media data according to the encoding parameters to obtain transcoded media data.
[0083] Encoding parameters, also known as transcoding parameters, include one or more of the following: encoding style, resolution, and bitrate. Additionally, in some examples, encoding parameters may include one or more of the following: video width, video height, landscape / portrait adaptation, frame rate, maximum number of frames per second (MPB), and peak bitrate.
[0084] For example, taking encoding parameters including encoding style, resolution, and bitrate as examples, the encoding style, resolution, and bitrate will be explained below.
[0085] Encoding style is used to indicate the encoding format and image style of the output media data.
[0086] The encoding format is used to indicate the standards and specifications for storing and transmitting media data. Taking images as an example, encoding formats include, but are not limited to, PNG, JPG, Graphics Interchange Format (GIF), and bitmap (BMP). Taking video as an example, encoding formats include, but are not limited to, MPEG-4 Part 2, H.264, H.256, VP9, and AV1.
[0087] Image style is used to indicate the visual characteristics of an image. Commonly used image styles include, but are not limited to: cartoon style, illustration style, color, etc.
[0088] Resolution is used to quantify image quality. Resolution includes spatial resolution and pixel resolution.
[0089] Spatial resolution indicates the sharpness of an image. Pixel resolution indicates the quality of pixels contained in the width and height of an image. For example, a pixel resolution of 1920*1080 means that the image has 1920 horizontal pixels and 1080 vertical pixels. In the following text, resolution generally refers to pixel resolution.
[0090] Bitrate is used to indicate the speed at which media data is transmitted and the file size. Generally, higher bitrate media data corresponds to larger files, which will occupy more storage space.
[0091] In addition, in some other implementations, AI processing can be added to the media conversion process shown in Figure 1(a). As shown in Figure 1(b), compared to the media conversion process shown in Figure 1(a), the media transcoding process shown in Figure 1(b) adds a stage 3 AI processing between stage ① and stage ②.
[0092] AI processing refers to using AI models to perform one or more of the following processes on the decoded media data in order to adjust the display effect of the media data: image style conversion, content generation, image enhancement, and resolution adjustment.
[0093] AI models include machine learning-based models and neural network-based models.
[0094] Media transcoding services based on cloud computing technology refer to media transcoding services deployed in a cloud environment.
[0095] Cloud environment: An entity that provides cloud services to users using basic resources under the cloud computing model. The cloud environment includes cloud data centers and cloud service platforms.
[0096] Cloud data centers encompass a vast amount of basic resources (including computing clusters, storage resources, and network resources) owned by cloud service providers. In this article, cloud data centers may also be referred to as cloud management platforms, cloud computing platforms, etc.
[0097] Host: A physical server deployed in a cloud management platform. The physical resources of a host include the physical central processing unit (CPU) and memory devices. Each host runs virtualization software, which virtualizes some physical resources into virtual resources for use by instances. For example, the virtualization software virtualizes the CPU into a virtual CPU (vCPU). The host also has resources such as memory channels, cache channels, caches, network input / output (I / O) bandwidth, and storage I / O bandwidth, which are shared by the instances running on the host.
[0098] Instance: A compute node running on a host machine. Common instances include virtual machines (VMs) or containers. Each instance consumes some or all of the host's virtual resources. Instance specifications include instance type (also known as flavor) and instance specifications. In this document, instances are used to perform media processing, such as media transcoding and content generation. Furthermore, in this document, instances may also be referred to as processing nodes, transcoding nodes, etc., and this application does not limit the terminology used.
[0099] Instance type indicates the resource characteristics of an instance. For example, Economy instances use cheaper CPUs, consume less computing resources, and have lower costs; Compute-Enhanced instances use high-performance CPUs, consume ample computing resources, and have higher costs; Network-Enhanced instances consume ample network resources, such as being configured with high IO bandwidth, and have higher costs compared to Economy instances. Different instance types represent the resource requirements of tenants for their services running on Flexible instances.
[0100] The instance specifications, also known as instance size or instance resource specifications, indicate the amount of resources an instance uses.
[0101] Specifications: Including the number of vCPUs and memory size (in gigabytes, GB).
[0102] Alternatively, specifications include the number of vCPUs, memory size, memory bandwidth, network bandwidth, number of graphics processing units (GPUs), and the size of non-volatile storage devices (typically high-speed storage media such as solid-state drives (SSDs) and NVMe SSDs). In this article, instance specifications include the number of vCPUs and memory size (gigabytes, GB).
[0103] Next, the method for processing media data provided in this application will be described in detail with reference to the accompanying drawings.
[0104] First, referring to Figure 2, which is a schematic diagram of the structure of a computer system provided in this application. As shown in Figure 2, the computer system includes a client 20, a cloud service platform 30, and a cloud management platform 10.
[0105] In the first optional implementation, the client 20 can be a computer running an application. This computer can be a physical machine or a virtual machine. For example, if the computer running the application is a physical computing device, it can be a host or a terminal. The terminal can also be called a terminal device, user equipment (UE), mobile station (MS), mobile terminal (MT), etc. Terminals can be mobile phones, tablets, laptops, desktop computers, virtual reality (VR) terminal devices, augmented reality (AR) terminal devices, wireless terminals in industrial control, wireless terminals in autonomous driving, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, etc. The embodiments of this application do not limit the specific technology or device form adopted by the client 20.
[0106] In the second alternative implementation, client 20 can be an application, such as a cloud computing application. Alternatively, client 20 can be a web client. Client 20 runs on a terminal device. The terminal device includes, but is not limited to, mobile terminals, tablets, personal computers, or laptops.
[0107] It should be noted that the above two implementation methods are merely different implementations of client 20. In practical applications, client 20 can also have other implementation methods. For example, client 20 can be a software module running on any one or more hosts in cluster 110. This application does not limit this.
[0108] In one alternative implementation, a media data processing system 120 is deployed in the cloud management platform 10.
[0109] In some implementations, the media data processing system 120 can be deployed independently within an instance of the cloud management platform 10. Alternatively, the media data processing system 120 can be deployed in a distributed manner across multiple instances of the cloud management platform 10.
[0110] As shown in Figure 2, the media data processing system 120 is abstracted into a cloud service by the cloud service provider on the cloud service platform 30 and provided to users. After the user purchases the cloud service through the client 20 on the cloud service platform 30 (pre-payment is possible, with settlement based on the final resource usage), the cloud environment utilizes the media data processing system 120 deployed on the cloud management platform 10 to provide the cloud service to the user. When using the cloud service, the user can determine the tasks to be executed and upload data to the cloud environment through the application program interface (API) or graphical user interface (GUI) in the client 20. The media data processing system 120 in the cloud environment receives the user's task information and data, performs data processing and corresponding tasks, and obtains the processing results. The media data processing system 120 returns the processing results of the tasks or the status information during the execution process to the user through the API or GUI. Among these tasks, there are including but not limited to media processing tasks, media transcoding tasks, live transcoding tasks, audio and video playback tasks, AI recognition tasks, etc. This application does not limit these.
[0111] In addition, in some other embodiments, the functions of the media data processing system 120 may be performed by the cloud management platform 10 or other components of the cloud management platform 10, and this application does not limit the implementation method.
[0112] In one alternative implementation, as shown in Figure 2, the cloud management platform 10 also includes a cluster 110.
[0113] Cluster 110 refers to a collection of computers connected via a local area network or the Internet, which includes multiple transcoding nodes. As shown in Figure 2, cluster 110 includes transcoding node 111 and transcoding node 112.
[0114] Cluster 110 is typically used to execute large tasks (also known as jobs). These jobs are usually large-scale operations requiring significant computing resources for parallel processing; this embodiment does not limit the nature or number of jobs. A job may contain multiple computational tasks, which can be distributed across multiple computing resources for execution. Most tasks are executed concurrently or in parallel, while some tasks depend on data generated by other tasks. Each computing device in Cluster 110 uses the same hardware and the same operating system; alternatively, different hardware and operating systems can be used on the hosts of Cluster 110 depending on business needs. Because tasks deployed using Cluster 110 can be executed concurrently, overall performance can be improved.
[0115] The media data processing system 120 establishes a communication connection with each transcoding node, distributes media processing tasks to instances on the transcoding nodes, receives processing results returned by the instances, and returns the processing results to the client 20.
[0116] It should be noted that the computer system architecture shown in Figure 2 is merely an example. The types or number of devices within the system can be configured according to actual needs, and this application embodiment does not limit this. For example, the computer system may also include more clusters 110.
[0117] The following description, with reference to Figure 2, provides an exemplary illustration of the various modules included in the media data processing system 120.
[0118] As shown in Figure 2, the media data processing system 120 includes an algorithm management unit 121 and a transcoding management unit 122 that are interconnected.
[0119] It is worth noting that the functions implemented by the algorithm management unit 121 and the transcoding management unit 122 can be performed by the cloud management platform 10 or other components of the cloud management platform 10, and this application does not limit their implementation.
[0120] The functions implemented by the algorithm management unit 121 and the transcoding management unit 122 are described below by way of example.
[0121] The transcoding management unit 122 is used to manage the resources used by the media processing service, schedule media processing tasks, and configure media processing templates.
[0122] Managing the resources used by media processing services includes: expanding and shrinking resources, and managing hardware with different computing power.
[0123] In this document, "scheduling of media processing tasks" can refer to scheduling media processing tasks to transcoding nodes that match the media processing tasks. "Transcoding nodes that match media processing tasks" can mean transcoding nodes whose resource specifications meet the load requirements of the media processing tasks. In some implementations, the transcoding nodes that match the media processing tasks can be determined with reference to the embodiments shown in Figure 6 below, which will not be described in detail here.
[0124] Configuring a media processing template can refer to configuring the AI model used in the media processing process, as well as the processing logic between different AI models. In some implementations, the method of configuring a media processing template can be referred to the embodiments provided in Figures 10 to 14 below, which will not be described in detail here.
[0125] The algorithm management unit 121 is used to manage the AI models that can be used in the media processing process. In this application, the AI models that can be used in the media processing process include at least two types: user-defined AI models and AI models pre-configured by the cloud management platform 10. The AI models pre-configured by the cloud management platform 10 include AI models provided by the cloud management platform 10 through a model call interface or AI models provided by the cloud management platform 10 through local files. These local files refer to files stored in the cloud platform used to execute the AI models.
[0126] The user-defined AI model can refer to the AI model that is uploaded to the cloud management platform through the client. Specifically, the implementation method of uploading the AI model by the user can be referred to the embodiments provided in Figures 5, 15A to 16B below, which will not be described in detail here.
[0127] The AI models provided by the cloud management platform 10 through local files include one or more of the following: AI models pre-uploaded by the cloud platform administrators, AI models generated using generative models deployed in the cloud platform, and AI models read from external storage devices.
[0128] The AI model provided by the cloud management platform 10 through the model call interface can refer to: the cloud management platform 10 provides one or more model call interfaces, and the cloud management platform 10 calls the corresponding AI model by calling the model call interface.
[0129] In the first example, the AI model provided by the model call interface can refer to an AI model provided by another platform. This other platform can refer to the platform used to provide the algorithm model, such as TensorFlow Hub, PyTorch Hub, GitHub, etc.
[0130] In the second example, the AI model provided by the model call interface can refer to an AI model provided by a database. This database is used to store multiple AI models.
[0131] It should be noted that the two examples above are only different forms of AI models provided by the model calling interface. In practical applications, the AI models provided by the model calling interface can also take other forms, such as AI models uploaded to the cloud platform by other users. This application does not limit this.
[0132] It should be noted that Figure 2 is merely an exemplary drawing and does not constitute a limitation on the media data processing method provided in the embodiments of this application. The naming and division of modules in the media data processing system 120 shown in Figure 2 are illustrative. In practical applications, the media data processing system 120 may have other grouping methods, which are not limited in this application. In addition, the media data processing system 120 may also be named a media data processing system 18. The media data processing system 18 may include modules different from those shown in Figure 2. The media data processing system 18 can be referred to the embodiments provided in Figure 18 below, which will not be described in detail here.
[0133] The implementation of the media data processing method provided in the embodiments of this application will be described next with reference to Figures 3 to 17.
[0134] Please refer to Figure 3, which is a flowchart illustrating the media data processing method provided in this embodiment. The media data processing method shown can be applied to the cloud management platform 10 shown in Figure 2 and executed by the cloud management platform 10. This cloud management platform 10 can be referred to as a cloud computing platform, management node, etc. Alternatively, the media data processing method can also be executed by the media data processing system 120 within the cloud management platform 10.
[0135] Taking the media data processing method provided in this application embodiment as an example, which is executed by the cloud management platform 10, as shown in Figure 3, the media data processing method includes steps S310 to S340.
[0136] S310, Cloud Management Platform 10 obtains algorithm configuration request.
[0137] Corresponding to the S310 process, client 20 sends an algorithm configuration request to cloud management platform 10.
[0138] The algorithm configuration request includes one or more first identifiers for indicating the AI model and one or more transcoding parameters.
[0139] The following provides an illustrative example of the first identifier included in the algorithm configuration request.
[0140] In one alternative implementation, the first identifier may refer to the model identifier of the AI model.
[0141] In some examples, the model identifier can be the model name of an AI model, such as Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), Generative Adversarial Network (GAN), etc.
[0142] In other examples, the model identifier can also be the type of processing implemented by the AI model, such as content generation, modality transformation, etc.
[0143] In some other examples, the model identifier can also be a model identifier specified by the user or the cloud management platform 10, such as Algorithm 1, Algorithm 2, etc.
[0144] In the embodiments of this application, the specific form of the model identifier is not limited.
[0145] When the algorithm configuration request includes a first identifier, the algorithm configuration request is used to instruct the cloud management platform 10 to configure an AI model during media processing.
[0146] In cases where the algorithm configuration request includes multiple first identifiers, the algorithm configuration request is used to instruct the cloud management platform 10 to configure multiple AI models during media processing.
[0147] It is worth noting that when configuring multiple AI models, it is also necessary to configure the execution logic between these AI models. The following is an example of configuring the execution logic between multiple AI models.
[0148] In the first alternative example, the execution logic between multiple AI models can be determined by utilizing the order among the multiple first identifiers included in the algorithm configuration request. For example, if the multiple first identifiers included in the algorithm configuration request are Algorithm 1, Algorithm 2, and Algorithm 3, then the execution logic between the multiple AI models would be Algorithm 1 → Algorithm 2 → Algorithm 3.
[0149] In a second alternative example, the algorithm configuration request may also include the execution order between multiple AI models indicated by a first identifier. For example, the algorithm configuration request may include a field “1: Algorithm 1; 2: Algorithm 3; 3: Algorithm 2”, which indicates the execution order between AI models as Algorithm 1 → Algorithm 3 → Algorithm 2.
[0150] The two examples above are merely implementation methods for configuring different execution logics among multiple AI models; other implementation methods are possible in practical applications. This application does not limit these methods.
[0151] The above explanation primarily uses the model identifier of the AI model as an example. In other embodiments, the first identifier may also indicate the identifier of the model template. The model template is used to indicate one or more AI models, and in the case of multiple AI models, the execution logic between them. The following explanation uses the identifier of the model template as an example to illustrate the first identifier and the AI model it indicates.
[0152] In one alternative implementation, a model template is used to perform a media processing, for example, model template 1 is used to perform cartoon style conversion, model template 2 is used to perform content generation, and model template 3 is used to perform high-resolution reconstruction.
[0153] When the algorithm configuration request includes an identifier for a model template, the algorithm configuration request is used to instruct the cloud management platform 10 to configure a model template during media processing.
[0154] When the algorithm configuration request includes identifiers for multiple model templates, the algorithm configuration request is used to instruct the cloud management platform 10 to configure multiple model templates during media processing.
[0155] Correspondingly, the execution logic between multiple model templates can refer to the execution logic between multiple AI models mentioned above, and this application will not elaborate on this.
[0156] The above primarily provides an illustrative explanation of the first identifier in the algorithm configuration request. The following provides an illustrative explanation of one or more transcoding parameters included in the algorithm configuration request.
[0157] In one alternative implementation, the transcoding parameters can be user-configurable. Alternatively, the transcoding parameters can be set by the client itself. This application does not limit this. The implementation of user-configurable transcoding parameters can be referred to the embodiments provided in Figures 12 to 14 below, which will not be described in detail here.
[0158] In one alternative implementation, the algorithm configuration request may carry a field indicating the numerical values of one or more transcoding parameters. For example, the first request may carry the following fields: xx1920*1080; xx4m / s, where xx1920*1080 indicates a resolution of 1920*1080 and xx4m / s indicates a bitrate of 4m / s.
[0159] In another implementation, the algorithm configuration request may also carry a second identifier to indicate transcoding parameters. The second identifier is illustrated below.
[0160] In one alternative implementation, the second identifier may refer to a transcoding parameter template identifier. A transcoding parameter template includes a set of transcoding parameters. For example, transcoding parameter template 1 includes encoding style and bitrate. Transcoding parameter template 2 includes encoding style, resolution, and bitrate.
[0161] The two optional implementation methods described above are merely illustrative examples of different forms of one or more transcoding parameters included in the algorithm configuration request. In practical applications, the algorithm configuration request may include transcoding parameters in other forms. This application does not limit this.
[0162] The following provides an example of how to obtain the algorithm configuration request.
[0163] In the first alternative implementation, the algorithm configuration request can be sent by the client.
[0164] For example, when the cloud management platform 10 receives a media processing task from a user through a client, the client obtains the first identifier and transcoding parameters input by the user, and sends an algorithm configuration request to the cloud management platform 10 based on the first identifier and transcoding parameters.
[0165] For example, the client sends an algorithm configuration request to the cloud management platform 10 based on the first and second identifiers input by the user.
[0166] In the second alternative implementation, the algorithm configuration request can also be triggered by the cloud management platform 10.
[0167] For example, a user sends a first identifier and a second identifier to the cloud management platform 10 via a client. The cloud management platform 10 triggers an algorithm configuration request based on the received first and second identifiers.
[0168] For example, a user sends a first identifier and transcoding parameters to the cloud management platform 10 via a client. The cloud management platform 10 then triggers an algorithm configuration request based on the received first identifier and transcoding parameters.
[0169] The two implementation methods described above are merely different triggering methods for algorithm configuration requests. In other embodiments, algorithm configuration requests may have other triggering methods, which this application does not limit. For example, algorithm configuration requests may also be triggered based on the interface. Specifically, the implementation method of triggering algorithm configuration requests through the interface can be referred to the embodiments provided in Figures 10 to 14 below, which will not be described in detail here.
[0170] S320, the cloud management platform 10 obtains the media processing template according to the algorithm configuration request.
[0171] The media processing template is used to indicate the media processing flow, the AI technology used in the media processing flow, and the encoding format of the output processing results.
[0172] In one alternative implementation, the cloud management platform 10 associates the AI model indicated by the first identifier with the transcoding parameters included in the algorithm configuration request and the AI model indicated by the first identifier to obtain a media processing template.
[0173] Since the algorithm configuration request includes one or more first identifiers and one or more transcoding parameters, the media processing template includes: one or more AI models corresponding to the first identifiers, and one or more transcoding parameters.
[0174] The following examples illustrate the contents of the media processing template using algorithm configuration requests under different circumstances.
[0175] In one alternative implementation, the algorithm configuration request includes a first identifier and multiple transcoding parameters.
[0176] In the first optional example, the first identifier is the model identifier of the AI model. The media processing template includes an AI model and multiple transcoding parameters. For example, the model identifier of the AI model is Algorithm 1, and the transcoding parameters are resolution 1920*1080 and bitrate 4m / s, as shown in Figure 4(a). The media processing template includes the following media processing flow: input media data → decoding → Algorithm 1 → encoding → output media (resolution 1920*1080, bitrate 4m / s).
[0177] In the second optional example, the first identifier is the identifier of the model template, and the second identifier is the identifier of the transcoding parameter template. When the model template indicated by the identifier of the model template includes multiple AI models, and the transcoding parameter template indicated by the identifier of the transcoding parameter template includes a set of transcoding parameters, the media processing template includes multiple AI models and multiple transcoding parameters. For example, the identifier of the model template is model template 1 (algorithm 1 → algorithm 2), and the identifier of the transcoding parameter template is transcoding parameter template 1 (resolution 1920*1080, bitrate 4m / s), as shown in Figure 4(b). The media processing template includes the following media processing flow: input media data → decoding → model template 1 (algorithm 1 → algorithm 2) → encoding → output media (transcoding parameter template 1 (resolution 1920*1080, bitrate 4m / s)).
[0178] In the third optional example, the first identifier is the model identifier of the AI model, and the second identifier is the transcoding parameter template identifier. When the transcoding parameter template indicated by the transcoding parameter template identifier includes a set of transcoding parameters, the media processing template includes an AI model and multiple transcoding parameters. For example, the model identifier of the AI model is Algorithm 1, and the transcoding parameter template identifier is Transcoding Parameter Template 1 (resolution 1920*1080, bitrate 4m / s), as shown in Figure 4(c). The media processing template includes the following media processing flow: input media data → decoding → Algorithm 1 → encoding → output media (transcoding parameter template 1 (resolution 1920*1080, bitrate 4m / s)).
[0179] The three examples above represent different forms of media processing templates when the algorithm configuration request includes a first identifier and multiple transcoding parameters. In other embodiments, the media processing template may have other forms. This application does not limit this.
[0180] Alternatively, in another optional implementation, the algorithm configuration request includes a first identifier and multiple second identifiers. Or, the algorithm configuration request includes multiple first identifiers and multiple second identifiers.
[0181] Taking an algorithm configuration request that includes a first identifier and multiple transcoding parameters as an example, where the first identifier is the model identifier of the AI model, the media processing template includes: an AI model and multiple transcoding parameters. For example, if the model identifier of the AI model is Algorithm 2, and the multiple transcoding parameters include: resolution 1920*1080, bitrate 4m / s, the media processing template includes: the media processing flow is: input media data → decoding → Algorithm 2 → output media (resolution 1920*1080, bitrate 4m / s).
[0182] Taking an algorithm configuration request that includes multiple first identifiers and multiple transcoding parameters as an example, where the first identifier is the model identifier of the AI model, the media processing template includes: multiple AI models and multiple transcoding parameters. For example, the model identifiers of the multiple AI models include: Algorithm 3, Algorithm 4, and Algorithm 5; the multiple transcoding parameters include: resolution 1920*1080, bitrate 4m / s. The media processing template includes the following media processing flow: input media data → decoding → Algorithm 3 → Algorithm 4 → Algorithm 5 → output media (resolution 1920*1080, bitrate 4m / s).
[0183] The above primarily uses algorithm configuration requests under different circumstances as examples to illustrate different contents of media processing templates. It should be understood that the various forms of media processing templates provided above do not constitute a limitation on the media data processing method provided in this application. In practical applications, media processing templates may also have other forms, which this application does not limit.
[0184] The following is an example of a scheme for associating the AI model indicated by the first identifier with the transcoding parameters.
[0185] The first optional implementation method is to bind the AI model indicated by the first identifier to the transcoding parameters.
[0186] The second optional implementation method is to write the AI model indicated by the first identifier and the transcoding parameters into the same data table, and use the data table to associate the AI model indicated by the first identifier with the transcoding parameters.
[0187] It should be noted that the above two implementation methods are only different ways of associating AI models and transcoding parameters. In other embodiments, there may be other implementation methods, which are not limited in this application.
[0188] The above mainly uses the method of associating the AI model indicated by the first identifier with the transcoding parameters as an example to illustrate the method of obtaining the media processing template. In other embodiments, other implementation methods can be used to obtain the media processing template, and this application does not limit them.
[0189] For example, the cloud management platform 10 creates an initial media processing template based on the transcoding parameters included in the algorithm configuration request, and adds the AI model indicated by the first identifier to the initial media processing template to obtain the media processing template.
[0190] For example, the cloud management platform 10 creates an initial media processing template based on the AI model indicated by the first identifier, and associates the initial media processing template with the transcoding parameters included in the algorithm configuration request to obtain the media processing template.
[0191] For example, the cloud management platform 10 provides a template configuration interface to the user based on the AI model and transcoding parameters indicated by the first identifier. This interface displays candidate media processing templates, which the user can modify or confirm. These candidate media processing templates are created based on the AI model and transcoding parameters indicated by the first identifier. When the user inputs a modification operation on the template configuration interface, the cloud management platform 10 responds by adjusting the candidate media processing template to obtain the desired media processing template. When the user inputs a confirmation operation on the template configuration interface, the cloud management platform 10 responds by confirming the candidate media processing template as the desired media processing template.
[0192] The "user input modification operation based on template configuration interface" can include at least one of the following: adding AI model, deleting AI model, replacing AI model, adjusting the execution logic between AI models, modifying transcoding parameters, deleting transcoding parameters, and adding transcoding parameters.
[0193] S330, cloud management platform 10 receives the first request from the client.
[0194] Corresponding to the S330 process, client 20 sends the first request to cloud management platform 10.
[0195] In one alternative implementation, the first request can be any media processing request in the media processing stream. The media processing stream, also referred to as a data stream, includes multiple media processing requests, each instructing the processing of first media data. Different media processing requests indicate different first media data.
[0196] The following three specific examples illustrate the first request.
[0197] In the first alternative example, the first request originates from a media processing stream sent from the same device. This device can refer to a terminal, user equipment, server, or other type of device.
[0198] In the second alternative example, the first request comes from a media processing stream sent by the same application.
[0199] For example, the application can be deployed on one or more devices that communicate with the cloud platform. This application can include, but is not limited to, cloud PC applications, video playback applications, live streaming applications, and game applications.
[0200] In a third alternative example, the first request originates from a media processing stream within the same task. This task could be a video playback task, a video upload task, a video distribution task, a live streaming task, a real-time video call task, a game, etc. This application does not limit this.
[0201] The three possible examples above are merely optional methods for the first request provided in the embodiments of this application. Multiple first requests belonging to the same media processing stream indicate that media data flows from one device to another. The direction of the media processing stream can be input (cloud management platform 10 receives media data) or output (cloud management platform 10 sends media data to other devices). In some optional cases, the media processing stream may also be called a request sequence, a data request stream, or other names, etc., and this application does not limit it in this regard.
[0202] In one alternative implementation, the first request indicates that the first media data can be implemented in multiple ways, for example:
[0203] In a first alternative example, the first request may carry first media data, which the cloud management platform 10 obtains by parsing the first request. For example, the message of the first request includes fields for the first media data.
[0204] In the second alternative example, the first request carries an access path used to retrieve the first media data. The cloud management platform 10 retrieves the first media data based on this access path.
[0205] The access path can refer to the cloud access path of the first media data on the cloud server, or it can refer to the local access path of the first media data on the local device. If the access path is the cloud access path, the user can upload the first media data to the cloud server through the client, obtain the cloud access path of the first media data on the cloud server, and then send a first request to the cloud management platform 10 based on the cloud access path of the first media data on the cloud server.
[0206] In a third alternative example, the first request carries a media data identifier, which indicates first media data. The cloud management platform 10 queries the database based on the media data identifier to retrieve the first media data that matches the media data identifier.
[0207] The three optional examples described above are merely different implementations of the first request indicating the first media data. In other embodiments, the first request indicating the first media data can also be implemented in other ways. For example, the first request may carry a storage address, which indicates the storage address of the first media data in a cloud server or a local server. The cloud management platform 10 uses the storage address carried in the first request to read the first media data. This application does not limit this.
[0208] The three examples described above are merely illustrative of the first request indicating the first media data. In other embodiments, the first request may also indicate a request type. Request types include offline processing requests and real-time processing requests.
[0209] The following sections explain offline processing requests and real-time processing requests respectively.
[0210] An offline processing request is used to instruct the cloud management platform 10 to perform offline processing of the media data. Specifically, it requests the cloud management platform 10 to process the entire first media data when the first media data is not in an active network stream state.
[0211] In the case that the first request is an offline processing request, the cloud management platform 10 provides the processing result of the first media data after the entire first media data processing is completed.
[0212] A real-time processing request is used to instruct the cloud management platform 10 to perform real-time processing of media data. Specifically, it requests the cloud management platform 10 to process the first media data during the transmission of the first media data and provide each processed frame of the first media data to the user as a data stream.
[0213] The following provides an exemplary description of how the first request indicates the request type is implemented.
[0214] In the first optional implementation, the first request carries a request type identifier code, and the cloud management platform 10 determines the request type of the first request based on the request type identifier code carried in the first request. This application does not limit the specific form of the request type identifier code. For example, if the request type identifier code carried in the first request is XX01, the request type of the first request is an offline processing request. As another example, if the request type identifier code carried in the first request is XX10, the request type of the first request is an online processing request.
[0215] In the second optional implementation, the request type is related to the latency of processing media data. Offline processing latency is greater than real-time processing latency. Therefore, the first request can carry a latency indicating its type. For example, the cloud management platform 10 obtains the latency carried by the first request. If the latency is less than or equal to a preset latency threshold, the cloud management platform 10 determines the first request type as a real-time processing request. If the latency is greater than the preset latency threshold, the cloud management platform 10 determines the first request type as an offline processing request. This application does not limit the specific value of the preset latency threshold.
[0216] In the third optional implementation, the request type is related to the task to which the first request belongs. If the first request carries a task identifier, and the task identifier indicates that the task to which the first request belongs is a live streaming task or a real-time video call task, the cloud management platform 10 determines that the request type of the first request is a real-time processing request. If the task identifier carried by the first request indicates that the task to which the first request belongs is a video upload task or a video distribution task, the cloud management platform 10 determines that the request type of the first request is an offline processing request.
[0217] The three optional implementations described above are merely different ways to indicate the request type of the first request. In other embodiments, the first request indicating the request type can also adopt other implementations. For example, the first request can carry an application identifier corresponding to the application that triggered the first request. If the application identifier indicates a live streaming application or a game application, the request type indicated by the first request is a real-time processing request. If the application identifier indicates a cloud computing application or a video playback application, the request type indicated by the first request is an offline processing request. This application does not limit this.
[0218] Furthermore, in some embodiments, after processing the first media data, the cloud management platform 10 needs to send or return the processed first media data. Therefore, the first request can also indicate the target terminal receiving the processed first media data. For example, the first request may carry a terminal identifier for indicating the target terminal, which may be the terminal's IP address, MAC address, device identifier, etc., and this application does not limit this.
[0219] The following two examples illustrate how a target terminal receives and processes the first media data.
[0220] In a first alternative example, the target terminal may be the device where the client that sent the first request to the cloud management platform 10 is located.
[0221] For example, in scenarios involving interaction between a single device and a cloud platform, such as video playback or video-on-demand, the first media data processed by the client that sends the first request is played.
[0222] In a second alternative example, the target terminal may also be a device containing a second client that is different from the client that sent the first request; this application does not limit this.
[0223] For example, in scenarios involving multiple device interactions, such as video distribution, video transmission, video calls, or live streaming, the first media data after processing is played by the target terminal where the second client is located.
[0224] The two optional examples above are only specific forms of the target terminal in different scenarios. In other embodiments, the target terminal may have other forms, which are not limited in this application.
[0225] The above provides an exemplary description of the content included in the first request. To better illustrate the first request, the following provides an exemplary description of how the first request is triggered.
[0226] In one alternative implementation, the first request can be triggered by a user interface. For example, the cloud management platform 10 provides a media processing interface to the client, which then displays the interface. The media processing interface displays function buttons that trigger media processing tasks. In response to the user's input to the function buttons for the media processing task, the client sends a first request to the cloud management platform 10. The input operation can be a contact-based operation, such as tapping, long-pressing, swiping, double-tapping, or clicking controls on the interface. Alternatively, the input operation can be a contactless operation, such as gesture input, physical button input, or voice input. Or, the input operation can be a typing operation, such as typing with a mouse or keyboard.
[0227] In some alternative implementations, the first request can be event-triggered. This event could be a client-triggered live streaming service request event, video playback request event, or a client-triggered media data broadcast event.
[0228] For example, when a client requests a live streaming service or a video playback service from the cloud management platform 10, the client sends a first request to the cloud management platform 10 based on the live streaming service request event or the video playback request event.
[0229] For example, when a user requests the publication of media data to the cloud management platform 10 through a client, the client triggers a media data broadcast event and sends the first request to the cloud management platform 10.
[0230] In addition, in some alternative implementations, the first request can also be triggered in a contextual manner. For example, if the cloud service requested by the client from the cloud management platform 10 includes audio and video playback, the client sends the first request to the cloud management platform 10.
[0231] It should be noted that the above implementation methods are only used to describe different triggering methods for the first request. In other embodiments, the first request may also have other triggering methods, such as triggering via voice. This application does not limit this.
[0232] S340, the cloud management platform 10 processes the first media data according to the media processing template to obtain the second media data.
[0233] In one alternative implementation, the cloud management platform 10 can read the first media data stored in the storage system, or the cloud management platform 10 can receive the first media data sent by the client or other devices. The storage system can refer to the storage system deployed on the cloud platform where the cloud management platform 10 resides, or a storage system deployed on other cloud platforms, or the storage system of the device where the client resides; this application does not limit this.
[0234] The following two specific examples illustrate how the cloud management platform 10 obtains the first media data under offline processing requests and real-time processing requests, respectively.
[0235] In a first alternative example, the first request is an offline processing request. The cloud management platform 10 responds to the first request and reads the first media data stored in the storage system. For example, the cloud management platform 10 accesses the storage system of the client's device and reads the first media data from the storage system. Another example is that the cloud management platform 10 reads the first media data from the storage system based on the access path carried in the first request. Yet another example is that the cloud management platform 10 reads the first media data from the storage system based on the storage address carried in the first request.
[0236] In the second alternative example, the first request is a real-time processing request. The cloud management platform 10 receives first media data sent by the client or other device. For example, the cloud management platform 10 responds to the first request by receiving the first media data carried in the first request. Alternatively, the cloud management platform 10 obtains the first media data sent by the client or other device via RTMP.
[0237] The two examples above are merely different implementations of the cloud management platform 10 obtaining the first media data, and should not be construed as limiting the media data processing method provided in this application. In other embodiments, the cloud management platform 10 may also use other implementations to obtain the first media data, and this application does not limit such implementations.
[0238] In this application, the second media data refers to the first media data after being processed by the media processing template. In some optional implementations, the second media data may be referred to as the processed first media data, the transcoded first media data, etc. This application does not limit this.
[0239] In one alternative implementation, the cloud management platform 10 may process the first media data with reference to the embodiments provided in Figures 6 and 7 below, which will not be described in detail here.
[0240] In one alternative implementation, after obtaining the second media data, the cloud management platform 10 can store the second media data in the cloud platform's storage system, return the second media data to the client, or send the second media data to the target terminal indicated by the terminal identifier.
[0241] Based on the embodiment shown in Figure 3, users configure the AI model during media task processing themselves, and the media processing template can be modified, thus increasing the flexibility and applicability of the media processing template. Furthermore, the AI models available to users include preset AI models in the cloud platform and user-defined AI models. By increasing the number of usable AI models, the adaptability between the media processing template and the media processing scenario can be improved, thereby ensuring the reliability of the media processing results. In addition, since users configure the AI model during media task processing themselves, there is no need to build an end-to-end media processing business system, thereby improving the efficiency of applying media processing algorithms to actual business operations.
[0242] In one alternative implementation, to achieve flexibility in media processing templates, users can upload user-defined AI models. In the media processing template creation implementation on the cloud management platform 10, users can select a corresponding AI model from uploaded user-defined AI models, user-defined AI models uploaded by other users, and pre-configured AI models in the cloud platform. The cloud management platform 10 then creates a media processing template based on the user-selected AI model.
[0243] The following example, with reference to Figure 5, illustrates a user-uploaded, user-defined AI model.
[0244] As shown in Figure 5, Figure 5 is a schematic diagram of the upload process of a user-defined AI model provided in the embodiment of this application. The upload process of the user-defined AI model shown includes steps S510 to S520.
[0245] S510, Cloud Management Platform 10 receives model upload request.
[0246] Corresponding to the S310 process, client 20 sends a model upload request to cloud management platform 10.
[0247] The model upload request is used to indicate the resource specifications required for the candidate AI model and the files needed to execute the candidate AI model.
[0248] The resource specifications include the hardware specifications and computing power specifications for executing candidate AI models.
[0249] The files used to execute the candidate AI model include, but are not limited to: the code file of the candidate AI model, the framework for executing the candidate AI model, environment parameters, and the API library used by the candidate AI model.
[0250] In the first alternative implementation, the model upload request includes: the resource specifications required for the candidate AI model and the files for executing the candidate AI model.
[0251] The client can package the code files, development framework, environment, and API libraries used by the development framework of a candidate AI model to generate a file for executing the candidate AI model. During the packaging process, the resource specifications required for executing the candidate AI model are defined. Based on the file for executing the candidate AI model and the required resource specifications, the client sends a model upload request to the cloud platform. This model upload request includes fields for the required resource specifications of the candidate AI model and the file for executing the candidate AI model.
[0252] For example, the client sends a model upload request to the cloud platform based on the resource specifications required for executing the candidate AI model as input by the user through the interface, and the file to be uploaded for executing the candidate AI model. Specifically, please refer to the embodiment shown in Figure 15B below; this application will not elaborate further on it.
[0253] In the second alternative implementation, the model upload request includes: the resource specifications required by the candidate AI model and the file access path for executing the candidate AI model.
[0254] The client obtains the file access path of the candidate AI model and the resource specifications required to execute the candidate AI model. Based on the file access path of the candidate AI model and the resource specifications required to execute the candidate AI model, the client sends a model upload request to the cloud platform. This model upload request carries fields for the resource specifications required by the candidate AI model and the file access path.
[0255] For example, the client sends a model upload request to the cloud platform based on the resource specifications required to execute the candidate AI model and the file access path for executing the candidate AI model, as input by the user through the interface. Specifically, please refer to the embodiment shown in Figure 15A below; this application will not elaborate further on it.
[0256] In one optional example, "file access path" can include: the storage address of the candidate AI model file on the cloud platform, the access address of the candidate AI model file on the client's device, or the access address of the candidate AI model file on other servers or other cloud platforms. Here, "other servers" can refer to servers storing the candidate AI model files, such as GitHub servers; or other servers can refer to servers used to provide multiple AI models, such as TensorFlow Hub servers or PyTorch Hub servers. "Other cloud platforms" can refer to cloud platforms that provide AI model storage functionality, or they can refer to cloud platforms that provide AI model creation, development, or training services.
[0257] It should be noted that the two optional implementation methods described above are merely alternative methods for the model upload request provided in this application. In other embodiments, the model upload request may have other forms, which are not limited in this application.
[0258] S520, the cloud management platform 10 responds to model upload requests and stores candidate AI models as user-defined AI models.
[0259] In one alternative implementation, the cloud management platform 10 can obtain the resource specifications of the candidate AI model based on the model upload request and obtain the candidate AI model, store the candidate AI model as a user-defined AI model in the cloud platform's storage system, and record the resource specifications of the candidate AI model.
[0260] For example, the cloud management platform 10 reads the file of the candidate AI model to be executed based on the model upload request, and stores the file of the candidate AI model to the cloud platform's storage system.
[0261] Furthermore, in some implementations, to improve the stability of user-defined AI models and thus ensure the stability of media processing templates, the cloud management platform 10 verifies the candidate AI models after acquiring them. This model verification improves the stability of user-defined AI models.
[0262] The following is an example of how the cloud management platform 10 verifies candidate AI models.
[0263] In one optional implementation, the cloud management platform 10 stores test cases and verifies candidate AI models based on these test cases. If the candidate AI model passes verification, the cloud management platform 10 stores it as a user-defined AI model. If the candidate AI model fails verification, the cloud management platform 10 sends a message to the client indicating that the model upload failed, prompting the client to re-upload the candidate AI model.
[0264] In this context, test cases refer to media test data. This media test data can be images, videos, or audio / video files, etc. This application does not impose any limitations on it.
[0265] In one optional example, "candidate AI model verification passed" can include: the candidate AI model runs normally or the candidate AI model can handle test cases. Here, "candidate AI model runs normally" can mean that the cloud platform did not encounter events such as file read failure, model execution interruption, or model runtime exceeding a time threshold during the execution of the candidate AI model. "Candidate AI model handle test cases" means that after test cases are input into the candidate AI model, output results are obtained.
[0266] In one optional example, if the candidate AI model fails validation, the cloud management platform 10 can also return a validation report of the candidate AI model to the client. This validation report includes, but is not limited to: the reasons for the candidate AI model's failure to validate, the code statements containing the anomalies, and suggested modifications.
[0267] The above primarily illustrates the use of test cases by the cloud management platform 10 to verify candidate AI models. In other implementations, the cloud management platform 10 can also use user-uploaded media data test sets to verify candidate AI models. The following describes the implementation method of the cloud management platform 10 using user-uploaded media data test sets to verify candidate AI models.
[0268] In one alternative implementation, the media data test set may include media data in various encoding formats or media data at various resolutions.
[0269] In one alternative implementation, the client can upload a media data test set to the cloud platform when uploading the AI model. For example, the media data test set can also be included in the model upload request.
[0270] Alternatively, in another implementation, after acquiring the candidate AI model uploaded by the user, the cloud management platform 10 sends a verification request to the client. The client responds to the verification request by sending a media data test set to the cloud platform.
[0271] Based on the embodiment provided in Figure 5, users can upload custom AI models to increase the richness and quantity of AI models in the cloud platform, thereby enhancing the diversity and flexibility of AI processing in the media processing process.
[0272] In one alternative implementation, after storing the candidate AI models as user-defined AI models, the cloud management platform 10 can provide a variety of available AI models to the client, allowing the user to select an AI model from the available models through the client. The cloud management platform 10 then creates a media processing template based on the user-selected AI model.
[0273] For example, cloud management platform 10 provides a client with an interface displaying identifiers for multiple available AI models. The user selects one or more first identifiers based on the available AI model identifiers displayed in the interface, and the client sends an algorithm configuration request to cloud management platform 10 based on the user-selected one or more first identifiers. Cloud management platform 10 creates a media processing template according to the algorithm configuration request. Specifically, embodiments for selecting AI models based on the interface can be found in the embodiments provided in Figures 10 and 11 below, which will not be described in detail here.
[0274] The above mainly explains the uploading method of user-defined AI models and the creation method of media processing templates. The following, with reference to Figures 6 and 7, explains how the cloud management platform 10 processes the first media data.
[0275] In one alternative implementation, the cloud management platform 10 can select a target transcoding node from the cluster 110, and the target transcoding node can process the first media data according to the media processing template to obtain the second media data.
[0276] The following explains how to select the target transcoding node.
[0277] In one optional implementation, the target transcoding node can be an idle transcoding node in the cluster. The cloud management platform 10 queries one or more candidate transcoding nodes currently in an idle state in cluster 110. Based on one or more candidate transcoding nodes, the target transcoding node is determined.
[0278] In the first example, cluster 110 has a candidate transcoding node, and cloud management platform 10 uses this candidate transcoding node as the target transcoding node.
[0279] In the second example, cluster 110 has multiple candidate transcoding nodes, and cloud management platform 10 randomly selects one candidate transcoding node from the multiple candidate transcoding nodes as the target transcoding node.
[0280] The two examples described above are merely different implementations of determining the target transcoding node from idle candidate transcoding nodes. In other embodiments, determining the target transcoding node from idle candidate transcoding nodes can also be implemented in other ways. For example, the cloud management platform 10 selects one or more target transcoding nodes from multiple candidate transcoding nodes in a load-balancing manner. This application does not limit this approach.
[0281] In another optional implementation, the cloud management platform 10 can also select a target transcoding node from multiple transcoding nodes in the cluster according to the resource specifications of the media processing template. As shown in Figure 6, which is a schematic diagram of the process for selecting a target transcoding node provided in this embodiment of the application, the process for selecting a target transcoding node includes steps S341 to S343.
[0282] S341, Cloud Management Platform 10 obtains the resource specifications of the media processing template.
[0283] In one alternative implementation, the resource specifications of the media processing template are related to the resource specifications of the AI model included in the media processing template.
[0284] In the first optional example, the resource specification of the media processing template is the maximum resource specification among the AI models included in the media processing template. For example, if the media processing template includes Algorithm 1 and Algorithm 2, and the resource specification of Algorithm 1 is (32vU, 64GB) and the resource specification of Algorithm 2 is (32vU, 32GB), then the resource specification of the media processing template is (32vU, 64GB).
[0285] In the second optional example, the resource specification of the media processing template is the product of a coefficient and the maximum resource specification of the AI model included in the media processing template. The coefficient is a real number greater than 1. This application does not limit the specific value of the coefficient; for example, the coefficient can be 1.5, 2, etc.
[0286] The two optional examples described above are merely alternative ways to obtain the resource specifications of the media processing template based on the resource specifications of the AI model. In other embodiments, the resource specifications of the media processing template may also be related to the number of concurrent media processing tasks and the resource specifications of the AI model included in the media processing template.
[0287] The number of concurrent media processing tasks can be set by the user when configuring the media processing template, or it can be configured by the cloud management platform 10 itself. This application does not limit this.
[0288] In the first optional example, the number of concurrent media processing tasks is 1, meaning the media processing template handles 1 media processing task. The resource specification of the media processing template is the maximum resource specification among the AI models included in the media processing template.
[0289] In the second optional example, the number of concurrent media processing tasks is multiple, that is, multiple media processing tasks are processed concurrently using a media processing template. The resource specification of the media processing template is the product of the number of concurrent media processing tasks and the maximum resource specification of the AI model included in the media processing template.
[0290] Furthermore, in other embodiments, a media processing template may have multiple resource specifications, and the same media processing template can concurrently process different numbers of media processing tasks under different resource specifications. The cloud management platform 10 can select the appropriate resource specification from the multiple resource specifications of the media processing template according to the number of media processing tasks that need to be processed concurrently.
[0291] The various resource specifications of the media processing template can be set by the user or configured by the cloud management platform 10. This application does not impose any limitations on this.
[0292] For example, the resource specifications of media processing template XX1 include resource specification 1 and resource specification 2, where resource specification 2 is greater than resource specification 1. When the resource specification of media processing template XX1 is resource specification 1, it can handle one media processing task. When the resource specification of media processing template XX1 is resource specification 2, it can handle two media processing tasks. When the number of media processing tasks requiring concurrent processing is one, the cloud management platform 10 uses resource specification 1 as the resource specification for media processing template XX1. When the number of media processing tasks requiring concurrent processing is two, the cloud management platform 10 uses resource specification 2 as the resource specification for media processing template XX1.
[0293] S342, the cloud management platform 10 selects the target transcoding node to perform media tasks from multiple transcoding nodes based on the resource specifications of the media processing template.
[0294] In some optional implementations, the "target transcoding node for performing media tasks" can refer to a transcoding node whose specifications are greater than or equal to those of the media processing template. As shown in Figure 6, the cloud management platform 10 will select transcoding node 111 from transcoding node 111 and transcoding interface 112, and use transcoding node 111 as the target transcoding node.
[0295] In a first alternative implementation, the cloud management platform 10 can select a candidate transcoding node with a resource specification greater than or equal to the media processing template from multiple candidate transcoding nodes in the cluster 110 that are in an idle state, and use it as the target transcoding node.
[0296] In a second optional implementation, if the specifications of multiple candidate transcoding nodes in the cluster that are idle are smaller than the resource specifications of the media processing template, the cloud management platform 10 can select a target transcoding node from the multiple transcoding nodes in the cluster through load balancing. For example, from the transcoding nodes whose specifications are greater than or equal to the resource specifications of the media processing template, the transcoding node with the smallest number of media processing tasks executed or the smallest load requirement of the media processing tasks executed can be selected as the target transcoding node.
[0297] The two implementation methods described above are merely different ways of selecting the target transcoding node to perform the media task. In other embodiments, there may be other ways to select the target transcoding node to perform the media task. For example, the cloud management platform 10 selects multiple target transcoding nodes from multiple transcoding nodes, and these multiple target transcoding nodes distribute and process the media processing task. This application does not limit this.
[0298] S343, the cloud management platform 10 schedules the media processing task to the target transcoding node, and controls the target transcoding node to process the first media data according to the media processing template to obtain the second media data.
[0299] In one optional implementation, the cloud management platform 10 schedules media processing tasks to target transcoding nodes through task scheduling. The cloud management platform 10 controls the target transcoding nodes to execute the embodiment provided in Figure 7 below, processing the first media data to obtain the second media data. As shown in Figure 6, the cloud management platform 10 schedules the media processing tasks to transcoding nodes 111 in the cluster 110, and controls the transcoding nodes 111 to process the first media data according to the media processing template to obtain the second media data.
[0300] Taking transcoding node 111 as an example, please refer to Figure 7. Figure 7 is a schematic diagram of the process of the target transcoding node performing media processing tasks according to the embodiment of this application. The process of performing media processing tasks shown includes steps S3431 to S3434.
[0301] S3431, transcoding node 111 obtains the first media data.
[0302] In one alternative implementation, the transcoding node 111 may obtain the first media data by referring to the implementation method of cloud management platform 10 obtaining the first media data in S340 above. This application does not limit this.
[0303] In addition, in some other implementations, transcoding node 111 can receive first media data sent by cloud management platform 10.
[0304] S3432, transcoding node 111 decodes the first media data to obtain the first intermediate media data.
[0305] Here, the first intermediate media data refers to the decoded first media data.
[0306] In one alternative implementation, transcoding node 111 can decode the first media data according to the data format of the first media data. For example, if the data format of the first media data is MP4, transcoding node 111 sequentially performs MP4 encapsulation format decoding and H.264 video decoding on the first media data to obtain the first intermediate media data.
[0307] S3433, transcoding node 111 calls the AI model included in the media processing template to process the first intermediate media data and obtain the second intermediate media data.
[0308] The second intermediate media data is media data processed by an AI model.
[0309] In the first optional implementation, the AI model is a super-resolution model. The second intermediate media data can be the first intermediate media data after resolution enhancement.
[0310] In the second alternative implementation, the AI model is a style transfer model. The second intermediate media data can be the style-transferred first intermediate media data. For example, if the style transfer model is used to convert an image to a cartoon style, then the second intermediate media data is cartoon-style media data.
[0311] In a third alternative implementation, the AI model can be a content generation model. The second intermediate media data can be media data generated based on the first intermediate media data. For example, if the content generation model is used to add text to an image, then the second intermediate media data is the first intermediate media data after the text has been added to the image.
[0312] The three optional implementation methods mentioned above are merely different options for processing the first intermediate media data under different AI models. In other implementation methods, there may be other ways to process the first intermediate media data, which this application does not limit.
[0313] The following example illustrates how the AI model included in the media processing template called by transcoding node 111 is implemented.
[0314] In the first alternative implementation, transcoding node 111 can invoke the AI model included in the media processing template using a model invocation interface.
[0315] In the second alternative implementation, transcoding node 111 can call the AI model included in the media processing template through the program interface of the AI model.
[0316] For example, the AI model is developed based on PyTorch, and transcoding node 111 calls the Python program interface to call the AI model.
[0317] The two optional implementation methods described above are merely different ways of implementing the AI model included in the media processing template. In other embodiments, the AI model included in the media processing template can also be implemented in other ways, such as by calling a file that executes the AI model. This application does not limit this.
[0318] S3434, transcoding node 111 encodes the second intermediate media data according to the transcoding parameters to obtain the second media data.
[0319] In the first optional implementation, the transcoding parameters include the encoding style, and the transcoding node 111 encodes the second intermediate media data according to the encoding style.
[0320] For example, when the encoding style is H.246, transcoding node 111 encodes the second intermediate media data using H.246.
[0321] For example, if the encoding style is quantization encoding, transcoding node 111 performs quantization encoding on the second intermediate media data.
[0322] In the second optional implementation, the transcoding parameters include encoding style and bitrate, and the transcoding node 111 encodes the second intermediate media data according to the encoding style and bitrate.
[0323] For example, with the encoding style being H.246 and the bitrate being xm / s, transcoding node 111 uses H.246 for the second intermediate media data and encodes the second intermediate media data according to xm / s.
[0324] For example, if the encoding style is quantization encoding and the bitrate is xm / s, transcoding node 111 quantizes the second intermediate media data at a bitrate of xm / s.
[0325] In the third optional implementation, the transcoding parameters include transcoding style, bitrate and resolution, and transcoding node 111 encodes the second intermediate media data according to the encoding style, bitrate and resolution.
[0326] The above three optional implementation methods are only different implementation methods for processing the second intermediate media data under different transcoding parameters. In other implementation methods, the transcoding node 111 can also use other implementation methods to process the second intermediate media data. This application does not limit this.
[0327] It is worth noting that the embodiment shown in Figure 7 above is illustrated using transcoding node 111 as the target transcoding node. In other embodiments, the target transcoding node performing the media processing task can also be other transcoding nodes in cluster 110. Alternatively, there can be multiple target transcoding nodes performing the media processing task, and these multiple target transcoding nodes perform the media processing task in a distributed processing manner. The specific form and number of target transcoding nodes are not limited in the embodiments of this application.
[0328] Based on the embodiment shown in Figure 7, the target transcoding node processes media by calling the AI model included in the media processing template, eliminating the need to build an end-to-end media processing business system. This improves the efficiency of applying the AI model to actual business operations, thereby enhancing the efficiency of the media processing method.
[0329] Based on the embodiment provided in Figure 6, the cloud management platform 10 selects the corresponding target transcoding node according to the resource specifications of the media processing template to ensure the compatibility between the target transcoding node for processing media processing tasks and the resource specifications of the media processing template, thereby ensuring the normal operation of the media processing tasks and thus guaranteeing the reliability of media processing.
[0330] To better illustrate the media data processing method provided in this application, the following uses the cloud management platform 10, which includes an algorithm management unit 121 and a transcoding management unit 122, as an example to explain the implementation process of the media data processing method under offline processing requests and real-time processing requests.
[0331] First, taking offline processing requests as an example, please refer to Figure 8. Figure 8 is a schematic flowchart of the media data processing method provided in this application embodiment. The media data processing flow shown includes steps S801 to S809.
[0332] S801, the algorithm management unit 121 receives the candidate AI model uploaded by the user.
[0333] Corresponding to process S801, the user uploads a candidate AI model to the algorithm management unit 121 through client 20. For example, the user uploads a candidate AI model through client 20 with reference to the embodiments provided in Figures 15A to 16B below; the embodiments of this application will not be described in detail here.
[0334] In one alternative implementation, the algorithm management unit 121 receives candidate AI models uploaded by the user according to the embodiment provided in Figure 5 above. Further details are omitted here.
[0335] S802, the algorithm management unit 121 stores the candidate AI model uploaded by the user as a user-defined AI model.
[0336] In one optional implementation, the algorithm management unit 121 stores the candidate AI model uploaded by the user as a user-defined AI model according to the embodiment provided in Figure 5 above. Further details are omitted here.
[0337] S803, the user sends an algorithm configuration request to the transcoding management unit 122 based on the user-defined AI model.
[0338] S804, the transcoding management unit 122 obtains the algorithm configuration request and creates a media processing template.
[0339] S805, the transcoding management unit 122 receives an offline processing request and creates an offline media processing task.
[0340] In the implementation of S805, the user sends an offline processing request to the transcoding management unit 122 through the client 20.
[0341] S806, the transcoding management unit 122 schedules offline media processing tasks to the target transcoding node based on the resource specifications of the media processing template.
[0342] In one alternative implementation, the transcoding management unit 122 schedules the offline media processing task to the target transcoding node with reference to the embodiment provided in FIG6 above, which will not be described in detail here.
[0343] S807, the target transcoding node reads the first media data from the storage system.
[0344] S808, the target transcoding node processes the first media data according to the media processing template to obtain the second media data.
[0345] In one alternative implementation, the target transcoding node processes the first media data with reference to the embodiment provided in Figure 7 above to obtain the second media data, which will not be described in detail here.
[0346] S809, the target transcoding node sends the second media data to the target terminal.
[0347] For details regarding the target terminal, please refer to the description of the target terminal in S330 above. This application embodiment will not repeat the details here.
[0348] Based on the embodiment shown in Figure 8, when the cloud management platform 10 schedules offline media processing tasks to the target transcoding node, the target transcoding node actively reads the first media data and processes it using a user-configured media processing template. Because the media processing template offers greater flexibility and applicability, the reliability of the media processing results can be guaranteed.
[0349] The following example uses real-time processing requests. Please refer to Figure 9, which is a flowchart illustrating the media data processing method provided in this embodiment. Compared to the media data processing method provided in Figure 8, the media data processing method provided in Figure 9 includes steps S901 to S905 after step S804.
[0350] S901, the transcoding management unit 122 receives a real-time processing request and creates a real-time media processing task.
[0351] In the implementation of S901, the user sends a real-time processing request to the transcoding management unit 122 deployed in the cloud management platform 10 through the client 20.
[0352] S902, the transcoding management unit 122 schedules real-time media processing tasks to the target transcoding node based on the resource specifications of the media processing template.
[0353] In one alternative implementation, the transcoding management unit 122 schedules the real-time media processing task to the target transcoding node with reference to the embodiment provided in FIG6 above, which will not be described in detail here.
[0354] S903, the target transcoding node receives the first media data.
[0355] In the first alternative implementation, client 20 sends first media data to the target transcoding node.
[0356] In the second optional implementation, the client 20 sends the first media data to the cloud management platform 10, and the cloud management platform 10 sends the first media data to the target transcoding node.
[0357] In a third optional implementation, other computing devices send the first media data to the target transcoding node. These other computing devices can be devices that store the media data, such as storage devices or media resource systems; they can also be devices that capture images, video, or audio, such as cameras or microphones. This application does not limit the specific form of the other computing devices.
[0358] S904, the target transcoding node processes the first media data according to the media processing template to obtain the second media data.
[0359] In one alternative implementation, the target transcoding node processes the first media data with reference to the embodiment provided in Figure 7 above to obtain the second media data, which will not be described in detail here.
[0360] S905, the target transcoding node sends the second media data to the target terminal.
[0361] For details regarding the target terminal, please refer to the description of the target terminal in S330 above. This application embodiment will not repeat the details here.
[0362] Based on the embodiment shown in Figure 9, upon receiving the first media data, the target transcoding node executes a media processing task, using a user-configured media processing template to process the first media data, thereby improving the timeliness of media processing. Furthermore, the greater flexibility and applicability of the media processing template ensures the reliability of the media processing results.
[0363] Figures 3 to 9 above mainly illustrate the technical solution provided in this application from the perspective of the interaction between the cloud management platform 10 and the client. The following, with reference to Figures 10 to 16B, provides an exemplary description of the interaction interface between the cloud management platform 10 and the client.
[0364] The following description, in conjunction with Figures 10 to 14, illustrates the interface for obtaining the algorithm configuration request.
[0365] As shown in Figure 10, Figure 10 is a schematic diagram of the interface for obtaining an algorithm configuration request provided in an embodiment of this application. The cloud management platform 10 provides a template configuration interface to the client, as shown in Figure 10(a), which displays a "template configuration" function control. When the user clicks the "template configuration" function control, the client sends a second request to the cloud management platform 10. This second request is used to instruct the cloud management platform 10 to perform media processing template configuration. The cloud management platform 10 responds to the second request and provides a first interface as shown in Figure 10(b). This first interface displays an algorithm input box.
[0366] The user inputs a first set of operations based on the first interface, entering one or more first identifiers into the algorithm input box of the first interface. The cloud management platform 10 responds to the first operation and provides a second interface as shown in Figure 10(c). As shown in Figure 10(c), the first identifiers displayed in the algorithm input box of the second interface include: Algorithm 1 and Algorithm 2, and the second interface also displays a confirmation control and a cancel control. The user clicks the "Confirm" control in the second interface to input a second operation into the cloud management platform 10. The cloud management platform 10 responds to the second operation and obtains an algorithm configuration request.
[0367] In the first alternative example, the cloud management platform 10 can receive an algorithm configuration request sent by the client 20.
[0368] When the user clicks the "Confirm" control on the second interface, the client 20 sends an algorithm configuration request to the cloud management platform 10 based on the transcoding parameters and the first identifier displayed in the algorithm input box on the second interface. The transcoding parameters can be configured by the client 20 itself.
[0369] In the second alternative example, the cloud management platform 10 can generate an algorithm configuration request.
[0370] When the user clicks the "Confirm" control on the second interface, the cloud management platform 10 obtains the transcoding parameters and the first identifier displayed in the algorithm input box on the second interface, and generates an algorithm configuration request based on the transcoding parameters and the first identifier. The transcoding parameters can be sent by the client 20 or configured by the cloud management platform 10 itself; this embodiment does not limit the specific parameters.
[0371] The two examples above are merely alternative methods for the cloud management platform 10 to obtain algorithm configuration requests. In other examples, the algorithm configuration request may have other implementation methods. This application embodiment does not limit this.
[0372] The first group of operations will be illustrated below with reference to Figure 11.
[0373] In one alternative implementation, the first set of operations includes a first sub-operation and a second sub-operation.
[0374] Taking the first sub-operation as an example, we will explain its implementation. Similarly, the implementation of the second sub-operation can be explained by referring to the implementation of the first sub-operation.
[0375] The first sub-operation can be a contact-based operation, such as tapping, long-pressing, swiping, double-tapping, or clicking on controls in the interface. Alternatively, the first sub-operation can be a contactless operation, such as gesture input, physical buttons, or voice input. Alternatively, the first sub-operation can also be a typing operation, such as typing with a mouse or keyboard.
[0376] The user clicks the algorithm input box on the first interface to input a first sub-operation. In response to the first sub-operation, the cloud management platform 10 provides identifiers for multiple available AI models. The user then inputs a second sub-operation based on the displayed identifiers of the multiple available AI models. In response to the second sub-operation, the cloud management platform 10 fills the first identifier selected by the second sub-operation into the algorithm input box, providing the second interface as shown in Figure 10(c).
[0377] The identification of AI models includes the model identifier or the identifier of the model template.
[0378] The following explanation, in conjunction with Figure 11, illustrates how the first identifier is selected.
[0379] In the first optional implementation, as shown in Figure 11(a), the user clicks the algorithm input box on the first interface to input a first sub-operation. The cloud management platform 10 responds to the first sub-operation by providing model identifiers for multiple available AI models on the first interface. As shown in Figure 11(b), the first interface displays model identifiers for multiple available AI models in a drop-down list: Algorithm 1, Algorithm 3, Algorithm 3, and Algorithm 4. The user inputs a second sub-operation, selecting Algorithm 1 and Algorithm 2 as the first identifier from the Algorithms 1, 3, 3, and 4 displayed on the first interface. In response to the second sub-operation, the cloud management platform 10 provides the second interface shown in Figure 10(c).
[0380] In the second optional implementation, compared to the first optional implementation, the cloud management platform 10 responds to the first sub-operation by providing identifiers for multiple model templates on the first interface. As shown in Figure 11(c), the identifiers for multiple model templates are displayed in a drop-down list on the first interface: Model Template 1: Algorithm 1, Algorithm 2; Model Template 2: Algorithm 4, Algorithm 6. The user inputs the second sub-operation, selecting Model Template 1 as the first identifier from Model Template 1 and Model Template 2 displayed on the first interface. In response to the second sub-operation, the cloud management platform 10 provides the second interface shown in Figure 10(c).
[0381] In the third optional implementation method, compared to the two optional implementation methods mentioned above, the cloud management platform 10 responds to the first sub-operation and provides a first sub-interface. The interface displayed on the client 20 jumps from the first interface shown in Figure 11(a) to the first interface shown in Figure 11(d). As shown in Figure 11(d), the cloud management platform 10 responds to the first group of operations and provides a first sub-interface, which displays a confirmation control and AI model identifiers: Algorithm 1, Algorithm 2, Algorithm 3, Algorithm 4, Algorithm 5, Algorithm 6, Algorithm 7, Algorithm 8, Algorithm 9, Algorithm 10, Algorithm 11, and Algorithm 12. The user selects the first identifiers from Algorithm 1, Algorithm 2, Algorithm 3, Algorithm 4, Algorithm 5, Algorithm 6, Algorithm 7, Algorithm 8, Algorithm 9, Algorithm 10, Algorithm 11, and Algorithm 12: Algorithm 1 and Algorithm 2, and clicks the confirmation control to input the second sub-operation. The cloud management platform 10 fills the first identifier selected by the user into the algorithm input box, providing the second interface shown in Figure 10(c).
[0382] It should be noted that the above three optional implementation methods are only different ways for the cloud management platform 10 to respond to the first group of operations. In another embodiment, the cloud management platform 10 may also have other implementation methods for responding to the first group of operations, which are not limited in this application embodiment.
[0383] In addition, in some optional scenarios, users can view the model information of any AI model in the interface (first interface or first sub-interface), which includes the AI model's type, function, application scenario, source, resource specifications, etc.
[0384] For example, taking the first sub-interface shown in Figure 11(d) as an example, as shown in Figure 11(e), when the cursor moves to Algorithm 2, the model information page of Algorithm 2 is displayed in the first sub-interface. The model information page of Algorithm 2 displays: Type: CNN, Function: Cartoon Style, Application Scenarios: Style Transfer, Source: Uploaded by User XX3, Resource Specifications: XXXX. This application does not limit the display position of the model information page in the first sub-interface. When the cursor moves away from Algorithm 2, the cloud management platform 10 closes the model information page of Algorithm 2 in the first sub-interface.
[0385] The above mainly uses the example of configuring the AI model in the media processing template through the interface for illustration. In addition, in some other implementations, the transcoding parameters included in the algorithm configuration request can also be input by the user based on the interface. As shown in Figure 12, Figure 12 is a schematic diagram of the interface for obtaining the algorithm configuration request provided in an embodiment of this application.
[0386] Compared to the interface diagram provided in Figure 10, the interface diagram provided in Figure 12 includes a parameter input box in the first interface provided by the cloud management platform 10, as shown in Figure 12(a). The user inputs a first set of operations based on the first interface, entering one or more first identifiers into the algorithm input box and transcoding parameters into the parameter input box. The cloud management platform 10 responds to the first set of operations and displays the second interface as shown in Figure 12(b). In the second interface shown in Figure 12(b), the first identifiers displayed in the algorithm input box include: Algorithm 1 and Algorithm 2. The parameter input box displays a resolution of 1920*1080 and a bitrate of 4m / s. The second interface also displays a "Confirm" control. The user clicks the "Confirm" control in the second interface to input a second operation into the cloud management platform 10. The cloud management platform 10 responds to the second operation and obtains the algorithm configuration request.
[0387] Figure 12 above is merely a schematic illustration of the interface for sending algorithm configuration requests and does not constitute a limitation on the media data processing method provided in this application. For example, in some other embodiments, the parameter input box of the second interface may also display a second identifier indicating transcoding parameters.
[0388] The first set of operations may include the user inputting a first identifier and the user inputting transcoding parameters. Similar to the first sub-operation described above, the user inputting the first identifier and the user inputting transcoding parameters may also be contact-based, contactless, or keystroke-based operations, and this application does not impose any limitations on these operations.
[0389] The first group of operations will be explained below with reference to Figures 13A to 13C.
[0390] In one alternative implementation, the first set of operations includes a third sub-operation and a fourth sub-operation. The third sub-operation refers to the user input of the first identifier. The fourth sub-operation refers to the user input of transcoding parameters.
[0391] As shown in Figure 13A(a), the user clicks the algorithm input box in the first interface to input the third sub-operation. In response to the third sub-operation, the cloud management platform 10 provides a second sub-interface as shown in Figure 13A(b), which includes an algorithm input box and a parameter input box, with Algorithm 1 and Algorithm 2 displayed in the algorithm input box. As shown in Figure 13A(b), the user clicks the parameter input box in the second sub-interface to input the third sub-operation. The cloud management platform responds to the fourth sub-operation. The transcoding parameters of the fourth sub-operation are filled into the parameter input box of the second sub-interface, providing the second interface as shown in Figure 13A(c).
[0392] The implementation of the input transcoding parameters will be illustrated below with reference to Figure 13B.
[0393] In one alternative implementation, similar to the implementation of selecting the first identifier, the transcoding parameters can also be obtained by selection in the implementation of input transcoding parameters.
[0394] In the first optional implementation, the cloud management platform 10 responds to the second sub-operation and provides a second sub-interface as shown in Figure 13B(a). This second sub-interface includes an algorithm input box and a parameter input box, with Algorithm 1 and Algorithm 2 displayed in the algorithm input box. The user clicks the parameter input box in the second sub-interface to input a fourth sub-operation. The cloud management platform 10 responds to the fourth sub-operation and provides multiple transcoding parameter template identifiers in the second sub-interface. As shown in Figure 13B(b), the second sub-interface displays multiple transcoding parameter template identifiers in a drop-down list: Format 1, Format 2, Format 3, and Format 4. The user selects a transcoding parameter template identifier from the displayed identifiers. The cloud management platform fills the transcoding parameters corresponding to the selected template into the parameter input box, providing the second interface as shown in Figure 12(b).
[0395] In the second optional implementation, compared to the first optional implementation, the user clicks the parameter input box in the second sub-interface to input a fourth sub-operation. The cloud management platform 10 responds to the fourth sub-operation and provides the first parameter configuration interface as shown in Figure 13B(c). The interface displayed on the client 20 jumps from the second sub-interface shown in Figure 13B(a) to the first parameter configuration interface shown in Figure 13B(c). This first parameter configuration interface displays multiple transcoding parameter template identifiers: Format 1, Format 2, Format 3, Format 4, Format 5, Format 6, Format 7, Format 8, Format 9, Format 10, Format 11, and Format 12, as well as a "Select" control. When the user clicks the "Select" control for Format 12 in the first parameter configuration interface shown in Figure 13B(c), the cloud management platform 10 fills the transcoding parameters from the transcoding parameter template corresponding to Format 12 selected by the user into the algorithm input box, providing the second interface as shown in Figure 12(b).
[0396] Additionally, in some optional scenarios, users can view the transcoding parameters corresponding to any transcoding parameter template identifier in the interface (either the second sub-interface or the first parameter configuration interface). For example, taking the first parameter configuration interface shown in Figure 13B(c) as an example, as shown in Figure 13B(d), when the cursor moves to format 12, the parameter page for format 12 is displayed in the first parameter configuration interface. The parameter page for format 12 includes: resolution 1920*1080, bitrate 4m / s. When the cursor moves away from format 12, the cloud management platform 10 closes the parameter page for format 12 in the first parameter configuration interface.
[0397] It should be noted that the two optional implementation methods described above are merely different ways of identifying the input transcoding parameters by selecting a transcoding parameter template. In other embodiments, there may be other ways to input transcoding parameters. For example, users can input transcoding parameters manually. Alternatively, users can input transcoding parameters by modifying the content of the transcoding parameter template provided by the cloud management platform; this is not limited in the embodiments of this application.
[0398] It is worth noting that the embodiments shown in Figures 13A and 13B above are illustrated using the order of the first group of operations: inputting the first identifier → inputting the transcoding parameters. In other embodiments, the first group of operations may also be in the order of inputting transcoding parameters → inputting the first identifier.
[0399] In the case where the first group of operations follows the sequence of inputting transcoding parameters → inputting the first identifier, the first group of operations can include a fifth sub-operation and a sixth sub-operation. As shown in Figure 13C, specifically in Figure 13C(a), the user clicks the parameter input box in the first interface to input the fifth sub-operation. The cloud management platform 10 responds to the fifth sub-operation and provides the third sub-interface shown in Figure 13C(b). As shown in Figure 13C(b), the third sub-interface includes a parameter input box and an algorithm input box, and the parameter input box displays a resolution of 1920*1080 and a bitrate of 4m / s. The user clicks the algorithm input box in the third sub-interface to input the sixth sub-operation. The cloud management platform 10 responds to the sixth sub-operation and provides the second interface shown in Figure 12(b).
[0400] The cloud management platform 10 can respond to the fifth sub-operation by referring to the embodiment provided in Figure 13B above, and provide the third sub-interface as shown in Figure 13C(b), which will not be described in detail here.
[0401] The cloud management platform 10 can respond to the sixth sub-operation with reference to the embodiment provided in Figure 11 above, providing the second interface shown in Figure 13C(c). Further details are omitted here.
[0402] Furthermore, Figures 12 to 13C above are merely examples of interactive interfaces with different input first identifiers and transcoding parameters. In other embodiments, there may be other interactive interfaces for inputting the first identifier and transcoding parameters. This application does not limit this.
[0403] For example, as shown in Figure 14, the cloud management platform 10 provides a first interface as shown in Figure 14(a). This first interface includes an algorithm input box. Following the embodiment shown in Figure 11, the cloud management platform 10 interacts with the user, providing a first intermediate interface as shown in Figure 14(b). This first intermediate interface includes first identifiers displayed in the algorithm input box, including Algorithm 1 and Algorithm 2, and also displays a "Confirm" control. When the user clicks the "Confirm" control in the first intermediate interface, the cloud management platform 10 provides a second intermediate interface as shown in Figure 14(c). The second intermediate interface displays a parameter input box. Following the embodiment shown in Figure 13B, the cloud management platform 10 interacts with the user, providing a third intermediate interface as shown in Figure 14(d). This third intermediate interface includes a parameter input box and a "Confirm" control. The parameter input box displays a resolution of 1920*1080 and a bitrate of 4m / s. When the user clicks the "Confirm" control in the third intermediate interface, the cloud management platform 10 provides a second interface as shown in Figure 12(b).
[0404] As shown in Figures 10 to 14, users can configure the AI model used in the media processing process and the transcoding parameters of the output media data through the interface. This improves the interactivity of the media processing target configuration process and enhances the flexibility and applicability of the media processing template, thereby ensuring the reliability of the media processing results.
[0405] The following examples, with reference to Figures 15A and 16B, illustrate the interface for users to upload user-defined AI models.
[0406] In one alternative implementation, the user sends the file access path of the candidate AI model to the cloud management platform 10 via an interface.
[0407] As shown in Figure 15A, the cloud management platform 10 provides a model upload interface, as shown in Figure 15A(a). This model upload interface displays a "Model Upload" function control. In response to the user's input operation based on the "Model Upload" control, the client sends a third request to the cloud management platform 10. This third request is used to request the cloud management platform 10 to upload a user-defined AI model. The cloud management platform 10 responds to the third request and provides a third interface as shown in Figure 15A(b). This third interface includes: a model identifier input box, a model file path input box, and a resource specification input box. Based on the third interface, the user inputs a second set of operations: entering the model identifier of the candidate AI model in the model identifier input box, entering the file access path of the candidate AI model in the model file path input box, and entering the resource specifications of the candidate AI model in the resource specification input box. The cloud management platform 10 responds to the second set of operations and provides a fourth interface as shown in Figure 15A(c). This fourth interface includes: a cancel control, an "Upload Model" control, a model identifier input box, a model file path input box, and a resource specification input box. The model identifier input box displays Algorithm 1. The model file path input box displays the file access path: XXX1:\0415\2408255\Algorithm1. The resource specifications input box displays the hardware specifications requirements: XX, hardware computing power requirements: CCC. When the user clicks "Upload Model" on the fourth interface and enters the third operation, the cloud management platform 10 responds to the third operation and obtains the model upload request.
[0408] In an optional example, the second set of operations includes an identifier input operation, a file access path input operation, and a resource specification input operation. The identifier input operation indicates the model identifier of the AI model uploaded by the user of cloud management platform 10. The file access path input operation indicates the file access path of the AI model uploaded by the user of cloud management platform 10. The resource specification input operation indicates the resource specifications of the AI model uploaded by the user of cloud management platform 10.
[0409] In this embodiment of the application, the order of the identifier input operation, the file access path input operation, and the resource specification input operation is not limited.
[0410] In addition, in some other implementations, users upload candidate AI model files to the cloud management platform 10 via an interface.
[0411] As shown in Figure 15B, compared to the embodiment provided in Figure 15A, the cloud management platform 10 provides a third interface as shown in Figure 15B(a) in the interaction diagram provided in Figure 15B. This third interface includes: a model identifier input box, a "Click to Upload" control, and a resource specification input box. The user inputs a second set of operations based on the third interface: inputting the model identifier of the candidate AI model into the model identifier input box, clicking the "Click to Upload" control to upload the candidate AI model file, and inputting the resource specifications of the candidate AI model into the resource specification input box. The cloud management platform 10 responds to the second set of operations and provides a fourth interface as shown in Figure 15B(b). This fourth interface includes: a model identifier input box, a model file input box, and a resource specification input box. The model identifier input box displays "Algorithm 1". The model file input box displays "File 1". The resource specification input box displays the hardware specifications requirements: (32vU, 64GB) and the hardware computing power requirements: GPU4060. The user clicks "Upload Model" in the fourth interface to input a third operation, and the cloud management platform 10 responds to the third operation by obtaining the model upload request.
[0412] In one optional example, the second set of operations includes an identification input operation, a file upload operation, and a resource specification input operation. The file upload operation instructs the cloud management platform 10 to receive the AI model file uploaded by the user.
[0413] In this embodiment of the application, the order of the identifier input operation, the file upload operation, and the resource specification input operation is not limited.
[0414] In one alternative implementation, the resource specification input operation includes a first input sub-operation and a second input sub-operation. Users can manually input the resource specifications of candidate AI models in a third interface.
[0415] As shown in Figure 16A, specifically in Figure (a), compared to the third interface shown in Figure (b) of Figure 15A, the resource specification input box in the third interface shown in Figure (b) of Figure 16A includes a hardware specification input sub-box and a hardware computing power input sub-box. When the user clicks the hardware specification input sub-box to input a first sub-operation, the cloud management platform 10 responds by filling the corresponding content into the hardware specification input sub-box, providing the first configuration sub-interface shown in Figure (b) of Figure 16A. This first configuration interface includes a hardware specification input sub-box and a hardware computing power input sub-box, and the hardware specification input sub-box displays: (32vU, 64GB). When the user clicks the hardware computing power input sub-box in the first configuration sub-interface to input a second sub-operation, the cloud management platform 10 responds by filling the corresponding content into the hardware computing power input sub-box, providing the second configuration sub-interface shown in Figure (c) of Figure 16A. This second configuration sub-interface includes both a hardware specification input sub-box and a hardware computing power input sub-box. The hardware specifications input box displays: (32vU, 64GB), and the hardware computing power input box displays GPU4060.
[0416] In another alternative implementation, users can select the resource specifications of candidate AI models on a third interface.
[0417] For example, the implementation of resource specifications for user-selected candidate AI models is explained by taking user-configured hardware specifications as an example. The implementation of user-configured hardware computing power can be referred to the implementation of user-configured hardware specifications.
[0418] As shown in Figure 16B, compared to the embodiment provided in Figure 16A, in the embodiment provided in Figure 16B, the user clicks the hardware specification input sub-box in Figure 16A(a) to input a first input sub-operation. The cloud management platform 10 responds to the first input sub-operation and provides a third configuration sub-interface as shown in Figure 16B(a). This third configuration sub-interface displays a "View" control and identifiers for multiple candidate hardware specifications: hardware specification 1, hardware specification 2, hardware specification 3, hardware specification 4, hardware specification 5, and hardware specification 6. The user clicks on the identifier of any candidate hardware specification in the third configuration sub-interface, and the cloud management platform 10 provides the hardware specification information for that candidate hardware specification. As shown in Figure 16B(a), the user clicks on hardware specification 6. The cloud management platform 10 provides a fourth configuration sub-interface as shown in Figure 16B(b), which displays a "Return" control, a "Select" control, and the hardware specifications of hardware specification 6: (32V U, 64GB). When the user clicks the "Select" control to input a selection, the cloud management platform 10 responds by filling the hardware specification 6 into the hardware specification input sub-box, providing the first configuration sub-interface shown in Figure 16B(c). This first configuration interface includes a hardware specification input sub-box and a hardware computing power input sub-box, and the hardware specification input sub-box displays: hardware specification 6.
[0419] As shown in Figures 15A and 16B, allowing users to upload their own custom AI models can increase the richness of AI models available in the media processing process, thereby improving the compatibility between media processing templates and media data, and thus enhancing the reliability of media processing.
[0420] Figures 3 to 16B above mainly illustrate the media data processing method provided in this application by taking the creation of a media processing template when a user requests a media processing service as an example. In some other optional implementations, in order to simplify the media data processing flow, the cloud management platform 10 can provide the client with multiple available candidate media processing templates. When the client requests a media processing service from the cloud platform, the user can determine the target media processing template from the multiple available candidate media processing templates on the client. The cloud management platform 10 processes the media data according to the target media processing template to obtain the processed media data.
[0421] The following three specific examples illustrate the candidate media processing templates provided by the cloud management platform 10.
[0422] In a first optional example, the candidate media processing template can be a media processing template already created by the second client in the cloud management platform 10. For example, the second client sends a second algorithm configuration request to the cloud management platform 10. The cloud management platform 10 receives second algorithm configuration requests from different second clients and executes the above steps S310 to S320 to create candidate media processing templates for different second clients. As another example, the second client requests media processing services from the cloud platform, and the cloud management platform 10 creates candidate media processing templates according to the input operations of the user corresponding to the second client.
[0423] This application does not limit the specific form of the second client. For example, the second client and the client requesting the media service may be deployed on different devices, or the second client and the client requesting the media service may correspond to different users.
[0424] In the second alternative example, the candidate media processing templates can be pre-created in the cloud management platform 10 by the client requesting the media processing service. For example, when a client logs into the cloud platform for the first time, or when a client registers or purchases a media processing service, the user creates one or more candidate processing templates in the cloud management platform 10 through the client.
[0425] In a third alternative example, the candidate processing templates can be automatically generated by the cloud management platform 10. For example, a large model is deployed in the cloud platform, and the cloud management platform 10 calls the large model to generate multiple candidate processing templates. The large model can be a large language model or other neural network models with content generation capabilities; this application does not limit this.
[0426] The above three examples are merely optional examples of the candidate media processing templates provided in this application. In other embodiments, the candidate media processing templates may also have other forms. For example, the candidate media template may also be read by the cloud management platform 10 from an external device. This application does not limit this.
[0427] The following description, in conjunction with Figure 17, illustrates another implementation of the media data processing method. Figure 17 is a flowchart illustrating the media data processing method provided in this embodiment. This media data processing method includes steps S171 to S173.
[0428] S171, the cloud management platform 10 receives the fourth request from the client.
[0429] Similar to the first request described above, the fourth request can be used to instruct the processing of third media data. In an alternative implementation, the fourth request instructs the third media data in a manner similar to the first request instructing the first media data, which will not be elaborated upon here.
[0430] S172, the cloud management platform 10 determines the target media processing template from multiple candidate media processing templates.
[0431] The following provides two optional implementation methods to illustrate how to determine the target media processing template.
[0432] In the first alternative implementation, the target media processing template is selected by the user.
[0433] For example, in response to the fourth request, cloud management platform 10 provides a template selection interface to the user. The template selection interface displays template identifiers for multiple candidate media processing templates, along with the media processing flow included in each candidate template. The user selects a target template identifier from the multiple candidate media processing template identifiers. Based on the target template identifier selected by the user, the client sends a fifth request to cloud management platform 10. Cloud management platform 10 responds to the fifth request and selects a target media processing template from the candidate media processing templates that matches the target template identifier.
[0434] For example, the fourth request carries a third identifier. This third identifier is used to indicate the target template selected by the user. The cloud management platform 10 responds to the fourth request and selects the target media processing template that matches the third identifier from the candidate media processing templates.
[0435] In the second optional implementation, the target media processing template is selected automatically by the cloud management platform 10.
[0436] For example, the cloud management platform 10 can select a target media processing template from multiple candidate media processing templates based on the scene information of the third media data, where the scene information of the third media data is matched.
[0437] The two optional implementation methods described above are merely different ways of determining the target media processing template. In other embodiments, determining the target media processing template may also have other implementation methods. For example, the user selects an initial media processing template from multiple candidate media processing templates. The user modifies the AI model included in the initial media processing template, the execution logic between AI models, or the transcoding parameters to obtain the target media processing template. This application does not limit this.
[0438] S173, the cloud management platform 10 processes the third media data according to the target media processing template to obtain the fourth media data.
[0439] In an alternative implementation, the cloud management platform 10 may process the third media data with reference to the embodiments described in S340 above or Figure 6. Further details are omitted here.
[0440] Based on the embodiment shown in Figure 17, in the implementation of a user requesting a media processing service, the cloud management platform 10 provides the user with media processing templates created by other users, enabling the user to determine the target media processing template from multiple candidate media processing templates. This allows for the reuse of multiple media processing templates, thereby improving media processing efficiency.
[0441] The above description primarily focuses on the interaction between the cloud management platform 10 and the client, outlining the technical solutions provided in this application. It is understood that the cloud management platform 10 includes the corresponding hardware structures and / or software modules for executing various functions. Those skilled in the art will readily recognize that, based on the algorithm steps of the examples described in the embodiments disclosed in this application, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0442] This application embodiment can group the cloud management platform 10 into functional modules according to the above method example. For example, each functional group can correspond to a functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the naming and grouping of modules in this application embodiment are illustrative and only represent one logical functional grouping. In actual implementation, there may be other grouping methods.
[0443] For example, the media data processing system 120 deployed in the cloud management platform 10 can also be named the media data processing system 18. As shown in Figure 18, which is a schematic diagram of the structure of the media data processing system provided in this embodiment of the application, the media data processing system 18 includes a communication module 181, a processing module 182, and a storage module 183.
[0444] The communication module 181 is used to acquire an algorithm configuration request and receive a first request from the client. The first request instructs the processing of first media data. The algorithm configuration request includes one or more first identifiers indicating the AI model and transcoding parameters. These transcoding parameters include one or more of the following: encoding style, bitrate, and resolution. The AI models corresponding to the multiple first identifiers include: AI models preset in the cloud platform and user-defined AI models. For example, the communication module 181 is used to execute S310 and S330 in Figure 3 above.
[0445] Processing module 182 is used to obtain a media processing template according to an algorithm configuration request, and to process the first media data according to the media processing template to obtain the second media data. For example, communication module 181 is used to execute S320 and S340 in Figure 3 above.
[0446] The storage module 183 is used to store executable program code, data from the media data processing system 18 during the media processing process, such as AI models preset in the cloud platform, user-defined AI models, transcoding parameters, first media data, and second media data.
[0447] The communication module 181, processing module 182, and storage module 183 can be implemented in software or in hardware. For example, the implementation of processing module 182 will be described below. Similarly, the implementation of communication module 181 and storage module 183 can refer to the implementation of processing module 182.
[0448] As an example of a software functional unit, processing module 182 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the aforementioned computing instance may be one or more. For example, processing module 182 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.
[0449] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.
[0450] As an example of a hardware functional unit, the processing module 182 may include at least one computing device, such as a server. Alternatively, the processing module 182 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.
[0451] The processing module 182 includes multiple computing devices that can be distributed within the same region or in different regions. Similarly, the communication module 181 includes multiple computing devices that can be distributed within the same Availability Zone (AZ) or in different AZs. Likewise, the processing module 182 includes multiple computing devices that can be distributed within the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0452] It should be noted that, in other embodiments, the processing module 182 can be used to execute any step in the media data processing method. The communication module 181 can be used to execute any step in the media data processing method. The storage module 183 can be used to execute any step in the media data processing method. The communication module 181, processing module 182, and storage module 183 can all be used to execute any step in the media data processing method. The steps implemented by the communication module 181, processing module 182, and storage module 183 can be specified as needed. By implementing different steps in the media data processing method through the communication module 181, processing module 182, and storage module 183, the full functionality of the media data processing system 18 is achieved.
[0453] This application also provides a computing device for performing the above-described media data processing method.
[0454] In one example, the computing device may include a media data processing system 18 as shown in FIG18. The media data processing system 18 includes a communication module 181, a processing module 182, and a storage module 183.
[0455] In another example, as shown in Figure 19, computing device 19 includes a bus 192, a processor 194, a memory 196, and a communication interface 198. The processor 194, memory 196, and communication interface 198 communicate via the bus 192. Computing device 19 can be a server or a terminal device. It should be understood that this application does not limit the number of processors 194 and memory 196 in computing device 19.
[0456] Bus 192 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, only one line is used in Figure 19, but this does not imply that there is only one bus or one type of bus. Bus 192 can include pathways for transmitting information between various components of computing device 19 (e.g., memory 196, processor 194, communication interface 198).
[0457] Processor 194 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0458] In this application, the processor 194 can execute the media data processing method provided in FIG3 above. For example, it configures a media processing template based on an algorithm configuration request sent by the client. Upon receiving a first request sent by the client, it processes the first media data indicated by the first request according to the media processing template to obtain second media data that satisfies the client's requirements.
[0459] Memory 196 may include volatile memory, such as random access memory (RAM). Processor 194 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0460] The memory 196 stores executable program code, which the processor 194 executes to implement the functions of the aforementioned communication module 181, processing module 182, and storage module 183, thereby realizing the media data processing method. That is, the memory 196 stores instructions for executing the media data processing method.
[0461] The communication interface 198 uses transceiver modules, such as, but not limited to, network interface cards and transceivers, to enable communication between the computing device 19 and other devices or communication networks.
[0462] The media data processing method disclosed in the above embodiments can be applied to, or implemented by, processor 194. Processor 194 can be an integrated circuit chip with signal processing capabilities.
[0463] In implementation, each step of the above method can be completed by the integrated logic circuits in the hardware of the processor 194 or by instructions in software form. The processor 194 can be a general-purpose processor, including a CPU, a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete vacuum tubes or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly embodied in the execution of the hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 196, and the processor 194 reads the information in memory 196 and completes the steps of the above method in combination with its hardware.
[0464] In one possible implementation, the processor 194 can also be used to execute a media data processing method. For specific implementation, please refer to the embodiments provided above for the media data processing method. The embodiments of this application will not be repeated here.
[0465] In this embodiment of the application, the chip system may be composed of chips or may include chips and other discrete devices.
[0466] This application also provides a computing device cluster 200 for performing the above-described media data processing method.
[0467] In one example, the computing device cluster 200 may include a media data processing system 18 as shown in FIG18, the media data processing system 18 including a communication module 181, a processing module 182 and a storage module 183.
[0468] In another example, as shown in Figure 20, the computing device cluster 200 includes at least one computing device 19 as shown in Figure 19. The computing device 19 includes a bus 192, a processor 194, a memory 196, and a communication interface 198. The processor 194, the memory 196, and the communication interface 198 communicate via the bus 192. The computing device 19 can be a server or a terminal device.
[0469] In one possible implementation, one or more computing devices in the computing device cluster 200 can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 21 illustrates one possible implementation. As shown in Figure 21, two computing devices 19A and 19B are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this type of possible implementation, the memory 196 in computing device 19A stores instructions for executing the functions of the communication module 181. Simultaneously, the memory 196 in computing device 19B stores instructions for executing the functions of the processing module 182 and the storage module 183.
[0470] The connection method between the computing device clusters shown in Figure 21 can be based on the media data processing method provided in this application. The media data processing involves receiving and processing media data streams, requiring the processing of a large amount of data. Therefore, it is considered that the functions implemented by the processing module 182 and the storage module 183 are executed by the computing device 19B, and the functions implemented by the communication module 181 are executed by the computing device 19A.
[0471] It should be understood that the functions of computing device 19A shown in Figure 21 can also be performed by multiple computing devices 19. Similarly, the functions of computing device 19B can also be performed by multiple computing devices 19.
[0472] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform the aforementioned media data processing method.
[0473] For example, when a computer program product is run on at least one computing device, the at least one computing device performs the media data processing method shown in Figure 4.
[0474] This application also provides a computer-readable storage medium. All or part of the processes in the above method embodiments can be implemented by a computer program instructing related hardware. This program can be stored in the computer-readable storage medium, and when executed, it can include the processes of the above method embodiments. The computer-readable storage medium can be a terminal of any of the foregoing embodiments, such as an internal storage unit including a data transmission end and / or a data receiving end, like a hard disk or memory of the terminal. The computer-readable storage medium can also be an external storage device of the terminal, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the terminal. Further, the computer-readable storage medium can include both the internal storage unit and the external storage device of the terminal. The computer-readable storage medium is used to store the computer program and other programs and data required by the terminal. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.
[0475] It should be noted that the terms "first" and "second," etc., in the specification, claims, and drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.
[0476] It should be understood that in this application, "at least one (item)" means one or more, "more than one" means two or more, "at least two (items)" means two or three or more, and "and / or" is used to describe the relationship between related objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the related objects before and after are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0477] It should be understood that in the embodiments of this application, "B corresponding to A" means that B is associated with A. For example, B can be determined based on A. It should also be understood that determining B based on A does not mean that B is determined solely based on A; B can also be determined based on A and / or other information. Furthermore, the term "connection" in the embodiments of this application refers to various connection methods, such as direct connection or indirect connection, to achieve communication between devices, and the embodiments of this application do not impose any limitations on this.
[0478] Unless otherwise specified, the term "transmission" in the embodiments of this application refers to bidirectional transmission, encompassing the actions of sending and / or receiving. Specifically, "transmission" in the embodiments of this application includes sending data, receiving data, or both sending and receiving data. In other words, data transmission here includes uplink and / or downlink data transmission. Data may include channels and / or signals; uplink data transmission refers to uplink channel and / or uplink signal transmission, and downlink data transmission refers to downlink channel and / or downlink signal transmission.
[0479] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for processing media data, characterized in that, The method is applied to a cloud management platform, and the method includes: Obtain an algorithm configuration request, which includes: transcoding parameters and one or more first identifiers for indicating the AI model; the transcoding parameters include one or more of the following: encoding style, bitrate, and resolution; the AI models corresponding to the multiple first identifiers include: AI models preset in the cloud management platform and user-defined AI models; The media processing template is requested according to the algorithm configuration request; Receive a first request from the client; the first request is used to instruct the processing of first media data; The first media data is processed according to the media processing template to obtain the second media data.
2. The method according to claim 1, characterized in that, The request to obtain the algorithm configuration includes: In response to a second request from the client, a first interface is provided to the client; the second request is used by the cloud management platform to configure media processing templates; the first interface includes: an algorithm input box; In response to a user's first set of operations on the first interface, a second interface is provided; in the second interface, the algorithm input box includes the one or more first identifiers; In response to the user's second operation on the second interface, the algorithm configuration request is obtained.
3. The method according to claim 2, characterized in that, The first group of operations includes a first sub-operation and a second sub-operation: The provision of a second interface in response to a user's first set of operations on the first interface includes: In response to the user's first sub-operation on the algorithm input box in the first interface, a first sub-interface is provided; the first sub-interface includes the identifiers of multiple AI models; the multiple AI models include: AI models preset in the cloud management platform, and user-defined AI models uploaded by different users; In response to a second sub-operation by the user on the identification of the plurality of AI models in the first sub-interface, the second interface is provided.
4. The method according to claim 2 or 3, characterized in that, The first interface also includes a parameter input box; in the second interface, the parameter input box includes the transcoding parameter or a second identifier for indicating the transcoding parameter.
5. The method according to claim 4, characterized in that, The first group of operations includes a third sub-operation and a fourth sub-operation; or the first group of operations includes a fifth sub-operation and a sixth sub-operation. The provision of a second interface in response to a user's first set of operations on the first interface includes: In response to the user's third sub-operation on the algorithm input box in the first interface, a second sub-interface is provided; The second sub-interface includes the parameter input box and the algorithm input box; the algorithm input box includes one or more first identifiers; In response to the user's fourth sub-operation on the parameter input box in the second sub-interface, the second interface is provided; Alternatively, in response to the user's fifth sub-operation on the parameter input box in the first interface, a third sub-interface is provided, the third sub-interface including the parameter input box and the algorithm input box; the parameter input box includes the second identifier or the transcoding parameter; In response to the user's sixth sub-operation on the algorithm input box in the third sub-interface, the second interface is provided.
6. The method according to any one of claims 1 to 5, characterized in that, Before obtaining the algorithm configuration request, the method further includes: Obtain the model upload request from the client, the model upload request including: the resource specifications required by the candidate AI model and the file access path for executing the candidate AI model; In response to the model upload request, the candidate AI model is stored as a user-defined AI model.
7. The method according to claim 6, characterized in that, The step of obtaining the model upload request from the client includes: In response to a third request from the client, a third interface is provided to the client; the third request is used to request the upload of a user-defined AI model; the third interface includes: a model file path input box and a resource specification input box; In response to the user's second set of operations on the third interface, a fourth interface is provided; in the fourth interface, the model file path input box includes the file access path of the candidate AI model, and the resource specification input box includes the resource specifications required by the candidate AI model; In response to the user's third operation on the fourth interface, the model upload request is obtained.
8. The method according to claim 6 or 7, characterized in that, The cloud management platform also stores test cases or media data test sets; the media data test sets are uploaded by the client. The step of storing the candidate AI model as a user-defined AI model includes: Based on the file access path, obtain the candidate AI model; The candidate AI model is validated based on the test cases or the media data test set. If the candidate AI model passes the verification, the candidate AI model is stored as the user-defined AI model.
9. The method according to any one of claims 1 to 8, characterized in that, The cloud management platform includes multiple transcoding nodes; the process of processing the first media data according to the media processing template to obtain the second media data includes: Based on the resource specifications of the media processing template, a target transcoding node for performing media tasks is selected from the plurality of transcoding nodes; the specifications of the target transcoding node are greater than or equal to the resource specifications of the media processing template; the resource specifications of the media processing template are related to the resource specifications of the AI model included in the media processing template. The media task is scheduled to the target transcoding node, and the target transcoding node is controlled to process the first media data according to the media processing template to obtain the second media data.
10. The method according to any one of claims 1 to 9, characterized in that, The cloud management platform also includes multiple candidate media processing templates; The candidate media processing template is created based on the algorithm configuration request from the second client; the method further includes: Receive a fourth request from the client; the fourth request instructs that the third media data be processed. A target media processing template is determined from the plurality of candidate media processing templates, and the third media data is processed according to the target media processing template to obtain the fourth media data.
11. A media data processing system, characterized in that, The system is deployed on a cloud management platform, and the system includes: A communication module is used to obtain an algorithm configuration request, the algorithm configuration request including: transcoding parameters and one or more first identifiers for indicating AI models; the transcoding parameters include one or more of the following: encoding style, bitrate, and resolution; the AI models corresponding to the multiple first identifiers include: AI models preset in the cloud management platform and user-defined AI models; The processing module is used to obtain the media processing template according to the algorithm configuration request; The communication module is further configured to receive a first request from a client; the first request instructs that the first media data be processed. The processing module is further configured to process the first media data according to the media processing template to obtain the second media data.
12. The system according to claim 11, characterized in that, The processing module is used for: In response to a second request from the client, a first interface is provided to the client; the second request is used by the cloud management platform to configure media processing templates. The first interface includes: an algorithm input box; In response to a user's first set of operations on the first interface, a second interface is provided; in the second interface, the algorithm input box includes the one or more first identifiers; In response to the user's second operation on the second interface, the algorithm configuration request is obtained.
13. The system according to claim 12, characterized in that, The first group of operations includes a first sub-operation and a second sub-operation; the processing module is used for: In response to the user's first sub-operation on the algorithm input box in the first interface, a first sub-interface is provided; the first sub-interface includes the identifiers of multiple AI models; the multiple AI models include: AI models preset in the cloud management platform 10, and user-defined AI models uploaded by different users; In response to a second sub-operation by the user on the identification of the plurality of AI models in the first sub-interface, the second interface is provided.
14. The system according to claim 11 or 12, characterized in that, The first interface also includes a parameter input box; in the second interface, the parameter input box includes the transcoding parameter or a second identifier for indicating the transcoding parameter.
15. The system according to claim 14, characterized in that, The first group of operations includes a third sub-operation and a fourth sub-operation; or the first group of operations includes a fifth sub-operation and a sixth sub-operation; the processing module is used for: In response to the user's third sub-operation on the algorithm input box in the first interface, a second sub-interface is provided; the second sub-interface includes the parameter input box and the algorithm input box; the algorithm input box includes one or more first identifiers; In response to the user's fourth sub-operation on the parameter input box in the second sub-interface, the second interface is provided; Alternatively, in response to the user's fifth sub-operation on the parameter input box of the first interface, a third sub-interface is provided, the third sub-interface including the parameter input box and the algorithm input box; the parameter input box includes the second identifier or the transcoding parameter; In response to the user's sixth sub-operation on the algorithm input box in the third sub-interface, the second interface is provided.
16. The system according to any one of claims 11 to 15, characterized in that, The processing module is further configured to: Obtain the model upload request from the client, the model upload request including: the resource specifications required by the candidate AI model and the file access path for executing the candidate AI model; In response to the model upload request, the candidate AI model is stored as a user-defined AI model.
17. The system according to claim 16, characterized in that, The processing module is also used for: In response to a third request from the client, a third interface is provided to the client; the third request is used to request the upload of a user-defined AI model; the third interface includes: a model file path input box and a resource specification input box; In response to the user's second set of operations on the third interface, a fourth interface is provided; in the fourth interface, the model file path input box includes the file access path of the candidate AI model, and the resource specification input box includes the resource specifications required by the candidate AI model; In response to the user's third operation on the fourth interface, the model upload request is obtained.
18. The system according to claim 16 or 17, characterized in that, The system includes: The storage module is used to store test cases or media data test sets; the media data test sets are uploaded by the client. The processing module is further configured to obtain the candidate AI model based on the file access path; verify the candidate AI model based on the test cases or the media data test set; and if the candidate AI model passes the verification, store the candidate AI model as the user-defined AI model.
19. The system according to any one of claims 11 to 18, characterized in that, The cloud management platform 10 includes multiple transcoding nodes; the processing module is further used for: Based on the resource specifications of the media processing template, a target transcoding node for performing media tasks is selected from the plurality of transcoding nodes; the specifications of the target transcoding node are greater than or equal to the resource specifications of the media processing template; the resource specifications of the media processing template are related to the resource specifications of the AI model included in the media processing template. The media task is scheduled to the target transcoding node, and the target transcoding node is controlled to process the first media data according to the media processing template to obtain the second media data.
20. The system according to any one of claims 11 to 19, characterized in that, The system includes: A storage module is used to store multiple candidate media processing templates; the candidate media processing templates are created based on the algorithm configuration request from the second client. The processing module is further configured to receive a fourth request from the client; the fourth request is used to instruct the processing of the third media data; and to determine a target media processing template from the plurality of candidate media processing templates, and process the third media data according to the target media processing template to obtain the fourth media data.
21. A computing device cluster, characterized in that, The system includes at least one computing device, each computing device including a processor, the processor of the at least one computing device being configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1 to 10.
22. A computer-readable storage medium, characterized in that, It includes computer program instructions, which, when executed by a cluster of computing devices, perform the method as described in any one of claims 1 to 10.
23. A computer program product containing instructions, characterized in that, When the instruction is executed by the computing device cluster, the computing device cluster performs the method as described in any one of claims 1 to 10.
Citation Information
Patent Citations
Streaming media transcoding method and device, memory medium and terminal device
CN108156485A
Data processing method and device applied to artificial intelligence platform
CN110659134A
Data processing method and server
CN111858041A
AI processing method and device for video stream
CN115086778A
Artificial intelligence and machine learning models management and / or training
WO2024010399A1