Video transcoding method and device for television station and medium

By using a pre-trained video transcoding evaluation model, dynamically selecting appropriate transcoding services for transcoding tasks is solved, and the load differences between cross-service providers in the existing technology are not fully considered, and more efficient and balanced utilization of transcoding resources is achieved.

CN120017913AInactive Publication Date: 2025-05-16浪潮智能终端有限公司
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510465826.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing transcoding systems fail to fully consider real-time load differences across service providers in terms of load balancing, resulting in low resource utilization and unbalanced load.

Method used

The pre-trained video transcoding evaluation model is used to evaluate the performance of each transcoding service based on time consumption, task queueing time and transcoding quality, and dynamically select the most suitable transcoding service for transcoding task allocation.

Benefits of technology

It improves the transcoding efficiency and quality, optimizes resource allocation, enhances load balancing capabilities, and avoids the problems of resource waste and uneven load of a single service provider.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017913A_ABST
    Figure CN120017913A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of video transcoding, and discloses a video transcoding method and device for a television station and a medium, and the method comprises the steps: obtaining an on-demand video of the television station, and setting the on-demand video as a to-be-transcoded video; based on a pre-trained video transcoding evaluation model, obtaining a transcoding score of each transcoding service on the to-be-transcoded video, the video transcoding evaluation model being a neural network model, the evaluation score being determined based on time consumption evaluation, task queuing duration evaluation and transcoding quality evaluation of each transcoding service; determining a specified transcoding service of the video to be transcoded based on the transcoding score; and carrying out a transcoding task on the to-be-transcoded video based on the specified transcoding service to obtain a transcoded video adapted to be played by the television station.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video transcoding technology, and in particular to a method, device and medium for video transcoding of a television station. Background Art

[0002] As the broadcasting and television industry develops towards multi-terminal and high-definition, video on demand services have put forward higher requirements for the real-time, adaptability and resource utilization of transcoding technology. TV stations usually need to transcode the original video to generate multi-bitrate, multi-format adaptive streaming media, and inject it into the content delivery network (CDN) to ensure smooth terminal playback. In order to improve transcoding efficiency and system disaster recovery capabilities, the industry generally adopts a multi-transcoding service provider collaboration model.

[0003] Although the current transcoding system has introduced distributed architecture (such as task allocation based on the Hadoop framework) or cloud computing technology (such as the MapReduce model) to improve the throughput of a single service provider cluster, its load balancing mechanism only performs simple polling or static priority allocation for the internal node resources (such as CPU and memory usage) of a single service provider, and does not fully consider the real-time load differences across service providers. Summary of the invention

[0004] One or more embodiments of the present specification provide a video transcoding method, device, and medium for a television station, which are used to solve the technical problems raised by the background technology.

[0005] One or more embodiments of this specification adopt the following technical solutions: One or more embodiments of this specification provide a video transcoding method for a television station, the method comprising: Obtaining a TV station's on-demand video, and setting the on-demand video as a video to be transcoded; Based on a pre-trained video transcoding evaluation model, a transcoding score of each transcoding service for the video to be transcoded is obtained, wherein the video transcoding evaluation model is a neural network model, and the evaluation score is determined based on a time consumption evaluation, a task queuing time evaluation, and a transcoding quality evaluation of each transcoding service; Determining a designated transcoding service for the video to be transcoded based on the transcoding score; The video to be transcoded is transcoded based on the designated transcoding service to obtain a transcoded video suitable for broadcasting by the TV station.

[0006] It should be noted that the embodiments of this specification have the following beneficial effects through the above contents: Improve transcoding efficiency: Through the pre-trained video transcoding evaluation model, the most suitable transcoding service can be quickly selected for the video to be transcoded, thereby reducing unnecessary transcoding attempts and improving transcoding efficiency.

[0007] Optimize resource allocation: The decision-making mechanism based on transcoding scores can allocate transcoding tasks more reasonably, making resources (such as CPU, memory, etc.) more effectively utilized and avoiding resource waste.

[0008] Improve transcoding quality: By comprehensively considering factors such as time consumption, task queue time and transcoding quality, the final output transcoded video is ensured to be not only efficient but also of high quality.

[0009] Enhanced load balancing: This method can take into account the real-time load differences across service providers, avoiding the problem of uneven load within a single service provider, thereby achieving more effective load balancing.

[0010] Furthermore, before obtaining the transcoding score of each transcoding service for the video to be transcoded, the method further includes: Constructing an initial video transcoding evaluation model, the initial video transcoding evaluation model includes a first neural network model corresponding to the time consumption evaluation of each transcoding service, a second neural network model for evaluating the task queuing time of each transcoding service, and a third neural network model for evaluating the transcoding quality of each transcoding service; The initial video transcoding evaluation model is trained based on the training set to obtain a video transcoding evaluation model that meets the requirements.

[0011] It should be noted that the embodiments of this specification have the following beneficial effects through the above contents: Improving evaluation accuracy: By building three neural network models that include time consumption evaluation, task queue duration evaluation, and transcoding quality evaluation, multiple aspects of video transcoding can be accurately evaluated, thereby improving the accuracy of the transcoding score.

[0012] Model generalization ability: By using the training set to train the model, the model can learn the characteristics and rules of various transcoding services, improving the generalization ability of the model and enabling it to be applied to a wider range of video transcoding scenarios.

[0013] Fast decision support: The trained evaluation model can quickly provide transcoding scores for the videos to be transcoded, which speeds up the decision-making process and reduces the waiting time for transcoding services.

[0014] Furthermore, the pre-trained video transcoding evaluation model is used to obtain the transcoding score of each transcoding service for the video to be transcoded, including: Based on the video transcoding evaluation model, obtaining a first score of the first neural network model corresponding to each of the transcoding services, a second score of the second neural network model corresponding to each of the transcoding services, and a third score of the third neural network model corresponding to each of the transcoding services; The first score, the second score, and the third score are input into a preset evaluation formula to obtain a transcoding score of each transcoding service for the video to be transcoded.

[0015] It should be noted that the embodiments of this specification have the following beneficial effects through the above contents: Automated decision-making process: By using a neural network model to automatically calculate transcoding scores, the need for manual decision-making is reduced, making the allocation of transcoding tasks more automated.

[0016] Improved evaluation speed: Pre-trained models can quickly process large amounts of data, thereby quickly obtaining transcoding scores for each transcoding service, improving the efficiency of the entire transcoding process.

[0017] Enhanced accuracy: The neural network model is used to conduct a comprehensive analysis of multiple evaluation dimensions (time consumption, task queue time, transcoding quality), which can more accurately predict the performance of the transcoding service.

[0018] Model interpretability: Although the neural network model itself may be difficult to explain, the calculation process of the score can be explained through the pre-set evaluation formula, which helps to understand the basis of the scoring.

[0019] Furthermore, the evaluation formula is: Score = α / Cost+ β / Delay+ γ·Quality; where, Score is the transcoding score, Cost is the first score, Delay is the second score, Quality is the third score, α+β+γ=1, and α, β, and γ are all positive numbers.

[0020] It should be noted that the embodiments of this specification have the following beneficial effects through the above contents: Weighted comprehensive evaluation: This formula introduces weights α, β, γ to weight time consumption, cost, and quality according to the specific needs of the TV station, thereby providing a comprehensive transcoding score.

[0021] Flexible adjustment: Since α, β, and γ can be adjusted independently, TV stations can adjust the weight of each factor according to their own business priorities (such as cost control, real-time requirements or quality standards) to improve the flexibility of evaluation.

[0022] Cost-benefit analysis: By using the inverse of Cost (1 / Cost) as the weight, this formula encourages the selection of more cost-effective transcoding services, which helps reduce operating costs.

[0023] Real-time guarantee: By assigning a weight β to the inverse of Delay (1 / Delay), this formula emphasizes the importance of real-time in transcoding service selection and ensures smooth video playback.

[0024] Quality priority: The Quality term is directly multiplied by γ, indicating that video quality is a key factor in the evaluation, which is conducive to ensuring the quality of the output video.

[0025] Furthermore, after performing the transcoding task on the video to be transcoded based on the designated transcoding service, the method further includes: Obtaining the processing status of the transcoding task by calling the status query interface of the specified transcoding service; If the processing status is that transcoding is completed, feedback that transcoding is completed is sent to the user terminal.

[0026] It should be noted that the embodiments of this specification have the following beneficial effects through the above contents: Real-time status update: By calling the status query interface of the specified transcoding service, the system can obtain the processing status of the transcoding task in real time, thereby providing users with the latest transcoding progress information.

[0027] Improve user experience: When users learn that a transcoding task has been completed, they can receive timely notifications, which improves user satisfaction and experience because users no longer need to wait or actively check the transcoding status.

[0028] Reduce user anxiety: Real-time status feedback reduces user anxiety about possible problems with transcoding tasks, and users can wait without worrying about progress.

[0029] Furthermore, if the processing status is that transcoding is not completed, the method further includes: Determine whether transcoding is abnormal; If so, re-transcoding the video to be transcoded based on the designated transcoding service.

[0030] It should be noted that the embodiments of this specification have the following beneficial effects through the above contents: Improve reliability: By determining whether transcoding is abnormal, the system can ensure the reliability of video transcoding and avoid incomplete transcoding results caused by errors or interruptions.

[0031] Automatic fault recovery: Once a transcoding anomaly is detected, the system automatically attempts to re-transcode, which reduces the need for manual intervention and improves the efficiency of fault recovery.

[0032] Reduce waiting time: Automatically retrying transcoding tasks can reduce the waiting time of users because users do not need to manually resubmit transcoding requests.

[0033] Furthermore, the method further comprises: Determine whether the number of task queues corresponding to the transcoding task is greater than a capacity expansion threshold, where the capacity expansion threshold is set based on a current concurrency capability; If it is determined that the number of task queues corresponding to the transcoding task is greater than the expansion threshold, determine whether the number of cluster transcoding instances has reached a maximum value; If it is determined that the number of transcoding instances has not reached the maximum value, the transcoding cluster is expanded.

[0034] It should be noted that the embodiments of this specification have the following beneficial effects through the above contents: Dynamic resource management: By monitoring the number of task queues in real time, the system can dynamically adjust resources according to load conditions to ensure that resources are used effectively.

[0035] Avoid resource bottlenecks: By setting expansion thresholds, the system can expand capacity before the number of task queues reaches a certain size, avoiding a backlog of transcoding tasks due to insufficient resources.

[0036] Improve transcoding efficiency: Timely expansion when the number of task queues increases can maintain the continuity of transcoding tasks and improve overall transcoding efficiency.

[0037] Furthermore, if it is determined that the number of task queues corresponding to the transcoding task is not greater than the capacity expansion threshold, the method further includes: Determine whether the number of task queues is less than a recycling threshold, where the recycling threshold is set based on the current concurrency capability; If it is determined that the number of task queues is less than the recycling threshold, determine whether the number of cluster transcoding instances has reached the minimum value; If it is determined that the number of cluster transcoding instances has not reached the minimum value, the transcoding cluster is recycled.

[0038] It should be noted that the embodiments of this specification have the following beneficial effects through the above contents: Optimized resource allocation: By setting the recycling threshold, the system can release unnecessary resources when the number of task queues is low, thereby optimizing resource allocation.

[0039] Reduce costs: Automatically reclaim resources during low-load periods to reduce unnecessary expenses and lower operating costs.

[0040] Improve resource utilization: By recycling idle transcoding instances, the system can use existing resources more effectively and improve resource utilization.

[0041] One or more embodiments of this specification provide a video transcoding device for a television station, including: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to: Obtaining a TV station's on-demand video, and setting the on-demand video as a video to be transcoded; Based on a pre-trained video transcoding evaluation model, a transcoding score of each transcoding service for the video to be transcoded is obtained, wherein the video transcoding evaluation model is a neural network model, and the evaluation score is determined based on a time consumption evaluation, a task queuing time evaluation, and a transcoding quality evaluation of each transcoding service; Determining a designated transcoding service for the video to be transcoded based on the transcoding score; The video to be transcoded is transcoded based on the designated transcoding service to obtain a transcoded video suitable for broadcasting by the TV station.

[0042] One or more embodiments of this specification provide a non-volatile computer storage medium storing computer executable instructions, which can achieve the following when executed by a computer: Obtaining a TV station's on-demand video, and setting the on-demand video as a video to be transcoded; Based on a pre-trained video transcoding evaluation model, a transcoding score of each transcoding service for the video to be transcoded is obtained, wherein the video transcoding evaluation model is a neural network model, and the evaluation score is determined based on a time consumption evaluation, a task queuing time evaluation, and a transcoding quality evaluation of each transcoding service; Determining a designated transcoding service for the video to be transcoded based on the transcoding score; The video to be transcoded is transcoded based on the designated transcoding service to obtain a transcoded video suitable for broadcasting by the TV station.

[0043] At least one of the above technical solutions adopted in the embodiments of this specification can achieve the following beneficial effects: Improve transcoding efficiency: Through the pre-trained video transcoding evaluation model, the most suitable transcoding service can be quickly selected for the video to be transcoded, thereby reducing unnecessary transcoding attempts and improving transcoding efficiency.

[0044] Optimize resource allocation: The decision-making mechanism based on transcoding scores can allocate transcoding tasks more reasonably, making resources (such as CPU, memory, etc.) more effectively utilized and avoiding resource waste.

[0045] Improve transcoding quality: By comprehensively considering factors such as time consumption, task queue time and transcoding quality, the final output transcoded video is ensured to be not only efficient but also of high quality.

[0046] Enhanced load balancing: This method can take into account the real-time load differences across service providers, avoiding the problem of uneven load within a single service provider, thereby achieving more effective load balancing. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art description. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative labor. In the drawings: Figure 1 A schematic flow chart of a video transcoding method for a television station provided in one or more embodiments of this specification; Figure 2 A diagram of a multi-service provider transcoding architecture provided for one or more embodiments of this specification; Figure 3 A flowchart of intelligent transcoding scheduling provided for one or more embodiments of this specification; Figure 4 A flowchart of a state feedback scheduling provided for one or more embodiments of this specification; Figure 5 A cluster expansion / recovery scheduling flow chart provided for one or more embodiments of this specification; Figure 6 A schematic diagram of the structure of a video transcoding device for a television station provided in one or more embodiments of this specification. DETAILED DESCRIPTION

[0048] The embodiments of this specification provide a video transcoding method, device and medium for a television station.

[0049] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.

[0050] Figure 1The present invention provides a flowchart of a video transcoding method for a television station in one or more embodiments of the present invention, and the flowchart can be executed by a video transcoding system of the television station. Some input parameters or intermediate results in the flowchart allow manual intervention and adjustment to help improve accuracy.

[0051] The method steps of the embodiment of this specification are as follows: S101, obtaining a TV station's on-demand video, and setting the on-demand video as a video to be transcoded.

[0052] In the embodiments of this specification, on-demand video is the on-demand video content provided by the TV station through multiple terminals (such as TVs and mobile terminals), which needs to adapt to different bit rates and formats according to the characteristics of the terminals. The video to be transcoded is the original video file that needs to be transcoded to generate a streaming media file adapted to different playback scenarios.

[0053] It should be noted that the embodiments of this specification can obtain on-demand video source files from the TV station's media resource library through API or file transfer protocols (such as FTP, HTTP), and support common formats (MPEG-4, H.264, etc.). Use the FFprobe tool to detect video metadata (resolution, frame rate, encoding format) to ensure file integrity and compliance. Build a queue of tasks to be transcoded, dynamically sort by priority (such as urgent live content > on-demand episodes), and support breakpoint resumption and abnormal retry S102, based on a pre-trained video transcoding evaluation model, obtain the transcoding score of each transcoding service for the video to be transcoded, the video transcoding evaluation model is a neural network model, and the evaluation score is determined based on the time consumption evaluation, task queuing time evaluation and transcoding quality evaluation of each transcoding service.

[0054] In an embodiment of the present specification, the video transcoding evaluation model can be based on a comprehensive evaluation system of a neural network, input video features and service provider status, and output three scores of time consumption, queue time, and quality. The time consumption evaluation is to predict the completion time of the transcoding task, which is affected by the service provider's hardware performance (GPU computing power) and the efficiency of the encoding algorithm (such as H.265 compression rate). The task queue time evaluation is to analyze the backlog of the service provider's task queue and the computing power margin in real time, and dynamically adjust the task allocation weight. The transcoding quality evaluation is to quantify the image quality loss based on VMAF (Video Multi-method Assessment Tool) and PSNR (Peak Signal-to-Noise Ratio), and balance the bandwidth cost in combination with the file size constraint (such as the bit rate compliance rate ≥ 95%).

[0055] In the embodiments of this specification, the time-consuming model can input video resolution, dynamic complexity (such as motion vectors) and service provider hardware parameters (GPU model), and use LSTM to predict task time, with an error reduction of 20%-30%. The queue length model can be combined with the service provider's real-time queue status (number of backlog tasks, CPU occupancy rate), dynamically perceive load fluctuations through the pulse neural network (SNN), and shorten the scheduling response time by 50%. The quality model can introduce multimodal features (HDR information, quantization matrix), dynamically adjust the encoding parameter weights through the attention mechanism, and reduce the VMAF score standard deviation from ±5 to ±1.5.

[0056] S103: Determine a designated transcoding service for the video to be transcoded based on the transcoding score.

[0057] In the embodiments of this specification, the transcoding service is a technical service provider that provides video transcoding capabilities, which may include a public cloud, a private cluster, or a hybrid deployment architecture. The designated transcoding service may be the best service provider selected by the evaluation model, taking into account efficiency, cost, and quality.

[0058] In the embodiments of this specification, for high-concurrency tasks (such as live sports events), service providers with a time-consuming score ≥ 0.8 can be given priority; for ultra-high-definition on-demand content, service providers with a quality score ≥ 0.9 can be selected. Real-time monitoring of the health status of service providers (such as API response delay) will automatically switch to the backup node if abnormal, and the disaster recovery response time will be ≤ 3 seconds. Encapsulate the differences in API protocols of different service providers and unify scheduling instructions through standardized JSON fields.

[0059] S104, performing a transcoding task on the video to be transcoded based on the designated transcoding service to obtain a transcoded video suitable for broadcasting by the TV station.

[0060] In the embodiments of this specification, the adapted transcoded video is to output a streaming media file that meets the terminal playback requirements, including multiple bitrates (such as 1080p@8Mbps, 720p@4Mbps), multiple formats (HLS, DASH) and DRM encryption support.

[0061] In the embodiments of this specification, the encoding parameters can be automatically matched according to the characteristics of the video content (such as HDR video using 10-bit color depth, animation content enabling inter-frame prediction optimization), and the bit rate saving rate is ≥50%. The video is divided into independent segments by GOP (group of pictures), and the throughput can be increased by 40%-60% through parallel transcoding of distributed clusters (such as Hadoop framework). FFmpeg is used for post-transcoding verification (such as black frame detection and audio and video synchronization); adaptive streams are injected through CDN edge nodes (such as AWS CloudFront), and dynamic bit rate switching (ABR) is supported.

[0062] It should be noted that this application uses a neural network model to evaluate the time consumption, task queuing time and transcoding quality of each transcoding service in real time, solving the problems of low resource utilization and unbalanced load across service providers caused by the traditional system's reliance on static allocation strategies. For example, when a service provider has a queue backlog due to high concurrent tasks, the model can dynamically perceive its remaining computing power and real-time load, and prioritize tasks to low-load service providers to avoid task delays. Compared to the traditional Hadoop framework that only optimizes the internal resource allocation of a single service provider, this method realizes dynamic scheduling of the global resource pool, and resource utilization can be improved by 30%-50%.

[0063] At the same time, the neural network model can predict the optimal encoding parameter combination (such as bit rate, resolution, GOP structure) for a specific video from different service providers by learning historical transcoding data (such as the correlation between service provider time consumption, quality indicators and video features). For example, for high dynamic range (HDR) videos, the model automatically recommends a high bit rate configuration to preserve details; while for low-complexity animations, the bit rate is reduced to save bandwidth costs. Combined with the test results of the TNG algorithm, this method can reduce storage space usage by 50%-70%, while reducing CDN bandwidth costs.

[0064] Moreover, by encapsulating heterogeneous APIs of multiple service providers (such as protocol differences and callback mechanisms) through a unified interface, this application can simplify the complexity of system integration. For example, referring to the cluster configuration solution of the EasyCVR platform, multiple service provider nodes can be deployed with one click and standardized task distribution can be achieved. The model can also support real-time monitoring and log analysis of transcoding tasks, making it easier to quickly locate service provider performance bottlenecks (such as abnormal time consumption due to hardware aging of a service provider).

[0065] In addition, when a transcoding service provider fails, the model can automatically switch to a backup service provider based on the real-time evaluation score, achieving disaster recovery switching in seconds. For example, the clustered deployment experience of the EasyCVR platform can avoid CDN injection interruptions caused by manual intervention through dynamic load monitoring and task migration mechanisms. At the same time, it supports unified adaptation of heterogeneous service provider interfaces to reduce the cost of collaborative development among multiple service providers.

[0066] Furthermore, before obtaining the transcoding scores of the video to be transcoded by each transcoding service, an initial video transcoding evaluation model can be constructed, wherein the initial video transcoding evaluation model includes a first neural network model corresponding to the time consumption evaluation of each transcoding service, a second neural network model for evaluating the task queuing time of each transcoding service, and a third neural network model for evaluating the transcoding quality of each transcoding service; the initial video transcoding evaluation model is trained based on the training set to obtain a video transcoding evaluation model that meets the requirements.

[0067] It should be noted that the embodiments of this specification solve the decision-making bias problem caused by the traditional single evaluation dimension (such as relying only on bit rate) by constructing a joint training framework of the first neural network model (time consumption evaluation), the second neural network model (task queuing time evaluation) and the third neural network model (transcoding quality evaluation).

[0068] For time consumption evaluation, a model design that predicts transcoding time consumption can be designed based on video features (such as resolution, frame rate, and encoding format). The first neural network can learn the correlation between the hardware performance of different service providers and the efficiency of encoding algorithms. For example, by extracting features such as the bit rate, resolution, and dynamic complexity of the video, combined with the FFprobe tool, the service provider's computing power status (such as CPU occupancy and the number of queue backlog tasks) can be obtained in real time.

[0069] For task queue duration evaluation, the second neural network can dynamically correct the queue duration prediction value by analyzing the real-time queue status of the service provider (such as the number of backlog tasks and computing power margin), which can avoid the imbalance of resource allocation caused by the static polling strategy. For example, for the comprehensive transcoding features after dimensional expansion and fusion (such as picture coding features and subjective quality indicators), the model can dynamically perceive the load fluctuations of different service providers, and the scheduling decision response time is greatly shortened.

[0070] For transcoding quality assessment, the third neural network can use VMAF (Video Multi-method Assessment Tool) as the core indicator, and combine it with the post-transcoding file size threshold constraint (score fs (w(x)) calculation rule) to ensure the balance between image quality and bandwidth cost. For example, when the size of the transcoded sample video file exceeds the original file, the parameter group is directly determined to be unavailable, and the image quality loss is quantified through VMAF scoring. The image quality of the model in complex scenes can be improved stably.

[0071] In addition, this application can realize automatic interface adaptation and intelligent fault switching by uniformly encapsulating the API protocol differences of different service providers.

[0072] For interface adaptation automation, the differences in service provider interfaces are converted into standardized feature vectors for input into the model. For example, the fourth-order cross multiplication operation is used to integrate the picture coding features, bit rate level and subjective quality indicators to avoid manual code adaptation.

[0073] For intelligent fault switching, when a service provider encounters an abnormality, the model quickly switches to a backup service provider based on historical load data.

[0074] It should be noted that the embodiments of this specification have the following beneficial effects through the above contents: Improving evaluation accuracy: By building three neural network models that include time consumption evaluation, task queue duration evaluation, and transcoding quality evaluation, multiple aspects of video transcoding can be accurately evaluated, thereby improving the accuracy of the transcoding score.

[0075] Model generalization ability: By using the training set to train the model, the model can learn the characteristics and rules of various transcoding services, improving the generalization ability of the model and enabling it to be applied to a wider range of video transcoding scenarios.

[0076] Fast decision support: The trained evaluation model can quickly provide transcoding scores for the videos to be transcoded, which speeds up the decision-making process and reduces the waiting time for transcoding services.

[0077] Furthermore, when the pre-trained video transcoding evaluation model is used to obtain the transcoding scores of the video to be transcoded by each transcoding service, the first score of the first neural network model corresponding to each transcoding service, the second score of the second neural network model corresponding to each transcoding service, and the third score of the third neural network model corresponding to each transcoding service can be obtained based on the video transcoding evaluation model; the first score, the second score, and the third score are input into a pre-set evaluation formula to obtain the transcoding scores of each transcoding service for the video to be transcoded.

[0078] It should be noted that sub-model training and data preparation can be implemented through the following specific implementation plans: Time evaluation (first neural network model): Input features: video attributes (resolution, frame rate, encoding format), service provider hardware parameters (GPU model, computing power margin).

[0079] Model selection: The xDeepFM model is used to capture high-order feature relationships (such as the nonlinear relationship between duration and slice size) through a compression interaction module, and the prediction error is reduced by 70% compared with manual rules. LSTM (Long Short-Term Memory) prediction can also be used to predict the time consumption of this transcoding based on the service provider's historical transcoding time consumption data.

[0080] Training data: Based on historical transcoding time data, a million-level sample set is constructed, which contains 54 features (such as file size and streaming transcoding tags).

[0081] Task queuing duration evaluation (second neural network model): Input features: service provider’s real-time queue status (number of backlog tasks, CPU usage), and task priority.

[0082] Model selection: Use spiking neural networks (SNNs) to dynamically sense load fluctuations and reduce scheduling response time by 50%.

[0083] Training data: Combine the service provider API with real-time monitoring data to build a time series data set and dynamically update queue status parameters.

[0084] Transcoding quality assessment (third neural network model): Input features: video content features (HDR information, dynamic complexity), encoding parameters (quantization matrix, bit rate level).

[0085] Model selection: The reference-free VSFA model is used, combined with the GRU network to model the temporal memory effect, and the standard deviation of the VMAF score is reduced to ±1.5. A picture quality assessment model based on the GAN (Generative Adversarial Network) network can also be constructed to evaluate the quality of the service provider's historical transcoded videos.

[0086] Training data: Use large-scale UGC video datasets such as LSVQ, which contains 110,000 diverse video samples.

[0087] When designing a comprehensive evaluation formula, the following specific implementation plans can be adopted: Normalization: The three scores are uniformly mapped to the range of 0-1 to eliminate the dimensional differences of service provider indicators.

[0088] Dynamic weight allocation: Dynamically adjust the weight coefficient according to the business scenario. For example: Live broadcast scenario: time weight (α=0.6), quality weight (γ=0.3); On-demand scenario: quality weight (γ=0.5), queue time weight (β=0.3).

[0089] It should be noted that the embodiments of this specification have the following beneficial effects through the above contents: Automated decision-making process: By using a neural network model to automatically calculate transcoding scores, the need for manual decision-making is reduced, making the allocation of transcoding tasks more automated.

[0090] Improved evaluation speed: Pre-trained models can quickly process large amounts of data, thereby quickly obtaining transcoding scores for each transcoding service, improving the efficiency of the entire transcoding process.

[0091] Enhanced accuracy: The neural network model is used to conduct a comprehensive analysis of multiple evaluation dimensions (time consumption, task queue time, transcoding quality), which can more accurately predict the performance of the transcoding service.

[0092] Model interpretability: Although the neural network model itself may be difficult to explain, the calculation process of the score can be explained through the pre-set evaluation formula, which helps to understand the basis of the scoring.

[0093] Furthermore, the evaluation formula is: Score = α / Cost+ β / Delay+ γ·Quality; where, Score is the transcoding score, Cost is the first score, Delay is the second score, Quality is the third score, α+β+γ=1, and α, β, and γ are all positive numbers.

[0094] It should be noted that the design of the sub-item model and the calculation of the scores can be implemented through the following specific implementation plan: (1) Cost (time consumption evaluation) Input features: Video attributes: resolution, frame rate, encoding format, dynamic complexity (such as motion vector density).

[0095] Service provider hardware parameters: GPU model, computing power margin, and current task queue length.

[0096] Model selection: Use LSTM or xDeepFM models to capture the nonlinear relationship between video features and hardware performance. For example, by compressing the interactive module, the prediction time error is reduced by 20%-30%.

[0097] Constraints: If the transcoding time exceeds the preset threshold (such as ≤5 seconds for live broadcast scenarios), set Cost = ∞ so that the score approaches zero.

[0098] (2) Delay (queue duration evaluation) Input features: Real-time queue status: number of backlog tasks, CPU / GPU usage, and historical load fluctuations of service providers.

[0099] Task priority: High-priority tasks (such as 4K live streaming) require dynamic queue interruption.

[0100] Model selection: Use a spiking neural network (SNN) to dynamically sense service provider load fluctuations and predict queue length response times, reducing response time by 50%.

[0101] Combined with dynamic programming algorithm, the task scheduling strategy is optimized to reduce the imbalance of resource allocation.

[0102] (3) Quality (transcoding quality assessment) Input features: Video content characteristics: HDR information, dynamic complexity, and quantization matrix parameters.

[0103] Encoding parameters: bit rate level, key frame interval, chroma offset.

[0104] Model selection: Use a no-reference quality assessment model (such as VSFA) or a generative adversarial network, combined with a GRU network to model temporal image quality stability, and the standard deviation of the VMAF score is reduced to ±1.5.

[0105] Constraint: If the size of the transcoded file exceeds the original file, the quality score is set to zero.

[0106] 2. Dynamic Weight Adjustment Strategy Scene Adaptive Mechanism: Live broadcast scenario: α=0.6 (time consumption priority), β=0.3 (low latency), γ=0.1 (basic image quality).

[0107] On-demand scenario: γ=0.5 (image quality first), α=0.3, β=0.2.

[0108] Disaster recovery scenario: Dynamically increase the β weight and give priority to service providers with idle queues.

[0109] Automatic parameter adjustment: The importance of each dimension in different scenarios is automatically learned through the attention mechanism. For example, high dynamic range (HDR) videos need to increase the γ weight to retain details.

[0110] Combine A / B testing to verify the actual effect of the weight combination, such as comparing user viewing time and freeze rate, and optimizing α, β, and γ.

[0111] It should be noted that the embodiments of this specification have the following beneficial effects through the above contents: Weighted comprehensive evaluation: This formula introduces weights α, β, γ to weight time consumption, cost, and quality according to the specific needs of the TV station, thereby providing a comprehensive transcoding score.

[0112] Flexible adjustment: Since α, β, and γ can be adjusted independently, TV stations can adjust the weight of each factor according to their own business priorities (such as cost control, real-time requirements, or quality standards) to improve the flexibility of evaluation.

[0113] Cost-benefit analysis: By using the inverse of Cost (1 / Cost) as the weight, this formula encourages the selection of more cost-effective transcoding services, which helps reduce operating costs.

[0114] Real-time guarantee: By assigning a weight β to the inverse of Delay (1 / Delay), this formula emphasizes the importance of real-time in transcoding service selection and ensures smooth video playback.

[0115] Quality priority: The Quality term is directly multiplied by γ, indicating that video quality is a key factor in the evaluation, which is conducive to ensuring the quality of the output video.

[0116] It should be noted that the traditional transcoding evaluation model only focuses on general indicators such as time consumption, queue, quality, etc., and does not consider the differences in the video content's own bit rate requirements (such as news broadcasts and portrait interviews have different sensitivities to image quality). Based on this, the embodiment of this specification can add a fourth neural network model to extract the semantic features of the video key frames (such as face ratio, motion intensity, text area density) through the pre-trained ResNet-50. The specific implementation can be through the following content: 1. Key frame screening mechanism Sliding window dynamic sampling: Based on the sliding window technology, the video is divided into overlapping sequence segments (the window length is recommended to be 5 seconds), and the pre-trained ResNet-50 is used to extract the spatial features of each frame.

[0117] Temporal feature fusion: Calculate the similarity between frames and select the key frames with the highest information entropy (such as face / text area mutation frames).

[0118] Among them, when calculating the similarity between frames, it can be based on the video displacement window multi-head self-attention mechanism (Video SW-MSA).

[0119] 2. Multi-dimensional semantic feature extraction Face ratio: Based on face detection technology, combined with OpenCV's Haar cascade classifier or DNN module, the face area ratio in the key frame is counted.

[0120] Motion intensity: Deep Feature Flow technology is used to calculate the displacement of feature maps of adjacent key frames through the optical flow method to quantify the motion intensity.

[0121] Text area density: Integrates text detection algorithms, uses Canny edge detection + projection analysis to locate text areas, and calculates the percentage of text pixels.

[0122] It should be noted that the fourth neural network module architecture design can integrate multi-dimensional semantic features with the original time-consuming, queuing, and quality assessment models to achieve dynamic weight allocation.

[0123] 3. Module structure Input layer: keyframe image (224x224 RGB) + temporal context features.

[0124] Backbone network: pre-trained ResNet-50 as encoder (freeze the weights of the first 3 layers), outputting a 1024-dimensional spatial feature vector.

[0125] Multi-task branch: Face branch: The RoI Align layer is connected after the conv5_x layer of ResNet-50 to locate the face area and calculate the proportion.

[0126] Motion branch: Based on the flow field propagation module, the optical flow field is calculated through FlowNet Half and the motion intensity score is output.

[0127] Text branch: Connect the projection analysis module after the conv4_x layer of ResNet-50 to generate a text area density map.

[0128] Feature fusion strategy: Dynamic weighted fusion: Based on the multi-attention mechanism (MACNet), a channel-spatial attention module is designed to automatically adjust the weights according to the importance of semantic features.

[0129] Formula definition: ; Among them, the semantic weight is calculated by the face ratio, motion intensity, and text density through the fully connected layer.

[0130] 4. Training and optimization strategies Synthetic dataset: Generate annotated virtual videos based on sample synthesis methods (such as superimposing text with different fonts, motion blur, and adjusting the size of human faces).

[0131] Adversarial training: Prevent overfitting of complex background scenes through Dropout and L2 regularization. Quantized deployment: The residual dual attention module is used to replace ResNet-50 with a lightweight ResNet-18, and TensorRT is applied for FP16 quantization.

[0132] Edge computing optimization: Based on the distributed processing architecture, the feature extraction module is deployed to the edge nodes to reduce the load on the central cluster.

[0133] Furthermore, after performing the transcoding task on the video to be transcoded based on the designated transcoding service, the processing status of the transcoding task can be obtained by calling the status query interface of the designated transcoding service; if the processing status is transcoding completion, feedback of transcoding completion is sent to the user terminal.

[0134] It should be noted that the designated transcoding service selects the best transcoding service provider through the evaluation model, which provides a standardized API interface to perform video transcoding tasks. That is, the service provider is dynamically selected based on the transcoding score (such as time consumption, quality, load, etc.), for example, A is suitable for high-concurrency live transcoding, and B is suitable for ultra-high-definition on-demand transcoding. For example, when the interface is adapted, A calls the a interface and needs to pass the aa parameter. B calls the b interface and needs to pass the bb parameter.

[0135] The status query interface is an API interface provided by the transcoding service provider, which is used to obtain the progress and result status of the transcoding task in real time. The transcoding task processing status is used to describe the complete life cycle status of the transcoding task from submission to completion, including progress percentage, error code, and output file information. Feedback to the user terminal on transcoding completion can be done through a callback mechanism or active notification, and the transcoding result (success / failure) and output file URL can be pushed to the user terminal.

[0136] It should be noted that the embodiments of this specification have the following beneficial effects through the above contents: Real-time status update: By calling the status query interface of the specified transcoding service, the system can obtain the processing status of the transcoding task in real time, thereby providing users with the latest transcoding progress information.

[0137] Improve user experience: When users learn that a transcoding task has been completed, they can receive timely notifications, which improves user satisfaction and experience because users no longer need to wait or actively check the transcoding status.

[0138] Reduce user anxiety: Real-time status feedback reduces user anxiety about possible problems with transcoding tasks, and users can wait without worrying about progress.

[0139] Further, if the processing status is that transcoding is not completed, it can be determined whether the transcoding is abnormal; if so, the transcoding task is re-performed on the video to be transcoded based on the designated transcoding service.

[0140] It should be noted that the embodiments of this specification have the following beneficial effects through the above contents: Improve reliability: By determining whether transcoding is abnormal, the system can ensure the reliability of video transcoding and avoid incomplete transcoding results caused by errors or interruptions.

[0141] Automatic fault recovery: Once a transcoding anomaly is detected, the system automatically attempts to re-transcode, which reduces the need for manual intervention and improves the efficiency of fault recovery.

[0142] Reduce waiting time: Automatically retrying transcoding tasks can reduce the waiting time of users because users do not need to manually resubmit transcoding requests.

[0143] Furthermore, the embodiments of this specification can also determine whether the number of task queues corresponding to the transcoding task is greater than the capacity expansion threshold, and the capacity expansion threshold is set based on the current concurrency capability; if it is determined that the number of task queues corresponding to the transcoding task is greater than the capacity expansion threshold, determine whether the number of cluster transcoding instances has reached the maximum value; if it is determined that the number of transcoding instances has not reached the maximum value, expand the transcoding cluster.

[0144] It should be noted that the embodiments of this specification have the following beneficial effects through the above contents: Dynamic resource management: By monitoring the number of task queues in real time, the system can dynamically adjust resources according to load conditions to ensure that resources are used effectively.

[0145] Avoid resource bottlenecks: By setting expansion thresholds, the system can expand capacity before the number of task queues reaches a certain size, avoiding a backlog of transcoding tasks due to insufficient resources.

[0146] Improve transcoding efficiency: Timely expansion when the number of task queues increases can maintain the continuity of transcoding tasks and improve overall transcoding efficiency.

[0147] Furthermore, if it is determined that the number of task queues corresponding to the transcoding task is not greater than the expansion threshold, determine whether the number of task queues is less than the recovery threshold, and the recovery threshold is set based on the current concurrency capacity; if it is determined that the number of task queues is less than the recovery threshold, determine whether the number of cluster transcoding instances has reached the minimum value; if it is determined that the number of cluster transcoding instances has not reached the minimum value, recycle the transcoding cluster.

[0148] It should be noted that the embodiments of this specification have the following beneficial effects through the above contents: Optimized resource allocation: By setting the recycling threshold, the system can release unnecessary resources when the number of task queues is low, thereby optimizing resource allocation.

[0149] Reduce costs: Automatically reclaim resources during low-load periods to reduce unnecessary expenses and lower operating costs.

[0150] Improve resource utilization: By recycling idle transcoding instances, the system can use existing resources more effectively and improve resource utilization.

[0151] It should be noted that the on-demand video in the broadcasting and television industry needs to be transcoded first and then injected into the CDN. In order to avoid technology monopoly, broadcasting and television companies sometimes introduce two or more transcoding service providers. The transcoding interfaces of different service providers are different and need to be connected in a targeted manner. In addition, transcoding load scheduling between multiple service providers requires a scheduling program. Generally, the load strategy is relatively simple and cannot fully utilize transcoding resources. This solution provides a multi-service provider transcoding resource collaboration method, which abstracts the transcoding resources of multiple service providers into a unified transcoding resource pool, implements transcoding resource scheduling through artificial intelligence algorithms, and fully utilizes the transcoding resources of multiple service providers.

[0152] Furthermore, the existing quality assessment model only supports SDR videos and cannot handle the special transcoding requirements of VR videos. Therefore, the embodiment of this specification predicts the viewpoint of the spherical projection area of ​​the VR video (using the user's head display posture data), adjusts the spherical projection center, and reduces the geometric distortion of the edge area. The specific implementation scheme can be as follows: Real-time collection of head display posture data: The user's head posture (yaw angle, pitch angle, roll angle) is obtained in real time through the built-in IMU sensor (gyroscope, accelerometer) of the VR device, and the LSTM (long short-term memory network) model is built in combination with the user's historical trajectory data to predict the center coordinates (θ, φ) and field of view coverage within the next 3 seconds. The dynamic projection center offset technology is introduced to adjust the spherical projection center according to the predicted viewpoint position, reduce the geometric distortion of the edge area, and improve the coding efficiency.

[0153] Spherical area heat modeling: Based on the user's viewpoint trajectory data and video content semantics (such as motion intensity, face / text distribution), the Gaussian probability distribution model is used to divide the core area (±30° from the center of the field of view), the transition area (±60° from the core area), and the edge area (the remaining area). Using spherical multi-projection technology, the spherical video is converted into a dynamic projection coordinate system centered on the viewpoint to reduce pixel redundancy in the edge area.

[0154] Core area: High-resolution encoding (such as H.266 / AVS3) can be used, with 1.5 times the base bit rate allocated, giving priority to retaining details (such as faces and text).

[0155] Transition zone: Adaptive QP value adjustment can be used (see page 10) to maintain the base bit rate and combine motion vector expansion technology to suppress ghosting.

[0156] Edge area: The intelligent tile stitching algorithm (see page 2) can be used to merge low-heat areas into low-resolution sub-blocks and allocate 0.8 times the base bit rate.

[0157] The present invention belongs to the technical field of multimedia data processing, and specifically relates to a multi-service provider transcoding resource collaboration method based on deep reinforcement learning.

[0158] This method includes two parts: multi-service provider transcoding resource pool and intelligent scheduling engine. Figure 2 , the specific contents are as follows: Multi-provider transcoding resource pool: pools transcoding resources from multiple providers and provides unified abstract management for the transcoding resources of the providers, including unified interface adaptation and containerized adaptation.

[0159] Unified interface adaptation: Due to the differences in interfaces between different service providers, it is necessary to implement interface abstraction adaptation for each connected service provider. Two major adaptations are completed: creating a transcoding task interface adaptation and transcoding task status query interface adaptation. The adapter interface code is as follows: public interface IJobService { / ** * Create a transcoding task * @param sourcePath original file address * @return jobId task ID * / String createJob(String sourcePath); / ** * Query the status of the transcoding task * @param jobId task ID * @return JobStatus Task status: in queue, transcoding, transcoding exception, transcoding completed * / String queryJobStatus(String jobId); } Containerization adaptation: If the service provider's transcoding service supports containerized deployment, the service provider's transcoding service can be further packaged in containers to support the expansion and recovery of the transcoding cluster in the resource pool.

[0160] Intelligent scheduling engine: includes intelligent transcoding scheduling, status feedback scheduling, and capacity expansion / recycling scheduling.

[0161] Intelligent transcoding scheduling: Initiate transcoding tasks and use artificial intelligence algorithms to decide on the service provider that performs the transcoding.

[0162] For the flowchart of intelligent transcoding scheduling, see Figure 3 , the specific contents are as follows: 1. Transcoding scheduling starts and tasks to be transcoded are obtained; 2. Intelligent scheduling algorithm: Build a three-dimensional evaluation model of cost-delay-quality to evaluate the transcoding of service providers. Prioritize service providers with high scores to perform transcoding tasks. The evaluation formula is Score = α / Cost+ β / Delay+ γ·Quality (α+β+γ=1, dynamically adjust the weight). Cost: The service provider's transcoding time estimate can be predicted using LSTM (Long Short-Term Memory) prediction to predict the current transcoding time based on the service provider's historical transcoding time data. Delay: Task queue duration, predict the transcoding task queue duration based on the service provider's current task queue, CPU, and memory usage. Quality: Transcoding quality, a picture quality evaluation model based on the GAN (Generative Adversarial Network) network can be built to evaluate the quality of the service provider's historical transcoded videos.

[0163] 3. Create a transcoding task: call the "Create Transcoding Interface" adapter to trigger the service provider's transcoding task, and transcoding will start.

[0164] 4. Transcoding scheduling ends: complete this scheduling and wait for the next scheduling.

[0165] Status feedback scheduling: Check whether the transcoding task has been completed, feedback the transcoding results to the user, and support transcoding retries from multiple service providers. Figure 4 , the specific contents are as follows: 1. Task status feedback scheduling begins.

[0166] 2. Query the transcoding task status: call the service provider's transcoding status query interface to obtain the current task status.

[0167] 3. Determine whether the transcoding is completed: If it is completed, the terminal user is fed back that the transcoding is completed, and then this scheduling ends. Otherwise, determine whether the transcoding is abnormal.

[0168] 4. Determine whether the transcoding is abnormal: if it is abnormal, the transcoding will be retried. Otherwise, the scheduling ends.

[0169] 5. Trigger transcoding retry to see if the maximum number of retries has been reached. If so, feedback indicates transcoding failure and the scheduling ends. Otherwise, initiate a transcoding retry.

[0170] 6. Transcoding retry: Try to change the service provider and re-initiate transcoding. This scheduling ends.

[0171] Expansion / recycling scheduling: When dealing with a large number of sudden transcoding tasks, this device supports dynamic cluster expansion. When the transcoding concurrency decreases, this device supports cluster recycling. For the cluster expansion / recycling scheduling flow chart, see Figure 5 , the specific contents are as follows: 1. Capacity expansion and recovery scheduling begins.

[0172] 2. Determine whether to expand capacity: Whether the number of task queues is greater than the expansion threshold. The expansion threshold = N * current concurrent capacity. N is a variable and can be set to 1.5 based on experience. If it is greater than the threshold, cluster expansion is triggered; otherwise, it is determined whether to recycle.

[0173] 3. Trigger cluster expansion to determine whether the number of transcoding instances has reached the maximum value. If it has reached the maximum value, cluster expansion cannot be performed and this scheduling ends. If it has not reached the maximum value, cluster expansion is performed and this scheduling ends.

[0174] 4. Determine whether to recycle: Whether the number of task queues is less than the recycling threshold, the recycling threshold = M * current concurrency capacity, M is a variable, based on experience can be set to 0.5. If it is less than the threshold, cluster recycling is triggered; otherwise, the current scheduling ends.

[0175] 5. Trigger cluster recycling to determine whether the number of transcoding instances has reached the minimum value. If it has reached the minimum value, cluster recycling cannot be performed and this scheduling ends. If it has not reached the minimum value, cluster recycling is performed and this scheduling ends.

[0176] Transcoding initiation & transcoding feedback: 1. The device integrates transcoding services of multiple service providers and completes the adaptation and encapsulation of the service provider's transcoding interface; 2. Transcoding scheduling is executed regularly to obtain transcoding tasks, determine the transcoding service provider through artificial intelligence algorithms, and send the transcoding tasks to the transcoding service provider to start transcoding; 3. Transcoding feedback scheduling is executed regularly to query the current transcoding task status and feedback the transcoding success or failure status to the end user; Cluster expansion and recovery: 1. Assume that the current cluster has 2 instances, and each instance has a concurrent capacity of 30 transcodes. The expansion threshold coefficient N = 1.5, the recovery threshold coefficient M = 0.5, calculate the expansion threshold = N * current concurrent capacity = 90, the recovery threshold = M * current concurrent capacity = 30, set the maximum number of cluster instances Max = 4, and the minimum number of instances Min = 2.

[0177] 2. The user initiates a transcoding task through the unified gateway interface to simulate large concurrent transcoding. When the number of task queues is greater than 90, cluster expansion is triggered, the number of cluster instances + 1 = 3 instances, the expansion threshold is adjusted to 135, and the recycling threshold is adjusted to 45.

[0178] 3. Continue to simulate large concurrent transcoding. When the number of task queues is greater than 135, the cluster expansion is triggered again. The number of cluster instances + 1 = 4 instances, the expansion threshold is adjusted to 180, and the recycling threshold is adjusted to 60.

[0179] 4. Continue to simulate large concurrent transcoding. When the number of task queues is greater than 180, the number of cluster instances has reached the maximum value of 4. Cluster expansion is not triggered this time.

[0180] 5. Continue to simulate large concurrent transcoding and appropriately reduce the number of task queues. When the number of queues is less than 60, cluster recycling is triggered. The number of cluster instances - 1 = 3, the expansion threshold is adjusted to 135, and the recycling threshold is adjusted to 45.

[0181] 6. Continue to simulate large concurrent transcoding and appropriately reduce the number of task queues. When the number of queues is lower than 45, the cluster is triggered to recycle the number of cluster instances - 1 = 2, the expansion threshold is adjusted to 90, and the recycling threshold is adjusted to 30.

[0182] 7. Continue to simulate large concurrent transcoding and appropriately reduce the number of task queues. When the number of queues is less than 30, cluster recycling is not triggered this time because the number of clusters is the minimum value of 2.

[0183] 1. Unified abstract management mechanism for transcoding resources of multiple service providers.

[0184] 2. Dynamic scheduling method for transcoding tasks based on cost-delay-quality three-dimensional evaluation model 3. Flexible transcoding service expansion and recovery mechanism to cope with sudden and large-scale transcoding tasks.

[0185] Figure 6 A schematic diagram of the structure of a video transcoding device of a television station provided for one or more embodiments of this specification includes: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to: Obtaining a TV station's on-demand video, and setting the on-demand video as a video to be transcoded; Based on a pre-trained video transcoding evaluation model, a transcoding score of each transcoding service for the video to be transcoded is obtained, wherein the video transcoding evaluation model is a neural network model, and the evaluation score is determined based on a time consumption evaluation, a task queuing time evaluation, and a transcoding quality evaluation of each transcoding service; Determining a designated transcoding service for the video to be transcoded based on the transcoding score; The video to be transcoded is transcoded based on the designated transcoding service to obtain a transcoded video suitable for broadcasting by the TV station.

[0186] One or more embodiments of this specification provide a non-volatile computer storage medium storing computer executable instructions, which can achieve the following when executed by a computer: Obtaining a TV station's on-demand video, and setting the on-demand video as a video to be transcoded; Based on a pre-trained video transcoding evaluation model, a transcoding score of each transcoding service for the video to be transcoded is obtained, wherein the video transcoding evaluation model is a neural network model, and the evaluation score is determined based on a time consumption evaluation, a task queuing time evaluation, and a transcoding quality evaluation of each transcoding service; Determining a designated transcoding service for the video to be transcoded based on the transcoding score; The video to be transcoded is transcoded based on the designated transcoding service to obtain a transcoded video suitable for broadcasting by the TV station.

[0187] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device, equipment, and non-volatile computer storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0188] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0189] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0190] In the embodiments provided in the present application, it should be understood that the disclosed devices / network equipment and methods can be implemented in other ways. For example, the device / network equipment embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0191] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0192] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above units may be implemented in the form of hardware or software.

[0193] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0194] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A video transcoding method for a television station, characterized in that: The method comprises: Obtaining a TV station's on-demand video, and setting the on-demand video as a video to be transcoded; Based on a pre-trained video transcoding evaluation model, a transcoding score of each transcoding service for the video to be transcoded is obtained, wherein the video transcoding evaluation model is a neural network model, and the evaluation score is determined based on a time consumption evaluation, a task queuing time evaluation, and a transcoding quality evaluation of each transcoding service; Determining a designated transcoding service for the video to be transcoded based on the transcoding score; The video to be transcoded is transcoded based on the designated transcoding service to obtain a transcoded video suitable for broadcasting by the TV station.

2. The method according to claim 1, characterized in that Before obtaining the transcoding score of each transcoding service for the video to be transcoded, the method further includes: Constructing an initial video transcoding evaluation model, the initial video transcoding evaluation model includes a first neural network model corresponding to the time consumption evaluation of each of the transcoding services, a second neural network model for evaluating the task queuing time of each of the transcoding services, and a third neural network model for evaluating the transcoding quality of each of the transcoding services; The initial video transcoding evaluation model is trained based on the training set to obtain a video transcoding evaluation model that meets the requirements.

3. The method according to claim 2, characterized in that The pre-trained video transcoding evaluation model is used to obtain the transcoding score of the video to be transcoded by each transcoding service, including: Based on the video transcoding evaluation model, obtaining a first score of the first neural network model corresponding to each of the transcoding services, a second score of the second neural network model corresponding to each of the transcoding services, and a third score of the third neural network model corresponding to each of the transcoding services; The first score, the second score, and the third score are input into a preset evaluation formula to obtain a transcoding score of each transcoding service for the video to be transcoded.

4. The method according to claim 3, characterized in that The evaluation formula is: Score = α / Cost+ β / Delay+ γ·Quality; where, Score is the transcoding score, Cost is the first score, Delay is the second score, Quality is the third score, α+β+γ=1, and α, β, and γ are all positive numbers.

5. The method according to claim 1, characterized in that After performing the transcoding task on the video to be transcoded based on the designated transcoding service, the method further includes: Obtaining the processing status of the transcoding task by calling the status query interface of the specified transcoding service; If the processing status is that transcoding is completed, feedback that transcoding is completed is sent to the user terminal.

6. The method according to claim 5, characterized in that If the processing status is that transcoding is not completed, the method further includes: Determine whether transcoding is abnormal; If so, re-transcoding the video to be transcoded based on the designated transcoding service.

7. The method according to claim 1, characterized in that The method further comprises: Determine whether the number of task queues corresponding to the transcoding task is greater than a capacity expansion threshold, where the capacity expansion threshold is set based on a current concurrency capability; If it is determined that the number of task queues corresponding to the transcoding task is greater than the expansion threshold, determine whether the number of cluster transcoding instances has reached a maximum value; If it is determined that the number of transcoding instances has not reached the maximum value, the transcoding cluster is expanded.

8. The method according to claim 7, characterized in that If it is determined that the number of task queues corresponding to the transcoding task is not greater than the capacity expansion threshold, the method further includes: Determine whether the number of task queues is less than a recycling threshold, where the recycling threshold is set based on the current concurrency capability; If it is determined that the number of task queues is less than the recycling threshold, determine whether the number of cluster transcoding instances has reached the minimum value; If it is determined that the number of cluster transcoding instances has not reached the minimum value, the transcoding cluster is recycled.

9. A video transcoding device for a television station, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to: Obtaining a TV station's on-demand video, and setting the on-demand video as a video to be transcoded; Based on a pre-trained video transcoding evaluation model, a transcoding score of each transcoding service for the video to be transcoded is obtained, wherein the video transcoding evaluation model is a neural network model, and the evaluation score is determined based on a time consumption evaluation, a task queuing time evaluation, and a transcoding quality evaluation of each transcoding service; Determining a designated transcoding service for the video to be transcoded based on the transcoding score; The video to be transcoded is transcoded based on the designated transcoding service to obtain a transcoded video suitable for broadcasting by the TV station.

10. A non-volatile computer storage medium, characterized in that: Computer executable instructions are stored, and when the computer executable instructions are executed by a computer, the following can be achieved: Obtaining a TV station's on-demand video, and setting the on-demand video as a video to be transcoded; Based on a pre-trained video transcoding evaluation model, a transcoding score of each transcoding service for the video to be transcoded is obtained, wherein the video transcoding evaluation model is a neural network model, and the evaluation score is determined based on a time consumption evaluation, a task queuing time evaluation, and a transcoding quality evaluation of each transcoding service; Determining a designated transcoding service for the video to be transcoded based on the transcoding score; The video to be transcoded is transcoded based on the designated transcoding service to obtain a transcoded video suitable for broadcasting by the TV station.

Citation Information

Patent Citations

  • Video conversion resource distribution method and system

    CN105992020A

  • Video transcoding method, device and system

    CN109788315A

  • Video transcoding scheduling method and device, computer equipment and storage medium

    CN114245139A

  • Cloud operation resource dynamic allocation system and method thereof

    TWI519967B

  • Method for allocating and scheduling task for maximizing video quality of transcoding server using heterogeneous processors

    US20210329279A1