Video transcoding method and device, electronic equipment, storage medium, program product and method for generating bit stream
By extracting video features and using a prediction model to dynamically adjust the frame rate, the frame rate adaptation problem of live content is solved, improving the user viewing experience and bandwidth utilization.
Patent Information
- Application Number
- CN202511079794.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-01
- Publication Date
- 2025-09-12
AI Technical Summary
The frame rate design of existing live broadcast content is fixed and cannot adapt to the differences in broadcast areas, networks and devices, resulting in a reduced viewer experience or bandwidth waste, affecting the user viewing experience.
By extracting video features and using a prediction model to predict image quality and frame rate mapping information, the optimal frame rate of video clips can be dynamically adjusted, and frame insertion or frame drop processing can be performed to optimize video transcoding.
It achieves flexible frame rate adjustment according to changes in live broadcast scenes, improves the audience's viewing experience, reduces bandwidth waste, and ensures smoothness and image quality.
Smart Images

Figure CN120640039A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of video processing. More specifically, the present disclosure relates to a video transcoding method, apparatus, electronic device, storage medium, program product, and method for generating a bitstream. Background Art
[0002] Live streaming of shows, e-commerce platforms, and games has become one of the hottest and most influential applications in multimedia. Most content providers offer live streaming content at one or more frame rates for tens of thousands of users to maximize clarity and smoothness.
[0003] Currently, most live content frame rates are designed with fixed rules, which set a fixed frame rate at the beginning of the live broadcast. However, due to differences in the region, network, and equipment of the anchor, the anchor's frame rate may not reach the set target frame rate in some scenarios (such as low-end models, weak network, etc.), which directly leads to a decrease in the audience's smoothness experience. At the same time, under the fixed frame rate setting rules, when the anchor is in a scene with more drastic changes, such as outdoor sports, the frame rate may be insufficient, affecting the smoothness. When the anchor is in a scene with more stable changes, such as indoors, the fixed frame rate rule will cause the frame rate to be too high, resulting in additional bit rate overhead, resulting in a certain amount of bandwidth waste, and even causing user lag under conditions of limited network bandwidth, affecting the user's smoothness experience. Summary of the Invention
[0004] Embodiments of the present disclosure provide a video transcoding method, apparatus, electronic device, storage medium, program product, and method for generating a bitstream, for solving at least one of the above-mentioned problems.
[0005] According to one aspect of the present disclosure, a video transcoding method is provided, which includes: extracting video features of currently received video clip data; processing the video features using a pre-trained prediction model to obtain image quality frame rate mapping information of the video clip data, wherein the image quality frame rate mapping information is used to represent a mapping relationship between a subjective image quality index and a frame rate; determining a preferred frame rate of the video clip data based on the image quality frame rate mapping information and a target subjective image quality index; and transcoding the video clip data according to the preferred frame rate to obtain video stream data corresponding to the video clip data.
[0006] Optionally, the prediction model is trained through the following steps: constructing a sample data set, wherein the sample data set includes sample video clip data of multiple different content types and multiple different frame rates; extracting the video features of each sample video clip data, and determining the actual subjective image quality index of each sample video clip data; using the video features as input parameters of the prediction model to be trained, and using the mapping relationship between the actual subjective image quality index and the frame rate as the training target of the prediction model to be trained, training the prediction model to be trained, and obtaining the prediction model.
[0007] Optionally, the video clip data is transcoded according to the preferred frame rate to obtain video stream data corresponding to the video clip data, including: based on the comparison result of the preferred frame rate and the source stream frame rate of the video clip data, the video clip data is transcoded including frame insertion or frame drop processing to obtain video stream data corresponding to the video clip data.
[0008] Optionally, the frame dropping process includes dropping other video frames except key frames.
[0009] Optionally, the video features include at least one of the following: content features, coding features, wherein the content features include at least one of the following: content type features, complexity features, aesthetic features, the content type features include at least one of the following: motion scenes, static scenes, the complexity features include at least one of the following: motion vectors, color change features, the aesthetic features include at least one of the following: composition features, color matching features, and the coding features include at least one of the following: bit rate, noise.
[0010] Optionally, the prediction model includes at least one of the following: XGBoost model, random forest, support vector machine, neural network, reinforcement learning network.
[0011] According to another aspect of the present disclosure, a video transcoding device is provided, which includes: an extraction unit configured to extract video features of currently received video clip data; a processing unit configured to process the video features using a pre-trained prediction model to obtain image quality frame rate mapping information of the video clip data, wherein the image quality frame rate mapping information is used to represent a mapping relationship between a subjective image quality index and a frame rate; a determination unit configured to determine a preferred frame rate of the video clip data based on the image quality frame rate mapping information and a target subjective image quality index; and a transcoding unit configured to perform transcoding processing on the video clip data according to the preferred frame rate to obtain video stream data corresponding to the video clip data.
[0012] Optionally, the prediction model is trained through the following steps: constructing a sample data set, wherein the sample data set includes sample video clip data of multiple different content types and multiple different frame rates; extracting the video features of each sample video clip data, and determining the actual subjective image quality index of each sample video clip data; using the video features as input parameters of the prediction model to be trained, and using the mapping relationship between the actual subjective image quality index and the frame rate as the training target of the prediction model to be trained, training the prediction model to be trained, and obtaining the prediction model.
[0013] Optionally, the transcoding unit is further configured to perform transcoding processing including frame insertion or frame drop processing on the video clip data based on a comparison result of the preferred frame rate and the source stream frame rate of the video clip data to obtain video stream data corresponding to the video clip data.
[0014] Optionally, the frame dropping process includes dropping other video frames except key frames.
[0015] Optionally, the video features include at least one of the following: content features, coding features, wherein the content features include at least one of the following: content type features, complexity features, aesthetic features, the content type features include at least one of the following: motion scenes, static scenes, the complexity features include at least one of the following: motion vectors, color change features, the aesthetic features include at least one of the following: composition features, color matching features, and the coding features include at least one of the following: bit rate, noise.
[0016] Optionally, the prediction model includes at least one of the following: XGBoost model, random forest, support vector machine, neural network, reinforcement learning network.
[0017] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and at least one memory storing computer-executable instructions, wherein when the computer-executable instructions are executed by the at least one processor, the at least one processor is prompted to execute the video transcoding method described above.
[0018] According to another aspect of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by at least one processor, the at least one processor is prompted to perform the video transcoding method as described above.
[0019] According to another aspect of the present disclosure, a computer program product is provided, comprising computer instructions, which implement the video transcoding method described above when the computer instructions are executed by at least one processor.
[0020] According to another aspect of the present disclosure, a method for generating a bitstream is provided, comprising: generating the bitstream according to the video transcoding method described above.
[0021] According to the video transcoding method, device, electronic device, storage medium, program product, and method for generating a bitstream according to the exemplary embodiments of the present disclosure, by analyzing the video features of each video clip data multiple times during the video transcoding process, and using a prediction model to obtain corresponding image quality and frame rate mapping information, and then determining the preferred frame rate, and performing transcoding processing based on this, it is possible to dynamically optimize the image quality and frame rate, thereby improving the user's viewing experience.
[0022] It is to be understood that both the foregoing general description and the following detailed description are examples only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate examples according to the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0024] Figure 1 is a flowchart illustrating an exemplary video transcoding method according to some embodiments of the present disclosure.
[0025] Figure 2 2 is a schematic diagram showing the transcoding process of a live broadcast system according to a specific embodiment of the present disclosure.
[0026] Figure 3 3 is a flow chart showing a video transcoding method according to a specific embodiment of the present disclosure.
[0027] Figure 4 is a block diagram illustrating an exemplary video transcoding apparatus according to some embodiments of the present disclosure.
[0028] Figure 5 is a diagram illustrating a computing environment coupled with a user interface according to some embodiments of the present disclosure. DETAILED DESCRIPTION
[0029] Reference will now be made in detail to the specific embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth to facilitate understanding of the subject matter presented herein. However, various alternatives may be used without departing from the scope of the claims, and the subject matter may be practiced without these specific details. For example, the subject matter presented herein may be implemented on many types of electronic devices with digital video capabilities.
[0030] It should be noted that the terms "first," "second," and the like in the description, claims, and drawings of the present disclosure are used to distinguish between objects, and are not used to describe any specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, such that the embodiments of the present disclosure described herein can be implemented in an order other than that shown in the drawings or described in the present disclosure.
[0031] Live streaming of shows, e-commerce, and games has become one of the hottest and most influential applications in multimedia. Most content providers offer live content at one or more different frame rates for tens of thousands of users to maximize clarity and smoothness. Currently, most live content frame rates are designed based on fixed rules, with a fixed frame rate set at the start of the broadcast. However, due to differences in regions, networks, and devices used by broadcasters, the broadcaster's frame rate may not meet the set target in some scenarios (such as low-end devices and weak network connections), directly reducing the smoothness of the viewer experience. Furthermore, with fixed frame rates, when broadcasters are engaged in highly dynamic scenes, such as outdoor sports, the frame rate may be insufficient, affecting smoothness. However, when broadcasters are in more stable scenes, such as indoors, the fixed frame rate may be too high, resulting in additional bitrate overhead and bandwidth waste. Even with limited network bandwidth, users may experience lag and experience poor smoothness.
[0032] While existing methods for network video transmission can adjust the frame rate based on network conditions, these approaches often neglect the user's subjective experience. This can lead to a significant degradation of the viewing experience despite fully utilizing network bandwidth in some cases. Furthermore, existing methods often rely on fixed frame rate adjustment strategies, lacking flexibility and failing to adapt well to the needs of diverse scenarios and content.
[0033] To address the above issues, exemplary embodiments of the present disclosure propose a video transcoding method that enables adaptive frame rate decision-making and transmission based on live content and user subjective perception of flow. This method dynamically optimizes image quality and frame rate, enhancing the user viewing experience and effectively addressing the subjective experience optimization and bandwidth resource waste issues associated with network video transmission. First, based on subjective image quality metrics, the present disclosure proposes a subjective frame rate scoring method. Through user subjective evaluation experiments, this method determines the differences in user subjective experience of different frame rates in different scenarios. Second, the present disclosure proposes an optimal frame rate prediction method based on live video features. This method extracts features from the live video and, using a machine learning model (i.e., a prediction model), learns user subjective ratings to predict the optimal frame rate. Furthermore, the present disclosure proposes a frame rate adjustment method that enables flexible frame reduction and insertion operations and supports multiple frame reduction and insertion algorithm modes. The present disclosure enables targeted optimization and dynamic, real-time adjustment of the frame rate of live content based on changes in the live scene during the live broadcast process. On the one hand, the present disclosure solves the problem of decreased frame rate of live broadcast content due to heating of the host-side model and network deterioration. On the other hand, it solves the problem of frame rate adaptation in different live broadcast scenarios. It can improve the frame rate in complex motion scenes and improve the smoothness of live broadcast content, thereby improving the audience's viewing experience. It can reduce the frame rate in static scenes, reduce unnecessary bandwidth waste, and ensure that the user experience is not affected.
[0034] Figure 1 This is a flowchart illustrating an exemplary video transcoding method according to some embodiments of the present disclosure. This video transcoding method can be implemented in an electronic device with sufficient computing power. It should be noted that the video transcoding method is used to transcode encoded video data into multiple versions with different bit rates and resolutions for selective playback by viewers, thereby meeting the viewing needs of different viewers.
[0035] Reference Figure 1 In step S101, video features of the currently received video clip data are extracted.
[0036] This step specifically receives the video segment data of the encoded video data. For example, in a live broadcast scenario, the video stream continuously shot and encoded by the host can be received. The currently received video segment data is the video stream segment currently received from the host, so that it can be executed multiple times during the push process of the host. Figure 1 The video transcoding method shown is used to dynamically adjust the transcoding frame rate.
[0037] For example, video features include at least one of the following: content features and encoding features. These features are extracted from two dimensions: the video's content and the encoding process performed by the video capture device. Specifically, content features include at least one of the following: content type features, complexity features, and aesthetic features. Content type features include at least one of the following: moving scenes or static scenes. Complexity features include at least one of the following: motion vectors (e.g., global motion vectors of objects / content in the video), color change features (e.g., inter-frame difference metrics such as inter-frame color distance, dynamic range metrics such as color gamut coverage, and temporal features such as color trend coefficients). Aesthetic features include at least one of the following: composition features (e.g., classic compositional features such as the rule of thirds, element coordinates, subject area proportion, and spatial dispersion), and color matching features (e.g., color relationship types such as contrasting color matching, color emotional labels such as highly saturated colors, and color harmony metrics). Content type features provide a simple summary of the content characteristics of video clip data. Complexity features further describe the changes in the video image. Aesthetic features describe the video content from the user's sensory perspective. These multi-dimensional content features help enrich the content of the video. Coding features include at least one of the following: bitrate and noise. These features reflect the encoding process performed by the video capture terminal and provide additional dimensions for describing the video. This multi-dimensional feature analysis method provides a more comprehensive description of the video's characteristics, providing richer data support for subsequent frame rate optimization decisions.
[0038] In step S102, a pre-trained prediction model is used to process video features to obtain image quality and frame rate mapping information of the video segment data.
[0039] Image quality and frame rate mapping information represents the mapping relationship between subjective image quality metrics and frame rate. For example, it can be a quality-to-frame rate curve, with the horizontal axis representing the frame rate and the vertical axis representing the subjective image quality metrics. This curve allows users to query the frame rate corresponding to the desired subjective image quality metrics. For different videos, even when using the same frame rate, the subjective image quality perceived by users often varies. By analyzing video features using a pre-trained prediction model, we can obtain image quality and frame rate mapping information appropriate for the current video clip data.
[0040] In step S103 , the preferred frame rate of the video segment data is determined according to the image quality frame rate mapping information and the target subjective image quality index.
[0041] The target subjective image quality index is the subjective image quality that is expected to be achieved. By querying the image quality frame rate mapping information based on this, the corresponding frame rate is determined as the preferred frame rate, and then transcoding processing is performed accordingly in the subsequent step S104. This can improve the subjective image quality of the video played after decoding the transcoded video stream data, and can ensure the user's viewing experience while reasonably utilizing bandwidth resources.
[0042] For example, since the impact of frame rate on video quality is primarily reflected in smoothness, the subjective image quality index used here can be an indicator for evaluating the degree of smoothness. In this case, if multiple definition levels are provided, the same target subjective image quality index can be assigned to each definition level to achieve the same preferred frame rate.
[0043] As an example, the target subjective image quality index can be the lowest value of the subjective image quality index. Step S103 can directly use the frame rate corresponding to the target subjective image quality index in the image quality frame rate mapping information as the preferred frame rate. At this time, the preferred frame rate is the lowest frame rate when the minimum requirement of subjective image quality is met, which can reduce the bit rate overhead. If the target subjective image quality index is a value range, indicating the subjective image quality range expected to be achieved, the frame rate with the best subjective image quality can be directly used as the preferred frame rate. Alternatively, based on the image quality frame rate mapping information, the frame rate range corresponding to the target subjective image quality index can be first determined, and then a frame rate can be determined from the frame rate range as the preferred frame rate. For example, other encoding parameters and corresponding optimization targets can be combined to determine a frame rate from the above frame rate range as the preferred frame rate. At this time, a machine learning algorithm, a deep learning algorithm, or a reinforcement learning algorithm can also be used to determine the excellent frame rate. These are all implementation methods of the present disclosure and fall within the scope of protection of the present disclosure.
[0044] In step S104, the video segment data is transcoded according to the preferred frame rate to obtain video stream data corresponding to the video segment data.
[0045] It should be understood that regardless of whether the preferred frame rates for multiple definition levels are the same, multiple definition versions of video stream data will be generated during the transcoding process. It should also be understood that the transcoded video stream data will be sent to the user terminal for decoding and playback, but not all definition versions will be sent. Instead, the video stream data corresponding to the definition level selected by the user terminal will be sent to the user terminal.
[0046] Next, a video transcoding method according to an exemplary embodiment of the present disclosure is further introduced.
[0047] Optionally, step S101 is executed in response to a preset condition being met, which means that in the process of receiving video data, each time the preset condition is met, step S101 needs to be executed. Figure 1 The process shown here updates the preferred frame rate and performs subsequent transcoding according to the updated preferred frame rate. This not only enables more reliable dynamic updates of the preferred frame rate, but also allows the frequency of preferred frame rate updates to be controlled by properly setting preset conditions, allowing sufficient processing time for related calculations and meeting the real-time requirements of video transcoding.
[0048] As an example, each time step S101 is executed, the extracted data may be a recent video segment of a certain length, or all video segment data received from the last execution of step S101 to the current execution of step S101, or a transcoding template segment extracted from all the video segment data.
[0049] As an example, the preset condition includes reaching a preset interval length, that is, the duration between the last execution of step S101 and the current execution of step S101 reaches the preset interval length. The preset interval length can remain unchanged during the process of receiving video data (it can be the same or different for different video data), for example, but not limited to, 5s or 10s, to achieve regular and stable updates of the preferred frame rate; it can also be negatively correlated to the degree of change in the preferred frame rate (for example, it can be the absolute value of the percentage change). In this case, if the change in the preferred frame rate is more significant, it indicates a higher demand for updating the preferred frame rate, and the update interval length can be shortened. Conversely, if the change in the preferred frame rate is smaller, it indicates a lower demand for updating the preferred frame rate, and the update interval length can be extended to save computing power.
[0050] Optionally, the prediction model is trained by the following steps: constructing a sample data set, wherein the sample data set includes sample video clip data of multiple different content types and multiple different frame rates; extracting video features of each sample video clip data, and determining the actual subjective image quality index of each sample video clip data; using the video features as input parameters of the prediction model to be trained, using the mapping relationship between the actual subjective image quality index and the frame rate as the training target of the prediction model to be trained, and training the prediction model to be trained to obtain the prediction model. By collecting sample video clip data of different frame rates, determining their actual subjective image quality index as the training target of the prediction model to be trained, and using the video features of the sample video clips as model input parameters, the prediction model can be guided to learn how to predict the mapping relationship between the frame rate and the subjective image quality index based on the video features. Specifically, each sample video clip data has its unique content type and frame rate. The sample data set includes sample video clip data of multiple different content types and multiple different frame rates, which means that the sample data set includes sample video clip data of multiple different content types, and the sample video clip data of the same content type covers multiple different frame rates, and can even include multiple sample video clip data with exactly the same content but different frame rates, so that the model can learn the impact of frame rate on subjective image quality in a targeted manner.
[0051] As an example, the actual subjective image quality index can be an index used in related technologies to measure subjective image quality, or it can be obtained by designing and implementing subjective survey experiments. For example, sample video clip data of different frame rate versions under different scenes, content vertical categories, and picture complexity can be played to different users, and these users are asked to score the played video clips. The different scores of the same sample video clip data are then cleaned, counted, and processed. The final statistical value is used as the actual subjective image quality index of the sample video clip, which can reduce the impact of individual differences. Such subjective image quality indicators are different from traditional objective quality assessment methods. They pay more attention to the user's actual viewing experience, thereby providing optimization solutions that are closer to user needs.
[0052] As an example, you can also use methods such as cross-validation and hyperparameter tuning during training to improve the performance of the prediction model.
[0053] As an example, the prediction model includes at least one of the following: XGBoost, Random Forest, Support Vector Machine (SVM), Neural Network, or Reinforcement Learning Network. These models can all predict image quality-to-frame rate mapping information and can be selected as needed, enhancing the flexibility of the solution. The XGBoost model has strong fitting capabilities and efficient training speed, helping to more accurately predict subjective image quality indicators for different videos at different frame rates, thereby guiding frame rate optimization decisions. It should be understood that the XGBoost model can be used to perform classification, regression, and ranking tasks. In this disclosure, regression tasks are specifically performed. Random Forest can be used for classification and regression tasks, improving prediction accuracy and robustness by integrating multiple decision trees. Support Vector Machines are suitable for data classification and regression tasks in high-dimensional spaces and can handle nonlinear relationships. Deep learning algorithms, such as Convolutional Neural Networks (CNN), are particularly suitable for feature extraction from image and video data, capturing both local and global features. Recurrent Neural Networks (RNNs) are suitable for processing sequential data and capturing temporal dependencies. Reinforcement learning networks, such as Q-learning, learn optimal strategies through trial and error and are suitable for decision optimization in dynamic environments. Deep Q-Network (DQN) combines deep learning and reinforcement learning and is suitable for decision optimization in complex environments.
[0054] Optionally, step S104 includes performing transcoding processing on the video clip data, including frame insertion or frame drop processing, based on a comparison result between the preferred frame rate and the source stream frame rate of the video clip data, to obtain video stream data corresponding to the video clip data. By comparing the preferred frame rate determined in step S103 with the source stream frame rate of the video clip data, it is possible to determine whether the source stream frame rate is too low or too high, and then perform frame insertion processing (i.e., inserting new video frames to increase the frame rate) or frame drop processing (i.e., discarding some existing video frames to reduce the frame rate) accordingly, to transcode the video clip data into video stream data having the preferred frame rate.
[0055] For example, frame interpolation can involve inserting a new frame at any position, simply by calculating the Presentation Time Stamp (PTS) position of the inserted frame. Frame interpolation includes three modes: simple copy mode, in which the inserted frame is identical to the previous or next frame; linear averaging mode, in which a weighted average of the previous and next frames is taken to obtain the inserted frame; and motion compensation mode, in which the content of the inserted frame is predicted based on motion estimation and compensation algorithms. The PTS of the inserted frame also includes three modes: source frame PTS mode, in which the PTS of the original frame remains unchanged and the inserted frame is directly inserted based on the calculated position of the frame rate; smooth PTS mode, in which the PTS of the inserted frame is evenly spaced from the previous and next frames; and all-frame smoothing mode, in which the PTS of all frames, including the original and inserted frames, are evenly spaced.
[0056] As an example, frame loss processing includes discarding video frames other than key frames, that is, key frames are not allowed to be discarded. Of course, the PTS of the key frames will not be changed, that is, the timestamp of the key frames rendered after decoding on the playback side (such as the user end in the live broadcast scenario) will not be changed, to ensure that the absolute position offset of the original video content does not occur, thereby ensuring the key frame alignment of video stream data at different clarity levels.
[0057] Next, take the live broadcast scene as an example, combined with Figure 2 and Figure 3 A video transcoding method according to a specific embodiment of the present disclosure is introduced.
[0058] In this specific embodiment, Figure 2As shown in the figure, the live streaming system includes the host, the origin server, the Content Delivery Network (CDN), and the viewer. When the host pushes the stream—that is, the host sends the captured and encoded video stream to the origin server—the origin server performs transcoding scheduling and initiates a bypass task to calculate real-time video features. It then uses a prediction model to process these features to obtain the quality and frame rate mapping information for the current video clip. Based on this information, it makes frame rate optimization decisions. Finally, it transcodes the video stream sent by the host according to the optimal frame rate determined. The transcoded video stream is then output to the CDN and the viewer, ensuring users receive the best viewing experience.
[0059] like Figure 3 As shown, this specific embodiment provides a method for subjective quality assessment of video transmission and adaptive frame rate optimization decision-making, including the following aspects.
[0060] 1. Design and implement subjective survey experiments to obtain subjective experience scores (e.g., subjective smoothness scores) from respondents when watching live streams in various scenarios, content categories, and image complexity at different frame rates. This score, as a measure of actual subjective image quality, can more accurately reflect users' actual viewing experience, thereby optimizing video transmission strategies and improving the user viewing experience. These scenarios may include, but are not limited to, gaming competitions, outdoor activities, and livestreaming e-commerce.
[0061] 2. Calculate the content features and coding features of different videos. See the above introduction for details and will not be repeated here.
[0062] Third, we established a machine learning model based on the XGBoost model structure as a prediction model, fitting the mapping relationship between video features and subjective image quality indicators. Through methods such as cross-validation and hyperparameter tuning, we improved the model's prediction accuracy and generalization capabilities.
[0063] 4. Adopt precise frame rate optimization decisions. Based on the target subjective image quality indicators, predict the image quality frame rate mapping information for live videos and their characteristics under different scenarios, content verticals, and picture complexity. Based on this, select the frame rate that optimizes the subjective image quality indicators to ensure the best visual effects in different content scenarios.
[0064] This specific embodiment also provides a video transmission subjective quality assessment and adaptive frame rate optimization decision and adjustment system, which includes a decision model and an adjustment module and can provide a flexible frame rate adjustment mechanism.
[0065] The decision module implements the aforementioned decision-making method. Based on an offline subjective fluency dataset, it uses sample enhancement, feature cross-pollination, and other methods to predict quality-to-frame rate mapping information and determine the optimal frame rate using a variety of model-based and rule-based architectures. These architectures may include, but are not limited to, random forests and neural networks. This multi-architecture decision-making mechanism can better adapt to different video content and network environments.
[0066] The adjustment module implements adaptive interpolation / dropping video processing filters based on the source stream frame rate and the preferred frame rate decision. It flexibly inserts frames in multiple modes and, when dropping frames, selects video frames other than keyframes while maintaining the PTS of keyframes, ensuring that keyframes are not lost or shifted. Based on this, the optimal frame rate and interpolation / dropping modes can be updated in real time based on the decision module's optimized results. This allows for flexible response to network changes without losing keyframes, maintaining smooth and stable video playback.
[0067] This specific embodiment, through adaptive frame rate determination, can avoid unnecessary bandwidth consumption while ensuring video quality. In particular, under poor network conditions, frame rate reduction reduces data transmission volume, thereby conserving bandwidth resources, improving network utilization, and achieving efficient resource utilization. Furthermore, this specific embodiment is not only applicable to different video content and scenarios, but also adapts to various network environments, providing optimized video transmission solutions for both broadband and mobile networks, ensuring that users can enjoy high-quality video services in all circumstances, thus possessing broad applicability.
[0068] Figure 4 is a block diagram illustrating an exemplary video transcoding apparatus according to some embodiments of the present disclosure.
[0069] Reference Figure 4 The video transcoding device 400 includes an extraction unit 401 , a processing unit 402 , a determination unit 403 , and a transcoding unit 404 .
[0070] The extraction unit 401 may extract video features of the currently received video segment data.
[0071] Optionally, the video features include at least one of the following: content features, coding features, wherein the content features include at least one of the following: content type features, complexity features, aesthetic features, content type features include at least one of the following: motion scene, static scene, complexity features include at least one of the following: motion vector, color change feature, aesthetic features include at least one of the following: composition feature, color matching feature, coding features include at least one of the following: bit rate, noise.
[0072] The processing unit 402 may process the video features using a pre-trained prediction model to obtain image quality frame rate mapping information of the video segment data, wherein the image quality frame rate mapping information is used to represent a mapping relationship between a subjective image quality index and a frame rate.
[0073] Optionally, the prediction model includes at least one of the following: XGBoost model, random forest, support vector machine, neural network, reinforcement learning network.
[0074] Optionally, the prediction model is trained through the following steps: constructing a sample data set, wherein the sample data set includes sample video clip data of multiple different content types and multiple different frame rates; extracting video features of each sample video clip data, and determining the actual subjective picture quality index of each sample video clip data; using the video features as input parameters of the prediction model to be trained, and using the mapping relationship between the actual subjective picture quality index and the frame rate as the training target of the prediction model to be trained, training the prediction model to be trained, and obtaining a prediction model.
[0075] The determining unit 403 may determine the preferred frame rate of the video segment data according to the image quality frame rate mapping information and the target subjective image quality index.
[0076] The transcoding unit 404 may perform transcoding processing on the video segment data according to the preferred frame rate to obtain video stream data corresponding to the video segment data.
[0077] Optionally, the transcoding unit 404 may further perform transcoding processing including frame insertion or frame drop processing on the video segment data based on a comparison result between the preferred frame rate and the source stream frame rate of the video segment data to obtain video stream data corresponding to the video segment data.
[0078] Optionally, the frame dropping process includes dropping other video frames except the key frames.
[0079] Regarding the apparatus in the above embodiment, the specific manner in which each unit performs operations has been described in detail in the embodiment of the method, and will not be elaborated on here.
[0080] Figure 5 is a block diagram of an electronic device according to an exemplary embodiment of the present disclosure.
[0081] Reference Figure 5 The electronic device 500 includes at least one memory 501 and at least one processor 502, wherein the at least one memory 501 stores a set of computer-executable instructions. When the computer-executable instruction set is executed by the at least one processor 502, a video transcoding method according to an exemplary embodiment of the present disclosure is executed.
[0082] As an example, electronic device 500 may be a PC, tablet device, personal digital assistant, smartphone, or other device capable of executing the aforementioned set of instructions. Here, electronic device 500 is not necessarily a single electronic device, but may also be any collection of devices or circuits capable of executing the aforementioned instructions (or instruction set) individually or in combination. Electronic device 500 may also be part of an integrated control system or system manager, or may be configured as a portable electronic device that interfaces with local or remote devices (e.g., via wireless transmission).
[0083] In electronic device 500, processor 502 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.
[0084] The processor 502 can execute instructions or codes stored in the memory 501, wherein the memory 501 can also store data. Instructions and data can also be sent and received over a network via a network interface device, wherein the network interface device can use any known transmission protocol.
[0085] The memory 501 may be integrated with the processor 502, for example, by placing RAM or flash memory within an integrated circuit microprocessor or the like. Furthermore, the memory 501 may comprise a separate device, such as an external disk drive, a storage array, or any other storage device usable by a database system. The memory 501 and the processor 502 may be operatively coupled or may communicate with each other, for example, via an I / O port, a network connection, or the like, such that the processor 502 can access files stored in the memory.
[0086] In addition, the electronic device 500 may further include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.) All components of the electronic device 500 may be connected to each other via a bus and / or a network.
[0087] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium may also be provided, which, when the instructions in the computer-readable storage medium are executed by at least one processor, prompts the at least one processor to perform the video transcoding method according to the exemplary embodiment of the present disclosure. Examples of computer-readable storage media include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disk storage, hard disk drive (HDD), solid state drive (SSD), card storage (such as a multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk and any other device configured to store a computer program and any associated data, data files and data structures in a non-transitory manner and provide the computer program and any associated data, data files and data structures to a processor or computer so that the processor or computer can execute the computer program. The computer program in the above-mentioned computer-readable storage medium can be executed in an environment deployed in a computer device such as a client, a host, an agent device, a server, etc. In addition, in one example, the computer program and any associated data, data files and data structures are distributed on a networked computer system so that the computer program and any associated data, data files and data structures are stored, accessed and executed in a distributed manner by one or more processors or computers.
[0088] According to an exemplary embodiment of the present disclosure, a computer program product may also be provided. The computer program product includes computer instructions. When the computer instructions are executed by at least one processor, the at least one processor is prompted to perform the video transcoding method according to the exemplary embodiment of the present disclosure.
[0089] According to an exemplary embodiment of the present disclosure, a method for generating a bitstream may further be provided, including: generating a bitstream according to the video transcoding method of the exemplary embodiment of the present disclosure.
[0090] The description of the present disclosure has been presented for purposes of illustration and is not intended to be exhaustive or limited to the present disclosure. Many modifications, variations, and alternative embodiments will be apparent to one of ordinary skill in the art having the benefit of the teachings presented in the foregoing description and the associated drawings.
[0091] Unless otherwise specifically stated, the order of steps of the method according to the present disclosure is intended to be illustrative only, and the steps of the method according to the present disclosure are not limited to the order specifically described above, but can be changed according to actual circumstances. In addition, at least one of the steps of the method according to the present disclosure can be adjusted, combined, or deleted according to actual needs.
[0092] The examples are chosen and described in order to explain the principles of the present disclosure and to enable others skilled in the art to understand the various embodiments of the present disclosure and to best utilize the basic principles and various embodiments with various modifications as are suited to the particular use contemplated. Therefore, it will be understood that the scope of the present disclosure is not limited to the specific examples of the embodiments disclosed and that modifications and other embodiments are intended to be included within the scope of the present disclosure.
Claims
1. A video transcoding method, characterized in that: The video transcoding method comprises: Extracting video features of the currently received video clip data; Processing the video features using a pre-trained prediction model to obtain image quality and frame rate mapping information of the video clip data, wherein the image quality and frame rate mapping information is used to represent a mapping relationship between a subjective image quality indicator and a frame rate; determining a preferred frame rate for the video clip data according to the image quality frame rate mapping information and a target subjective image quality indicator; The video segment data is transcoded according to the preferred frame rate to obtain video stream data corresponding to the video segment data.
2. The video transcoding method according to claim 1, wherein: The prediction model is trained through the following steps: Constructing a sample data set, wherein the sample data set includes sample video segment data of multiple different content types and multiple different frame rates; Extracting the video features of each sample video segment data and determining an actual subjective image quality index of each sample video segment data; The video features are used as input parameters of the prediction model to be trained, the mapping relationship between the actual subjective image quality index and the frame rate is used as the training target of the prediction model to be trained, and the prediction model to be trained is trained to obtain the prediction model.
3. The video transcoding method according to claim 1, wherein: The step of transcoding the video segment data according to the preferred frame rate to obtain video stream data corresponding to the video segment data includes: According to the comparison result between the preferred frame rate and the source stream frame rate of the video segment data, transcoding processing including frame insertion processing or frame drop processing is performed on the video segment data to obtain video stream data corresponding to the video segment data.
4. The video transcoding method according to claim 3, wherein: The frame dropping process includes dropping other video frames except key frames.
5. The video transcoding method according to any one of claims 1 to 4, wherein: The video features include at least one of the following: content features, coding features, wherein the content features include at least one of the following: content type features, complexity features, and aesthetic features, the content type features include at least one of the following: motion scenes and static scenes, the complexity features include at least one of the following: motion vectors and color change features, the aesthetic features include at least one of the following: composition features and color matching features, and the coding features include at least one of the following: bit rate and noise; and / or The prediction model includes at least one of the following: XGBoost model, random forest, support vector machine, neural network, reinforcement learning network.
6. A video transcoding device, characterized in that: The video transcoding device comprises: an extraction unit configured to extract video features of currently received video clip data; a processing unit configured to process the video features using a pre-trained prediction model to obtain image quality and frame rate mapping information of the video segment data, wherein the image quality and frame rate mapping information is used to represent a mapping relationship between a subjective image quality indicator and a frame rate; a determining unit configured to determine a preferred frame rate of the video segment data according to the image quality frame rate mapping information and a target subjective image quality index; The transcoding unit is configured to perform transcoding processing on the video segment data according to the preferred frame rate to obtain video stream data corresponding to the video segment data.
7. An electronic device, characterized in that: include: at least one processor; at least one memory storing computer-executable instructions, When the computer executable instructions are executed by at least one processor, the computer executable instructions prompt the at least one processor to execute the video transcoding method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that When the instructions in the computer-readable storage medium are executed by at least one processor, the at least one processor is prompted to perform the video transcoding method according to any one of claims 1 to 5.
9. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by at least one processor, the video transcoding method according to any one of claims 1 to 5 is implemented.
10. A method for generating a bit stream, characterized in that: include: A bit stream is generated according to the video transcoding method according to any one of claims 1 to 5.
Citation Information
Cited By
Data processing method and device, equipment, storage medium and program product
CN121486595A