Video processing method and device, electronic equipment, storage medium and program product
By comprehensively considering the multi-dimensional features of video in the evaluation model, adjusting the brightness, contrast, saturation, and color temperature of the video, the problem of insufficient improvement in video quality and aesthetics in existing technologies is solved, and the user's subjective experience and network adaptability are optimized.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
- Filing Date
- 2026-02-11
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies are insufficient to comprehensively improve video quality and aesthetics, neglect user subjective experience, and cannot adapt to the adjustment of visual parameters in real live streaming scenarios or content.
By comprehensively considering multiple dimensions of video features, such as content features, complexity features, aesthetic features, and coding features, the evaluation model is used to adjust multiple visual parameters of the video, such as brightness, contrast, saturation, and color temperature, to optimize video quality frame by frame.
It achieves optimal video quality and user experience, adapts to various live streaming scenarios and network environments, and ensures smooth and stable video playback.
Smart Images

Figure CN122053883A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of video processing, and more particularly to a video processing method and apparatus, electronic device, storage medium, and program product. Background Technology
[0002] With the development of network technology, live streaming has become one of the most influential application scenarios in the multimedia field. In order to serve millions, tens of millions, or even hundreds of millions of heterogeneous network users and provide them with high-definition, high-quality live streaming services, most content providers optimize and adjust the brightness and color of the live stream to avoid overexposure or underexposure, graying colors, or color differences, thereby improving the picture quality and aesthetic appeal of the live stream content and maximizing the user's subjective viewing experience.
[0003] Currently, while some methods exist to adjust the visual parameters of video images based on network conditions, these methods cannot comprehensively improve video quality and aesthetics. Furthermore, they neglect the user's subjective experience, resulting in situations where, although video images are adjusted, the actual image quality is not improved, and the user's viewing experience is significantly diminished. Therefore, current methods are ill-suited to addressing the visual parameter issues of various real-world live streaming scenarios or content. Summary of the Invention
[0004] This disclosure provides a video processing method and apparatus, electronic device, storage medium, and program product to at least solve the problem that related technologies are unable to cope with the visual parameters of various real live streaming scenarios or real live streaming content.
[0005] According to a first aspect of the present disclosure, a video processing method is provided, comprising: for each frame in a streaming video, performing the following adjustments: inputting a video segment containing the current frame and features of the video segment in multiple dimensions into an evaluation model to obtain viewing experience information for multiple visual parameters of the current frame, wherein the video segment is a segment of video in the streaming video and the current frame is the last frame of the video segment; determining an adjustment strategy for the current frame based on the viewing experience information of the multiple visual parameters of the current frame; adjusting the multiple visual parameters of the current frame respectively based on the adjustment strategy; and performing live streaming based on the adjusted streaming video.
[0006] Optionally, the evaluation model is obtained as follows: a training dataset is obtained, which includes multiple video segments and real viewing experience information for multiple visual parameters of the last frame in each video segment; the multiple video segments and the features of each video segment in multiple dimensions are input into the initial evaluation model to obtain the predicted viewing experience information for multiple visual parameters of the last frame in each video segment; the initial evaluation model is updated based on the predicted viewing experience information and the real viewing experience information to obtain the evaluation model.
[0007] Optionally, based on viewing experience information of multiple visual parameters of the current frame, an adjustment strategy for the current frame is determined, including: for each visual parameter, determining an adjustment value for the current visual parameter of the current frame based on the viewing experience information of the current visual parameter; and determining the adjustment values of all visual parameters as the adjustment strategy for the current frame.
[0008] Optionally, based on the adjustment strategy, multiple visual parameters of the current frame are adjusted separately, including: directly adjusting multiple visual parameters of the current frame according to the adjustment strategy to obtain the adjusted current frame; or, inputting the adjustment strategy and the video segment into the adjustment model to obtain the adjusted current frame.
[0009] Optionally, the multi-dimensional features include at least two of the following features: content features, complexity features, aesthetic features, and coding features of the video segment, wherein the content features contain content information of the video segment, the complexity features contain visual complexity information of the video segment, the aesthetic features contain information about at least one aesthetic element of the video segment, and the coding features contain coding information of the video segment.
[0010] Optionally, the multiple visual parameters include at least two of the following visual parameters: brightness, contrast, saturation, and color temperature.
[0011] According to a second aspect of the present disclosure, a video processing apparatus is provided, comprising: an adjustment unit configured to perform the following adjustments for each frame in a streaming video: inputting a video segment containing the current frame and features of the video segment in multiple dimensions into an evaluation model to obtain viewing experience information for multiple visual parameters of the current frame, wherein the video segment is a segment of video in the streaming video and the current frame is the last frame of the video segment; determining an adjustment strategy for the current frame based on the viewing experience information for the multiple visual parameters of the current frame; and a live streaming unit configured to perform live streaming based on the adjusted streaming video.
[0012] Optionally, the evaluation model is obtained as follows: a training dataset is obtained, which includes multiple video segments and real viewing experience information for multiple visual parameters of the last frame in each video segment; the multiple video segments and the features of each video segment in multiple dimensions are input into the initial evaluation model to obtain the predicted viewing experience information for multiple visual parameters of the last frame in each video segment; the initial evaluation model is updated based on the predicted viewing experience information and the real viewing experience information to obtain the evaluation model.
[0013] Optionally, the adjustment unit is also configured to, for each visual parameter, determine the adjustment value of the current visual parameter for the current frame based on the viewing experience information of the current visual parameter; and determine the adjustment values of all visual parameters as the adjustment strategy for the current frame.
[0014] Optionally, the adjustment unit is also configured to directly adjust multiple visual parameters of the current frame according to the adjustment strategy to obtain the adjusted current frame; or, the adjustment strategy and video segment are input into the adjustment model to obtain the adjusted current frame.
[0015] Optionally, the multi-dimensional features include at least two of the following features: content features, complexity features, aesthetic features, and coding features of the video segment, wherein the content features contain content information of the video segment, the complexity features contain visual complexity information of the video segment, the aesthetic features contain information about at least one aesthetic element of the video segment, and the coding features contain coding information of the video segment.
[0016] Optionally, the multiple visual parameters include at least two of the following visual parameters: brightness, contrast, saturation, and color temperature.
[0017] According to a third aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and at least one memory storing computer-executable instructions, wherein the computer-executable instructions, when executed by the at least one processor, cause the at least one processor to perform the video processing method as described above.
[0018] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided that, when instructions in the computer-readable storage medium are executed by at least one processor, causes at least one processor to perform the video processing method as described above.
[0019] According to a fifth aspect of the present disclosure, a computer program product is provided, including computer instructions that, when executed by a processor, implement the video processing method described above.
[0020] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects: According to the video processing method, apparatus, electronic device, storage medium, and program product disclosed herein, multiple visual parameters of the current frame can be uniformly optimized and adjusted based on the multi-dimensional features of the video segment containing the current frame. This enables the overall video quality and user subjective experience to reach their optimal levels. Furthermore, this disclosure can adjust multiple visual parameters of the streaming video frame by frame, better responding to network fluctuations and ensuring the smoothness and stability of video playback. Therefore, this disclosure can address the visual parameter issues of various real-world live streaming scenarios or content.
[0021] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0022] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0023] Figure 1 This is a flowchart illustrating a video processing method according to exemplary embodiments of the present disclosure; Figure 2 This is a schematic diagram illustrating a live streaming process according to exemplary embodiments of the present disclosure; Figure 3 This is a schematic diagram illustrating a video optimization process according to exemplary embodiments of the present disclosure; Figure 4 This is a block diagram illustrating a video processing apparatus according to exemplary embodiments of the present disclosure; Figure 5 This is a diagram illustrating a computing environment coupled to a user interface according to exemplary embodiments of the present disclosure. Detailed Implementation
[0024] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0025] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following examples do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0026] It should be noted that the phrase "at least one of several items" in this disclosure refers to three parallel cases: "any one of the several items", "a combination of any number of the several items", and "all of the several items". For example, "including at least one of A and B" includes the following three parallel cases: (1) including A; (2) including B; (3) including A and B. Another example is "performing at least one of step one and step two", which means the following three parallel cases: (1) performing step one; (2) performing step two; (3) performing both step one and step two.
[0027] Adjusting the visual parameters of video images typically involves adjusting dimensions such as brightness, contrast, saturation, and color temperature. For example, too low brightness will make the image dark and difficult to see, while too high brightness will give the user a feeling of overexposure; improper contrast will make the image blurry and the edges shallow; improper saturation will make the image appear grayish or greenish; and excessively high or low color temperature will cause color shifts and color differences.
[0028] However, in live streaming scenarios, on the one hand, the real-time requirements of the live content necessitate real-time decisions regarding adjustments to the visual parameters of the live stream; on the other hand, limitations of the live streaming system mean that the content may frequently change significantly, and the system cannot anticipate these changes in complexity and adjust its image adjustment strategies accordingly. Therefore, current methods for adjusting the visual parameters of live streams typically utilize fixed parameters or rules designed through offline analysis. As the complexity of the live content increases, how to more accurately design visual parameter adjustment strategies for the live stream under limited resource conditions (transcoding resources, storage resources) to optimize the overall experience for all users is one of the most critical issues in the audio-visual field.
[0029] Currently, although there are some methods to adjust the visual parameters of video images based on network conditions, two methods are briefly introduced below: 1) A method for calculating target image features by generating a normalized target histogram and adaptively adjusting the tone mapping function by combining a brightness assessment value. Specifically: a normalized target histogram of the display image that meets preset conditions is generated based on a preset scene of the video image; target image features are calculated based on the normalized target histogram; a brightness assessment value is calculated based on the video image and its exposure information; the overall brightness control parameters in the tone mapping function are confirmed based on the brightness assessment value; and tone mapping is performed on the video image using the confirmed tone mapping function.
[0030] 2) Adjusting video brightness using a trained model. Specifically: Obtain a training dataset, which includes multiple sets of training data. Each set includes visual information from historical video playback on the terminal, as well as terminal data. The terminal location includes whether the terminal is indoors, outdoors, or in motion. Based on the training dataset, obtain a simulation matrix. This simulation matrix is used to simulate the missing probability of target type data in the training dataset. Based on the training dataset, train the video brightness prediction model. During training, process the output matrix of the attention layer in the video brightness prediction model based on the simulation matrix to obtain the input matrix of the next layer. This allows for the training of a video brightness prediction model with high prediction accuracy, which can accurately predict the video brightness corresponding to different scenes, and adjust the video brightness accordingly.
[0031] However, the above method 1) has the following drawbacks: A) Incomplete input features: This method only refers to the exposure information in the video image, considers only a single type of feature, and has overly simplistic preset scenarios, making it difficult to make appropriate adjustments for various actual live broadcasts or video content. B) Limited adjustment dimensions: This method only analyzes and adjusts the visual parameter of brightness, making it difficult to improve the overall video quality and the user's subjective experience. C) Simple adjustment methods: This method only uses tone mapping functions as a video image adjustment method, making it difficult to adapt to changing scenes and complex video content.
[0032] The above method 1) has the following drawbacks: A) Limited dataset: This method only refers to visual information from video images, considering only a limited range of features, making it difficult to make appropriate adjustments for various actual live streams or video content. B) Limited adjustment dimension: This method only analyzes and adjusts the visual parameter of brightness, making it difficult to improve the overall video quality and user experience. C) Simple adjustment method: This method only adjusts by directly specifying the target video brightness, which is too simplistic and prone to introducing overexposure and other adverse phenomena, making it difficult to adapt to changing scenes and complex video content. D) Simple prediction model: This method only uses a simulated matrix prediction model to predict the target video brightness. The model structure and prediction results are relatively simple, making it difficult to make reasonable analyses and decisions for various complex scenes and video content.
[0033] To address the aforementioned issues, this disclosure comprehensively considers multiple dimensions of video characteristics, such as content features, complexity features, aesthetic features, and coding features, enabling better differentiation of various scenes and content. By using these multiple dimensions as input to the evaluation model, it can more comprehensively analyze and predict the actual performance and degree of defects of video images under visual parameters. Furthermore, this disclosure comprehensively considers multiple visual parameters, such as brightness, contrast, saturation, and color temperature, and uniformly optimizes and adjusts these parameters to achieve optimal overall video quality and user subjective experience.
[0034] Furthermore, this disclosure also obtains the actual subjective experience scores of survey subjects when watching live videos in various scenarios, content categories, and image complexity under different visual parameters. The evaluation model is trained using these actual subjective experience scores as the target to fit the mapping relationship between video features and subjective experience scores. This allows the evaluation model to comprehensively consider information from offline datasets (i.e., mapping relationships) and actual online live streaming data when it is actually used, so as to output a subjective experience score that more accurately reflects the user's actual viewing experience. Based on this subjective experience score, the optimal adjustment strategy for adjusting the live video can be obtained to ensure that the best visual effect can be provided under different network conditions, thereby improving the user's viewing experience.
[0035] Furthermore, this disclosure allows for the adjustment of visual parameters of a video frame by frame, and also provides various modes for adjusting visual parameters, such as directly adjusting the visual parameters of the video, adjusting visual parameters using a model, etc., so that visual parameters can be flexibly adjusted and appropriate adjustments can be made to various video images.
[0036] The video processing method and apparatus, electronic device, storage medium, and program product according to exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings.
[0037] Figure 1 This is a flowchart illustrating a video processing method according to exemplary embodiments of the present disclosure, such as... Figure 1 As shown, the video processing method includes the following steps: In step S101, for each frame in the streaming video, the following adjustments are performed: the video segment containing the current frame and the features of the video segment in multiple dimensions are input into the evaluation model to obtain viewing experience information for multiple visual parameters of the current frame, wherein the video segment is a segment of video in the streaming video and the current frame is the last frame of the video segment; based on the viewing experience information of multiple visual parameters of the current frame, an adjustment strategy for the current frame is determined; based on the adjustment strategy, the multiple visual parameters of the current frame are adjusted respectively.
[0038] In step S102, live streaming is performed based on the adjusted push video.
[0039] As an example, the evaluation model mentioned above can be a machine learning model based on the XGBoost model structure, or it can be a machine learning model based on other machine learning algorithms, deep learning algorithms, reinforcement learning algorithms, etc. Specifically, other machine learning algorithms can include, but are not limited to, Random Forest, Support Vector Machine (SVM), etc. Random Forest can be used for classification and regression tasks, improving the accuracy and robustness of predictions by integrating multiple decision trees; SVM is suitable for classification and regression tasks in high-dimensional space data and can handle nonlinear relationships. Deep learning algorithms can include, but are not limited to, Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), etc. CNN is particularly suitable for feature extraction of image and video data, and can capture local and global features; RNN is suitable for processing sequence data and can capture temporal dependencies. Reinforcement learning algorithms can include, but are not limited to, Q-learning, Deep Q-Network (DQN), etc. Among them, Q-learning learns the optimal strategy through trial and error and is suitable for decision optimization in dynamic environments; DQN combines deep learning and reinforcement learning and is suitable for decision optimization in complex environments.
[0040] It should be noted that when evaluating machine learning models based on the XGBoost model structure, the experiments showed higher prediction accuracy and faster training speed. Therefore, the XGBoost model structure has significant advantages in practical applications.
[0041] As an example, the viewing experience information mentioned above could be the Mean Opinion Score, which is not limited in this disclosure. It should be noted that the Mean Opinion Score (MOS) is a standard score for measuring the subjective quality of media such as audio / video. Its core is the Absolute Category Rating (ACR) of 1-5 points, which is specified by standards such as ITU-T P.800 and is widely used in telecommunications, VoIP, audio coding and other fields.
[0042] According to an exemplary embodiment of this disclosure, the features in step S101 above, encompassing multiple dimensions, include, but are not limited to, at least two of the following features: content features, complexity features, aesthetic features, and encoding features of the video segment. The content features include content information of the video segment; the complexity features include information about the image complexity of the video segment; the aesthetic features include information about at least one aesthetic element of the video segment; and the encoding features include encoding information of the video segment. Through this embodiment, by comprehensively considering the content features, complexity features, aesthetic features, and encoding features of the video, the characteristics of the video can be described more comprehensively, enabling better differentiation between various scenes and content, and providing richer data support for subsequent decisions on brightness and color optimization of the live stream.
[0043] As an example, the content information of the aforementioned video segment indicates the type of video content, such as whether the video is in a moving scene or a static scene. More specifically, it may be in a game scene, a conversation scene, a sales scene, or a game-playing scene, etc., which this disclosure does not limit. The aforementioned image complexity information indicates the image complexity of the video, which may include, but is not limited to, motion vectors, color changes, etc. The aforementioned aesthetic elements may include, but are not limited to, composition, color matching, etc. The aforementioned encoding information may be the encoding parameters of the video, which may include, but are not limited to, bitrate, noise, etc.
[0044] According to exemplary embodiments of this disclosure, the plurality of visual parameters includes, but is not limited to, at least two of the following visual parameters: brightness, contrast, saturation, and color temperature. By adjusting these four visual parameters, a better visual effect in the video can be ensured.
[0045] As an example, visual parameters may also include color gamut, color depth, white balance, etc., which are not limited in this disclosure.
[0046] According to an exemplary embodiment of this disclosure, the evaluation model in step S101 above can be obtained in the following manner: obtaining a training dataset, wherein the training dataset includes multiple video segments and real viewing experience information for multiple visual parameters of the last frame in each video segment; inputting the multiple video segments and the features of multiple dimensions of each video segment into the initial evaluation model respectively to obtain estimated viewing experience information for multiple visual parameters of the last frame in each video segment; updating the initial evaluation model based on the estimated viewing experience information and the real viewing experience information to obtain the evaluation model.
[0047] This embodiment acquires real viewing experience information for multiple video segments under multiple visual parameters, and uses this real viewing experience information as the target to train an initial evaluation model, resulting in a trained evaluation model. Compared to related technologies, this embodiment places greater emphasis on the user's actual viewing experience. The trained evaluation model can provide a more accurate reflection of the user's actual viewing experience, thus enabling subsequent optimization solutions that are closer to user needs to be provided based on accurate subjective experience scores, thereby improving the user's viewing experience.
[0048] As an example, a subjective survey experiment can be designed and implemented to obtain the actual subjective experience scores (i.e., the aforementioned actual viewing experience information) of multiple survey subjects when watching live stream videos in various scenarios, content categories, and levels of visual complexity, under different visual parameter values. The aforementioned different scenarios may include, but are not limited to, e-sports, outdoor activities, and live-streaming e-commerce; the aforementioned survey subjects may be ordinary people or professionals with specialized knowledge, and this disclosure does not limit this.
[0049] Then, using the parameters of the live video and its multiple dimensions, the initial evaluation model is input to obtain the estimated subjective experience score (i.e., the estimated viewing experience information) for the last frame of each live video. The loss between the estimated subjective experience score and the corresponding real subjective experience score for the last frame of each live video is calculated. The parameters of the initial evaluation model are adjusted in a way that minimizes the loss to obtain the evaluation model, which fits the mapping relationship between the live video features and the subjective experience score.
[0050] It should be noted that cross-validation, hyperparameter tuning, and other methods can be used during training to improve the model's prediction accuracy and generalization ability.
[0051] According to an exemplary embodiment of this disclosure, the step S101 above, which determines the adjustment strategy for the current frame based on the viewing experience information of multiple visual parameters of the current frame, may include: for each visual parameter, determining an adjustment value for the current visual parameter of the current frame based on the viewing experience information of the current visual parameter; and determining the adjustment values of all visual parameters as the adjustment strategy for the current frame. This embodiment allows for convenient and rapid determination of the adjustment strategy for the current frame.
[0052] As an example, suppose the live streaming scenario is a game live stream, the visual parameter is brightness, and the MOS score is 3. This means that the brightness is insufficient at this time, and the brightness can be appropriately increased, such as by increasing the original brightness value by 10%. Then, +10% of the original brightness value will be used as the brightness adjustment value. This adjustment value means increasing the original brightness value by 10%. In other words, the adjustment strategy for the current frame includes increasing the original brightness value by 10% on the basis of the original brightness value.
[0053] According to an exemplary embodiment of this disclosure, step S101, which adjusts multiple visual parameters of the current frame based on an adjustment strategy, may include: directly adjusting multiple visual parameters of the current frame according to the adjustment strategy to obtain the adjusted current frame; or, inputting the adjustment strategy and the video segment into an adjustment model to obtain the adjusted current frame. Through this embodiment, video parameters can be flexibly adjusted using multiple modes.
[0054] As an example, the above-mentioned adjustment of multiple visual parameters of the current frame directly according to the adjustment strategy can be achieved by directly providing the adjusted parameter values to adjust the corresponding visual parameters of the current frame, or by providing the adjustment value in the adjustment strategy, inputting the adjustment value into the function to obtain the adjusted parameter value, and then adjusting the corresponding visual parameters of the current frame through the adjusted parameters. This disclosure does not limit this approach.
[0055] As an example, the adjustment model described above can be a LUT (Look-Up Table) model, which is not limited in this disclosure. Assuming the adjustment model is a LUT (Look-Up Table) model, the adjustment strategy and video segment can be input into the LUT model, and the model directly outputs the adjusted current frame.
[0056] To better understand this disclosure, the following is in conjunction with... Figure 2 and Figure 3 Provide a systematic explanation.
[0057] Figure 2 This demonstrates a live streaming process, such as Figure 2 As shown, after the broadcaster starts streaming, the streamed video (i.e., the encoded audio and video data) is uploaded to the origin server. After transcoding and scheduling, the streamed video is optimized, which means it enters the next stage. Figure 2 The bypass task in the process first calculates the content features, complexity features, aesthetic features, and encoding features of each video segment corresponding to each frame in the streaming video. Then, multiple video segments and their corresponding multi-dimensional features are input into the evaluation model to obtain the viewing experience information of each visual parameter in brightness, contrast, saturation, and color temperature for each frame. Based on the viewing experience information of each visual parameter, the optimal adjustment strategy for each frame is obtained. Next, the optimal adjustment strategy is used to optimize the transcoded streaming video frame by frame to obtain the optimized streaming video. The optimized streaming video is then transcoded for live streaming and returned to the origin server. Finally, when the viewer opens the live streaming page, a pull request is sent to the origin server. The origin server sends the optimized streaming video to the viewer via a Content Delivery Network (CDN) according to the received pull request. After buffering locally, the viewer's player decodes and renders the picture in real time, forming a smooth live streaming viewing experience.
[0058] Figure 3 This demonstrates a video optimization workflow, namely Figure 2 The bypass task in the model first calculates the content features, complexity features, aesthetic features, and coding features of each video segment corresponding to each frame in the streaming video. Then, the features of multiple video segments and their corresponding content, complexity, aesthetic, and coding features are input into the XGBoost model, outputting a defect score vector. This defect score vector includes brightness defect score, contrast defect score, saturation defect score, and color temperature defect score, thus obtaining the viewing experience information for each visual parameter in brightness, contrast, saturation, and color temperature of each frame. Based on the viewing experience information of each visual parameter, the defective visual parameters are identified, and the optimal adjustment strategy for each frame is obtained. Then, the optimal adjustment strategy is used to optimize the transcoded streaming video frame by frame, resulting in an optimized streaming video that achieves the optimal user subjective experience score.
[0059] In summary, this embodiment provides a method for adaptive brightness and color decision-making and transmission of video based on user subjective image quality and aesthetic perception. It utilizes the XGBoost model structure to establish a mapping relationship between video features and subjective experience scores. This allows for the selection of the optimal adjustment strategy based on specific video features (such as content features, complexity features, aesthetic features, and encoding features). With the user's subjective experience score as the target, it comprehensively adjusts the brightness, contrast, saturation, and color temperature of the video image to ensure optimal visual effects under different network conditions, thus accurately and effectively solving the problem of subjective experience optimization during network video transmission. Specifically: This embodiment can be optimized based on the user's subjective experience. That is, by designing and implementing subjective survey experiments, the subjective experience scores of live videos under different brightness, contrast, saturation and color temperature versions under different scenarios, content categories and image complexity are obtained, so that more attention is paid to the user's actual viewing experience, thereby providing an optimization solution that is closer to the user's needs. This embodiment can perform analysis based on multi-dimensional features, that is, it not only considers the content features of the video, but also comprehensively considers the complexity features, aesthetic features and coding features. This multi-dimensional feature analysis method can more comprehensively describe the characteristics of the video and provide richer data support for subsequent decisions on brightness and color optimization of live broadcast images.
[0060] This embodiment employs an advanced machine learning model—the XGBoost model—which establishes a mapping relationship between video features and subjective experience scores. The XGBoost model boasts powerful fitting capabilities and efficient training speed, enabling it to more accurately predict the subjective experience scores of different videos under varying brightness, contrast, saturation, and color temperature, thereby guiding brightness and color optimization decisions.
[0061] It should be noted that this disclosure is not only applicable to different video content and scenarios, but also adaptable to various network environments. Whether it is a broadband network or a mobile network, it can provide optimized video transmission solutions to ensure that users can enjoy high-quality video services under any circumstances.
[0062] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.
[0063] Figure 4 This is a block diagram illustrating a video processing apparatus according to exemplary embodiments of the present disclosure. (Refer to...) Figure 4 The device includes an adjustment unit 40 and a live broadcast unit 42.
[0064] The adjustment unit 40 is configured to perform the following adjustments for each frame in the streaming video: inputting a video segment containing the current frame and features of the video segment in multiple dimensions into the evaluation model to obtain viewing experience information for multiple visual parameters of the current frame, wherein the video segment is a segment of video in the streaming video and the current frame is the last frame of the video segment; based on the viewing experience information of multiple visual parameters of the current frame, determining the adjustment strategy for the current frame; the live streaming unit 42 is configured to perform live streaming based on the adjusted streaming video.
[0065] According to an exemplary embodiment of this disclosure, the evaluation model is obtained as follows: a training dataset is obtained, wherein the training dataset includes multiple video segments and real viewing experience information for multiple visual parameters of the last frame in each video segment; the multiple video segments and the features of multiple dimensions of each video segment are respectively input into an initial evaluation model to obtain estimated viewing experience information for multiple visual parameters of the last frame in each video segment; the initial evaluation model is updated based on the estimated viewing experience information and the real viewing experience information to obtain the evaluation model.
[0066] According to an exemplary embodiment of this disclosure, the adjustment unit 40 is further configured to, for each visual parameter, determine an adjustment value for the current visual parameter for the current frame based on viewing experience information of the current visual parameter; and determine the adjustment values of all visual parameters as an adjustment strategy for the current frame.
[0067] According to an exemplary embodiment of this disclosure, the adjustment unit 40 is further configured to directly adjust multiple visual parameters of the current frame according to the adjustment strategy to obtain the adjusted current frame; or, input the adjustment strategy and the video segment into the adjustment model to obtain the adjusted current frame.
[0068] According to an exemplary embodiment of this disclosure, the features of multiple dimensions include at least two of the following features: content features, complexity features, aesthetic features, and encoding features of the video segment, wherein the content features contain content information of the video segment, the complexity features contain image complexity information of the video segment, the aesthetic features contain information of at least one aesthetic element of the video segment, and the encoding features contain encoding information of the video segment.
[0069] According to exemplary embodiments of this disclosure, the plurality of visual parameters includes at least two of the following visual parameters: brightness, contrast, saturation, and color temperature.
[0070] As an example, this disclosure provides a video transmission subjective quality assessment and adaptive brightness and color optimization decision and adjustment system, including: The decision-making module, based on offline datasets of subjective image quality and aesthetic perception, as well as online datasets of actual live stream video and image sources, uses various model-based and rule-based architectures, including sample augmentation and feature cross-referencing, to predict subjective experience scores and determine the optimal adjustment strategy. This decision-making mechanism can better adapt to different video content and network environments. It should be noted that the aforementioned architectures may include, but are not limited to, random forests and neural networks.
[0071] The adjustment module applies video processing filters to adjust video brightness, contrast, saturation, and color temperature frame by frame, and can flexibly adjust these parameters according to various modes. It's worth noting that the adjustment module can update its adjustment strategy in real time based on the optimal adjustment strategy from the decision module. This dynamic adjustment mechanism better handles network fluctuations, ensuring smooth and stable video playback.
[0072] Therefore, in this embodiment, the decision-making module predicts the user's subjective experience score and decides the optimal adjustment strategy based on the offline dataset of image quality and aesthetic perception, as well as the online actual live source video and image dataset; while the adjustment module adjusts the video brightness and color in real time according to the decision results, flexibly responding to network changes and maintaining the clarity, image quality and user perception of video playback.
[0073] Figure 5 A computing environment 510 coupled to a user interface 550 is shown. The computing environment 510 may be part of a data processing server. The computing environment 510 includes a processor 520, memory 530, and input / output (I / O) interface 540.
[0074] Processor 520 typically controls the overall operation of computing environment 510, such as operations associated with display, data acquisition, data communication, and image processing. Processor 520 may include one or more processors for executing instructions to perform all or some of the steps in the methods described above. Furthermore, processor 520 may include one or more modules that facilitate interaction between processor 520 and other components. The processor may be a central processing unit (CPU), microprocessor, microcontroller, graphics processing unit (GPU), etc.
[0075] Memory 530 is configured to store various types of data to support the operation of computing environment 510. Memory 530 may include predefined software 532. Examples of such data include instructions for any application or method operating on computing environment 510, video datasets, image data, etc. Memory 530 can be implemented using any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0076] I / O interface 540 provides an interface between processor 520 and peripheral interface modules (such as keyboard, click wheel, buttons, etc.). Buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. I / O interface 540 can be coupled to encoders and decoders.
[0077] In an embodiment, the computing environment 510 may be implemented by one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), FPGAs, GPUs, controllers, microcontrollers, microprocessors, or other electronic components for performing the methods described above.
[0078] According to embodiments of the present disclosure, an electronic device may be provided, the electronic device including at least one memory and at least one processor, wherein the at least one memory stores a set of computer-executable instructions, and when the set of computer-executable instructions is executed by the at least one processor, a video processing method according to embodiments of the present disclosure is performed.
[0079] As an example, the electronic device can be a PC, tablet, personal digital assistant, smartphone, or other device capable of executing the aforementioned set of instructions. Here, the electronic device 1000 is not necessarily a single electronic device, but can be any collection of devices or circuits capable of executing the aforementioned instructions (or instruction sets) individually or in combination. The electronic device can also be part of an integrated control system or system manager, or can be configured to interconnect with a portable electronic device locally or remotely (e.g., via wireless transmission) through an interface.
[0080] In addition, electronic devices may include video displays (such as liquid crystal displays) and user interaction interfaces (such as keyboards, mice, touch input devices, etc.). All components of the electronic device may be interconnected via buses and / or networks.
[0081] According to embodiments of this disclosure, a computer-readable storage medium may also be provided, wherein when instructions in the computer-readable storage medium are executed by at least one processor, the at least one processor causes the at least one processor to perform the video processing method of the embodiments of this disclosure. Examples of computer-readable storage media include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc storage, hard disk drive (HDD), solid-state drive (SSD), card storage (such as multimedia cards, secure digital (SD) cards, or ultra-fast digital (XD) cards), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid-state drive, and any other device configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and to provide the computer program and any associated data, data files, and data structures to a processor or computer so that the processor or computer can execute the computer program. The computer program in the aforementioned computer-readable storage medium can run in an environment deployed in computer devices such as clients, hosts, agent devices, servers, etc. Furthermore, in one example, the computer program and any associated data, data files, and data structures are distributed across a networked computer system, such that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner through one or more processors or computers.
[0082] According to embodiments of this disclosure, a computer program product is also provided, comprising, for example, a plurality of programs stored in a memory 530, which can be executed by a processor 520 in a computing environment 510 to perform the methods described above. For example, the computer program product may include a non-transitory computer-readable storage medium.
[0083] Unless otherwise specifically stated, the order of steps in the method according to this disclosure is intended to be illustrative only, and the steps of the method according to this disclosure are not limited to the specific order described above, but may be changed according to actual circumstances. Furthermore, at least one step in the method according to this disclosure may be adjusted, combined, or omitted as needed.
[0084] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the appended claims.
[0085] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A video processing method, characterized in that, include: For each frame in the streaming video, perform the following adjustments: The evaluation model is input with a video segment containing the current frame and features of the video segment in multiple dimensions to obtain viewing experience information for multiple visual parameters of the current frame. The video segment is a segment of the streaming video and the current frame is the last frame of the video segment. Based on the viewing experience information of multiple visual parameters of the current frame, an adjustment strategy is determined for the current frame; Based on the adjustment strategy, multiple visual parameters of the current frame are adjusted respectively; Live streaming will be conducted based on the adjusted push video.
2. The video processing method as described in claim 1, characterized in that, The evaluation model is obtained in the following way: Obtain a training dataset, wherein the training dataset includes multiple video segments and real viewing experience information for multiple visual parameters of the last frame in each video segment; The multiple video segments and the features of each video segment in multiple dimensions are input into the initial evaluation model to obtain the estimated viewing experience information for multiple visual parameters of the last frame in each video segment. Based on the estimated viewing experience information and the actual viewing experience information, the initial evaluation model is updated to obtain the evaluation model.
3. The video processing method as described in claim 1, characterized in that, The method for determining an adjustment strategy for the current frame based on viewing experience information from multiple visual parameters of the current frame includes: For each visual parameter, based on the viewing experience information of the current visual parameter, determine the adjustment value of the current visual parameter for the current frame; The adjustment values of all visual parameters are determined as the adjustment strategy for the current frame.
4. The video processing method as described in claim 1, characterized in that, The adjustment of multiple visual parameters of the current frame based on the adjustment strategy includes: According to the adjustment strategy, multiple visual parameters of the current frame are directly adjusted to obtain the adjusted current frame; Alternatively, the adjustment strategy and the video segment can be input into the adjustment model to obtain the adjusted current frame.
5. The video processing method according to any one of claims 1 to 4, characterized in that, The features of the multiple dimensions include the following features At least two of the following are included: content features, complexity features, aesthetic features, and encoding features of the video segment, wherein the content features include content information of the video segment, the complexity features include image complexity information of the video segment, the aesthetic features include information on at least one aesthetic element of the video segment, and the encoding features include encoding information of the video segment.
6. The video processing method according to any one of claims 1 to 4, characterized in that, The plurality of visual parameters includes at least two of the following visual parameters: brightness, contrast, saturation, and color temperature.
7. A video processing apparatus, characterized in that, include: The adjustment unit is configured to perform the following adjustments for each frame in the streaming video: The evaluation model is input with a video segment containing the current frame and features of the video segment in multiple dimensions to obtain viewing experience information for multiple visual parameters of the current frame. The video segment is a segment of the streaming video and the current frame is the last frame of the video segment. Based on viewing experience information from multiple visual parameters of the current frame, an adjustment strategy is determined for the current frame; Based on the adjustment strategy, multiple visual parameters of the current frame are adjusted respectively; The live streaming unit is configured to stream live based on the adjusted push video.
8. An electronic device, characterized in that, include: At least one processor; At least one memory that stores computer-executable instructions. The computer-executable instructions, when executed by the at least one processor, cause the at least one processor to perform the video processing method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by at least one processor, the at least one processor causes the at least one processor to perform the video processing method as described in any one of claims 1 to 6.
10. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the video processing method as described in any one of claims 1 to 6.