An audio and video processing method, device and equipment of streaming media

By constructing a random forest prediction model and a hidden effect algorithm to optimize the audio and video encoding process, the problems of low efficiency and poor stability in audio and video data processing are solved, achieving efficient and stable audio and video data transmission and improving user experience.

CN116827921BActive Publication Date: 2026-02-24CHINA MOBILE ONLINE SERVICES CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210277623.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-21
Publication Date
2026-02-24
Estimated Expiration
2042-03-21

AI Technical Summary

Technical Problem

Audio and video data processing is inefficient and unstable, especially when transmission demand is high and network congestion occurs, which can easily lead to stuttering and congestion, affecting user experience.

Method used

By constructing a random forest prediction model and combining it with an image quality assessment algorithm for the occlusion effect, the encoding module is automatically adjusted to optimize the encoding process of audio and video data. This allows for early prediction of quality change trends and timely adjustment of key parameters, thereby optimizing the encoding module to improve encoding quality and reduce bandwidth.

Benefits of technology

It improves audio and video encoding quality, reduces bandwidth usage at the same frame rate and resolution, enhances audio and video data processing efficiency and stability, reduces transmission costs, extends device battery life, and ensures the continuity and high quality of video services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116827921B_ABST
    Figure CN116827921B_ABST
Patent Text Reader

Abstract

The application discloses an audio and video processing method, device and equipment of streaming media, and belongs to the technical field of computers. The method mainly comprises the following steps: obtaining to-be-processed data frames in audio and video data of streaming media and target data of an encoding module; based on the target data, a random forest prediction model corresponding to the target data is constructed, and the random forest prediction model is used to determine a structure similarity prediction evaluation value of the to-be-processed data frames; based on the structure similarity prediction evaluation value, an image quality evaluation algorithm of a concealment effect is used to determine a target structure similarity prediction evaluation value of the audio and video data; and in the case that the target structure similarity prediction evaluation value meets a preset condition, the random forest prediction model is adjusted to obtain a target encoding module, so that the to-be-processed data frames are encoded through the target encoding module, and the problems of low processing efficiency and poor stability of the audio and video data can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer technology, and specifically relates to a method, apparatus and device for audio and video processing of streaming media. Background Technology

[0002] With the adoption of Orthogonal Frequency Division Multiple Access (OFDMA) technology in wireless communication networks, the network bandwidth of 4G and 5G technologies has been significantly enhanced, expanding the transmission capacity and service scope of multimedia value-added services such as audio, video, and animation. To meet the needs of people's daily work and life, audio and video services have diversified in form, with short videos, online conferencing, remote training, and video surveillance playing important roles in various application areas.

[0003] However, in related technologies, the size or dimensions of the images generated by the audio and video streaming modules vary, and the transmission bandwidth, latency, and real-time network conditions are also different, which affects the user's perception of service experience. Especially under conditions of high transmission demand and high network congestion, audio and video stuttering and congestion occur frequently, reducing the efficiency of audio and video streaming processing. Summary of the Invention

[0004] The purpose of this application is to provide a method, apparatus, and device for audio and video processing of streaming media, which can solve the problems of low efficiency and poor stability in audio and video data processing.

[0005] In a first aspect, embodiments of this application provide a method for audio and video processing of streaming media, characterized in that it includes:

[0006] The module acquires the data frames to be processed from the audio and video data of the streaming media and the target data of the encoding module. The encoding module is the module that encodes the audio and video data. The target data is the data required by the encoding module to encode historical data frames. The historical data frames are the data frames that have been encoded in the audio and video data.

[0007] Based on the target data, a random forest prediction model corresponding to the target data is constructed. The random forest prediction model is used to determine the structural similarity prediction evaluation value of the data frame to be processed.

[0008] Based on the structural similarity prediction evaluation value, the target structural similarity prediction evaluation value of audio and video data is determined by an image quality assessment algorithm with concealment effect.

[0009] If the target structure similarity prediction evaluation value meets the preset conditions, the random forest prediction model is adjusted to obtain the target encoding module, which is then used to encode the data frame to be processed.

[0010] Secondly, embodiments of this application provide an audio and video processing apparatus for streaming media, characterized in that it includes:

[0011] The acquisition module is used to acquire the data frames to be processed in the audio and video data of the streaming media and the target data of the encoding module. The encoding module is the module that encodes the audio and video data. The target data is the data required by the encoding module to encode historical data frames. The historical data frames are the data frames that have been encoded in the audio and video data.

[0012] The building module is used to construct a random forest prediction model corresponding to the target data based on the target data. The random forest prediction model is used to determine the structural similarity prediction evaluation value of the data frame to be processed.

[0013] The determination module is used to determine the target structural similarity prediction evaluation value of audio and video data based on the structural similarity prediction evaluation value and through the image quality evaluation algorithm of the concealment effect.

[0014] The adjustment module is used to adjust the random forest prediction model when the target structure similarity prediction evaluation value meets the preset conditions, so as to obtain the target encoding module, and encode the data frame to be processed through the target encoding module.

[0015] Thirdly, embodiments of this application provide a computer device, which includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor. When the program or instructions are executed by the processor, they implement the steps of the streaming media audio and video processing method as described in the first aspect.

[0016] Fourthly, embodiments of this application provide a readable storage medium storing a program or instructions, which, when executed by a processor, implement the steps of the streaming media audio and video processing method as described in the first aspect.

[0017] Fifthly, embodiments of this application provide a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the steps of the streaming media audio and video processing method as shown in the first aspect.

[0018] In this embodiment, the process involves acquiring the data frames to be processed from the audio and video data of the streaming media, as well as the target data of the encoding module. The encoding module is used to encode the audio and video data, and the target data is the data required for encoding historical data frames in the encoding module. The historical data frames are the encoded data frames in the audio and video data. Based on the target data, a random forest prediction model corresponding to the target data is constructed. The random forest prediction model is used to determine the structural similarity prediction evaluation value of the data frames to be processed. Based on the structural similarity prediction evaluation value, an image quality assessment algorithm with camouflage effect is used to determine the target structural similarity prediction evaluation value of the audio and video data. When the target structural similarity prediction evaluation value meets preset conditions, the random forest prediction model is adjusted to obtain the target encoding module, which is then used to encode the data frames to be processed. Therefore, key parameters affecting audio and video quality, such as bitrate, quadtree partitioning depth, frame data coding blocks, quantization parameters, and rate distortion values, can be extracted to establish a random forest prediction model. This model can adaptively adjust the weight values ​​in the random forest prediction model based on actual network conditions to achieve the quality assessment process of Structure Similarity in Simulation (SSIM). It can also automatically determine the structural differences between the model and the real image. If significant distortion exists, the target coding module for encoding audio and video data is readjusted, and then the data frame to be processed is encoded through the target coding module. This allows for the prediction of audio and video quality trends before encoding the audio and video data, and the adjustment of the coding module to restore quality when audio and video data is distorted and key parameters need to be adjusted in time. In this way, while improving the quality of audio and video encoding, the bandwidth occupied by streaming media at the same frame rate and resolution is significantly reduced, thus solving the problems of low efficiency and poor stability in audio and video data processing. Attached Figure Description

[0019] Figure 1 A schematic diagram of an audio and video processing architecture for streaming media provided in an embodiment of this application;

[0020] Figure 2 This application provides a schematic diagram illustrating the relationship between target data and audio / video encoding results in an embodiment of the present application.

[0021] Figure 3 A schematic diagram illustrating a monitoring audio and video data stream encoding method provided in an embodiment of this application;

[0022] Figure 4 A flowchart illustrating an audio and video processing method for streaming media provided in an embodiment of this application;

[0023] Figure 5 A schematic diagram of quadtree partitioning based on the concealment effect is provided for an embodiment of this application;

[0024] Figure 6A schematic diagram of the structure of an audio and video processing device for streaming media provided in an embodiment of this application;

[0025] Figure 7 A schematic diagram of the structure of an audio and video processing device for streaming media provided in an embodiment of this application;

[0026] Figure 8 This is a schematic diagram of the hardware structure of an audio and video processing device for streaming media provided in an embodiment of this application. Detailed Implementation

[0027] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0028] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0029] In related technologies, to address the issues of low efficiency and poor stability in audio and video data processing, efficient audio and video data stream encoding and compression (hereinafter referred to as encoding) methods can be considered. This involves compressing the original codeword stream to the minimum bitrate without altering image clarity. Here, the main method of audio and video data stream encoding is to change the density between codewords, compressing redundant information in both spatial and temporal dimensions to the maximum extent, and removing related errors and interference noise data to improve data quality, thereby indirectly reducing the data stream bandwidth and bitrate. However, in the selection of audio and video data stream encoding methods, distortion must be considered. Simply removing data redundancy, while increasing the bitrate, can lead to excessively high speeds, potentially resulting in the loss of valid data and affecting the integrity of the data stream.

[0030] Furthermore, to improve the efficiency and stability of audio and video data processing, this application embodiment considers adjusting the audio and video data stream encoding method by monitoring the encoding quality of the audio and video data stream. Currently, the encoding quality of the audio and video data stream can be determined in the following two ways.

[0031] Firstly, there's the method of judging the encoding quality of audio and video data streams subjectively. This method relies on a person's subjective judgment of the audio and video data images or sounds displayed at the receiving end. The judgment criteria are determined by individual subjective assumptions. Therefore, due to the lack of a unified standard and the differences in individual sensory abilities, the final evaluation conclusion will vary, and the subjective judgment will differ from the actual result. Furthermore, the subjective judgment process is time-consuming, making this method unsuitable when the demand for quality assessment is high.

[0032] Secondly, there are methods for judging the encoding quality of audio and video data streams through objective means. These methods can be divided into three types: full-reference, half-reference, and no-reference. Full-reference and half-reference methods involve receiving all or part of the original audio and video key data at the encoding / decoding end (electronic device or server) as a basis and comparing it with the decoded audio and video. When the image distortion exceeds a standard threshold, the audio and video data stream encoding is considered invalid. No-reference methods monitor and analyze key parameters affecting video quality during the audio and video data stream encoding process. When certain key parameters are abnormal, it indicates low encoding efficiency and poor image quality. However, full-reference and half-reference methods require consuming transmission network resources to retransmit the original audio and video data stream. For large audio and video data streams, this can lead to a waste of transmission resources, and waiting for the original video to be transmitted increases the quality assessment time. Generally, no-reference judgment methods do not require the original video data stream as a reference and do not occupy space resources. However, on the one hand, no-reference judgment methods can classify, train, and evaluate compressed image data, identify and extract blurry areas in the image, and give qualitative analysis conclusions. However, if methods such as pixel value probability distribution fitting or peak signal-to-noise ratio maximum likelihood estimation are used to comprehensively judge image quality, all encoding and decoding are required before comparison and analysis, which is time-consuming and the test results differ significantly from the actual human visual perception results. Therefore, this method lacks universality. On the other hand, key parameters can be predicted and evaluated based on the coding quality assessment of the bitstream compression. However, the prediction results of different compression coding methods have large deviations, and the evaluation strategy is often only applicable to a certain compression method, resulting in poor universality. Therefore, the final evaluation results may also be inconsistent with reality.

[0033] To effectively address the problems arising from the aforementioned methods, this application provides a streaming media audio and video processing method. This method can provide an automatic monitoring link for the audio and video encoding process of High Efficiency Video Coding (HEVC). The monitoring link can incorporate key parameters such as bitrate and quantization parameters. The training samples for the random forest model in the training monitoring link are the parameters (QP), rate-distortion value, and quadtree partitioning depth. A regression tree is constructed based on the training samples. The structural similarity prediction evaluation value of each leaf node is calculated through the regression tree, which is the predicted value. The structural similarity prediction evaluation value is weighted and adjusted with the intervention of the concealment effect. Finally, the target structural similarity prediction evaluation value of the audio and video data is obtained, which is the structural similarity (SSIM) of the final compressed audio and video stream image. This realizes the quality assessment process of SSIM and automatically judges the structural differences between the actual image and the real image. If there is a large distortion, the target encoding module for encoding the audio and video data is readjusted, and then the data frame to be processed is encoded through the target encoding module. This allows the change trend of audio and video quality to be predicted in advance before encoding the audio and video data. When the audio and video data is distorted and key parameters need to be adjusted in time, the encoding module is adjusted to restore the quality. In this way, while improving the audio and video encoding quality, the bandwidth occupied by the streaming media at the same frame rate and resolution is significantly reduced, thus solving the problems of low efficiency and poor stability of audio and video data processing.

[0034] Furthermore, by achieving service standard requirements in terms of transmission efficiency and video image quality, and by adding steps to monitor audio and video data stream encoding and determine encoding effect evaluation, the encoding quality is ensured to be optimized. Here, utilizing the audio and video data stream encoding step can effectively improve the processing efficiency of audio and video stream data. It can also modify the encoding method or parameters according to real-time business transmission needs when large-concurrency data stream transmission occurs, increasing the encoding rate, maximizing the utilization of wireless channel resources, reducing transmission costs, and ensuring uninterrupted and smooth audio and video data display. Simultaneously, monitoring the audio and video data stream encoding quality can accurately locate audio and video stream transmission faults, rationally plan and optimize idle transmission resources, and provide users with high-quality video service. Due to the adoption of the audio and video data stream encoding process, the encoding efficiency is improved, reducing the power consumption of encoding processing and related equipment, extending equipment battery life, and enhancing the duration of service to users, thus improving service quality. Based on this, the streaming media audio and video processing method provided in this application embodiment can be applied to the application environment of large-scale 5G network deployment, and can solve problems such as poor audio and video data application and service quality and poor stability in this scenario.

[0035] Based on this, the following is in conjunction with the appendix Figures 1-2 The audio and video processing method for streaming media provided in this application will be described in detail through specific embodiments and application scenarios.

[0036] This application proposes an audio and video processing architecture for streaming media, such as... Figure 1 As shown, the audio and video processing architecture of this streaming media may include a computer device 10. In one example, the computer device 10 may include a partitioning module, a prediction module, a transformation module, a quantization module, and an entropy coding module.

[0037] The following is combined Figure 1 This application provides a detailed description of the audio and video processing method for streaming media, as detailed below. Here, the audio and video processing method for streaming media provided in this application includes an audio and video data stream encoding process and a process for monitoring the audio and video data stream encoding.

[0038] First, combined Figure 1 The audio and video data stream encoding process is explained.

[0039] like Figure 1 As shown, for the audio and video data encoding process, this application embodiment adopts a new high-efficiency video compression standard, namely HEVC, to replace the original H.264 / AVC encoding standard. After the data frames in the original audio and video data are processed through the audio and video data encoding process, they are compressed to the target bitstream security state that matches the transmission channel. This fully saves channel resources to transmit the maximum saturation amount of audio and video data and achieves the best efficiency transmission. Based on this, the specific audio and video data stream encoding process is as follows.

[0040] The partitioning module acquires data frame 1 (such as audio frames or image frames) from the audio and video data and partitions it using a quadtree, specifically by dividing data frame 1 into tree units (CTUs) of equal size using HEVC. Then, the CTUs are continuously partitioned into coding units (CUs) and prediction units (PUs). It should be noted that the depth of the quadtree's tree structure directly affects the accuracy and efficiency of the entire audio and video data encoding process.

[0041] Here, this application embodiment presents a quality assessment method based on rate-distortion value and data coding complexity. Specifically, it uses HEVC and the SSIM evaluation method to analyze degradation factors between different pixels, incorporates a visual concealment effect discrimination method to improve audio and video service quality, extracts target data, constructs a random forest prediction model corresponding to the target data, and utilizes a random decision tree algorithm to obtain a more accurate quality assessment method, thus improving the accuracy of evaluating the coding quality of audio and video data streams. It should be noted that the SSIM evaluation method in this application embodiment improves accuracy compared to the Peak Signal-to-Noise Ratio (PSNR) evaluation method, which is easily affected by external environmental interference. The SSIM method in this application embodiment avoids the problem of discrepancies between the evaluation conclusions generated by the PSNR evaluation method and actual subjective judgment, making it unaffected by external brightness, contrast, etc. It calculates the similarity between two frames, and the calculated evaluation conclusion is basically consistent with actual subjective judgment, improving the automatic quality assessment effect and enhancing accuracy and practicality.

[0042] The prediction module is used to predict the depth of the quadtree tree structure partitioning result (e.g., the partitioning depth of each sub-unit output by the partitioning module, i.e., the coding unit (CU) and / or the prediction unit (PU). Further, the prediction relies on the entropy-encoded output value of the previous data frame (e.g., data frame 0) as a reference frame. It predicts the audio / video data quality loss value using the motion residual between the current frame (e.g., data frame 1) and the reference frame (e.g., data frame 0). This loss value is used to adjust the encoding prediction process of the current frame. Further, the prediction is divided into inter-frame prediction and intra-frame prediction. Inter-frame prediction involves calculating the residual between the quadtree tree structure partitioning depth of the current frame and the quadtree tree structure partitioning depth of encoded historical frames. Intra-frame prediction involves calculating the residual between all predicted sub-units within the quadtree tree structure partitioning depth of the current frame and all preset sub-units within the quadtree tree structure partitioning depth of the reference frame. It should be noted that the deeper the sub-unit's quadtree tree structure partitioning depth and the smaller the partitioned unit, the higher the audio / video quality. Therefore, this step directly affects the clarity of the encoded audio / video.

[0043] The transformation module is used to acquire the partitioning results of each sub-unit (i.e., coding unit CU and / or prediction unit PU) output by the partitioning module, and the prediction results of each sub-unit (i.e., coding unit CU and / or prediction unit PU) output by the prediction module. It then performs Discrete Cosine Transform (DCT) or Discrete Sine Transform (DST) on the partitioning and prediction results. The main purpose of this transformation is to perform a Fourier transform on the data stream, ensuring that the audio and video data stream format conforms to the network format requirements for subsequent quantization and entropy coding. This transforms the audio and video data stream from the spatial domain to the image transform domain, reducing spatial redundancy and improving coding efficiency and quality.

[0044] A quantization module (and / or reordering module) is used to convert continuous Fourier transform values ​​into discrete values, and to set the quantization coefficients of high-frequency signals such as noise to zero to enhance the compression coding effect. In the embodiments of this application, the quantization process refers to formula (1):

[0045]

[0046] Where xi is the quantized value output by the quantization module, and ai are the Fourier transform coefficients. QP is the quantization parameter, and α is the quantization parameter. i ε is the output value of the transformation module, ε is the Gaussian parameter (here, an integer is taken), and floor() represents the floor function. It should be noted that, as shown in formula (1), the larger QP is, the smaller the quantization parameter, the worse the denoising effect, and the worse the final audio and video data quality. It can be seen that QP directly affects the quality of compressed video; the smaller the QP, the better the quantization distortion quality.

[0047] The entropy coding module describes the minimum bit count requirement for encoding and compression of audio and video data without data loss. During data compression, the entropy of the message is minimized based on a probabilistic model of the source message, and the audio and video data is restored using this minimum entropy. It should be noted that the encoding process is related to the depth and size of the quadtree structure of the audio and video data.

[0048] Therefore, as Figure 2 As shown, factors affecting audio and video coding efficiency and quality can include bitrate, QP value, rate-distortion parameter, structural similarity prediction evaluation value (SSIM evaluation value), and quadtree tree structure such as the partitioning depth of the quadtree. The following section uses... Figure 2 Taking the target data as an example, we analyze the factors affecting audio and video distortion and degradation, so as to provide a detailed explanation of the process of encoding the monitoring audio and video data stream in the embodiments of this application.

[0049] (1) The influence of code rate data: The code rate control parameter controls the coding rate based on the actual channel transmission quality. It allocates and limits the number of coded bits based on bits. The channel rate value directly determines the average number of bits allocated. The higher the channel rate, the larger the real-time allocated bit block, and the code rate automatically increases.

[0050] (2) The impact of QP value data: The impact of QP value has been described earlier. The larger the QP, the worse the video quality.

[0051] (3) The influence of rate-distortion parameters, i.e. rate-distortion optimization, is as follows: In the HEVC encoding process, sometimes excessive compression can lead to video quality distortion. A rate-distortion optimization model is introduced to control the compression. Specifically, it can be based on formula (2):

[0052] J = D + βW·2 (QP-12) / 3 (2)

[0053] Where J is the encoding cost, D is the video distortion, and β and W are the calculation model parameters of the bitrate weight coefficient, such as the parameters of the initial random forest prediction model, which are determined by the encoder. The larger the QP value, the greater the encoding cost, the higher the distortion rate, and the worse the video encoding quality.

[0054] (4) The influence of the quadtree tree structure, i.e., the similarity evaluation index of the quadtree tree structure. Based on the above, this embodiment uses the SSIM quality evaluation method to evaluate the encoding quality of audio and video data by assessing the similarity of image pixels in terms of structure, brightness, and contrast. Here, the SSIM evaluation value ranges from 0 to 1. The larger the value, the less the loss. When there is no distortion, the pixels are completely restored, and the SSIM evaluation value is 1. It should be noted that the SSIM quality evaluation method has a very similar ability to the human visual system. Therefore, the SSIM quality evaluation method can reflect both subjective sensory information and accurately describe objective structural similarity through calculation.

[0055] Since the above-mentioned factors all affect the audio and video data stream encoding process and are deeply interconnected, we will still refer to... Figure 2 As shown, the SSIM evaluation value is directly related to the bitrate and QP quantization value. The higher the bitrate, the larger the data volume, and the better the quality of the audio and video images. Within a certain bitrate range, it can be directly proportional to the SSIM evaluation value. Different audio and video streams have different complexities, and the resulting curve relationships are not entirely consistent. The QP quantization parameter is inversely proportional to the SSIM evaluation value. The larger the quantization parameter, the more severe the video distortion and the lower the structural similarity.

[0056] Based on this, the embodiments of this application provide a monitoring audio and video data stream encoding process to address issues affecting video quality, in order to predict the changing trends of audio and video quality in advance. When the SSIM evaluation value falls below a certain specified threshold, it indicates that the audio and video data has been distorted, and key parameters need to be adjusted in a timely manner to restore its quality. The specific monitoring audio and video data stream encoding process can be combined with... Figure 3 A detailed explanation is provided. Additionally, it should be noted that the module executing the monitoring audio and video data stream encoding process can be located within the partitioning module and / or the prediction module, or it can be located between the bitrate impact and partitioning modules.

[0057] Based on this, such as Figure 3 As shown, audio and video data are acquired, and combined with... Figure 2 The target data shown can affect video coding quality. This target data may include bitrate, quadtree partitioning depth, frame data coding blocks, quantization parameters, and rate-distortion values. The target data is used to construct a random forest prediction model corresponding to the target data. This model can adaptively adjust the weights in the random forest prediction model based on the actual network conditions to jointly complete the SSIM quality assessment process. It automatically judges the structural differences between the model and the real image. If significant distortion exists, the model needs to be readjusted. Figure 1 The coding coefficients in each module of the monitoring link shown are adjusted to the weights in the random forest prediction model to avoid continuous distortion.

[0058] Therefore, this application provides an automated monitoring process for the audio and video encoding process of the HEVC protocol. It combines extracted target data and uses it as a reference value for weight calculation in the training of a random forest prediction model to construct a regression tree. The structural similarity prediction evaluation value for each leaf node is calculated, and the structural similarity prediction evaluation value is weighted and adjusted with the intervention of the concealment effect. Then, the target structural similarity prediction evaluation value of the audio and video data is determined as the SSIM of the final compressed audio and video stream image. Furthermore, to improve the quality of SSIM evaluation, a video concealment effect method is integrated into the intra-frame and inter-frame prediction processes. This maximizes the partitioning depth of the CU in each CTU unit, uses the number of CU blocks in different frames to determine the data complexity, optimizes the evaluation parameters of the SSIM evaluation process, and effectively improves the intra-frame and inter-frame prediction performance. The specific monitoring process can be combined with... Figure 4 and Figure 6 The encoding content of the monitoring audio and video data stream in the embodiments of this application will be described in detail.

[0059] It should be noted that the target module for encoding the monitoring audio and video data stream can be set in the partitioning module and / or prediction module, and the following should be executed: Figure 4 The audio and video processing method for streaming media shown in this application is combined with... Figures 4-6 The audio and video processing method for streaming media provided in the embodiments of this application will be described in detail.

[0060] Figure 4 This is a flowchart illustrating an audio and video processing method for streaming media provided in an embodiment of this application.

[0061] like Figure 4 As shown, this streaming media audio and video processing method can be applied to the above-mentioned... Figure 1 The audio and video processing architecture for streaming media shown can specifically include the following steps:

[0062] Step 410: Obtain the data frames to be processed and the target data of the encoding module from the audio and video data of the streaming media. The encoding module is the module that encodes the audio and video data, and the target data is the data required for encoding historical data frames in the encoding module. Historical data frames are the data frames that have already been encoded in the audio and video data. Step 420: Based on the target data, construct a random forest prediction model corresponding to the target data. The random forest prediction model is used to determine the structural similarity prediction evaluation value of the data frames to be processed. Step 430: Based on the structural similarity prediction evaluation value, determine the target structural similarity prediction evaluation value of the audio and video data using an image quality assessment algorithm with veil effect. Step 440: If the target structural similarity prediction evaluation value meets preset conditions, adjust the random forest prediction model to obtain the target encoding module, which is then used to encode the data frames to be processed.

[0063] Therefore, key parameters affecting audio and video quality, such as bitrate, quadtree partitioning depth, frame data coding blocks, quantization parameters, and rate distortion values, can be extracted to establish a random forest prediction model. This model can adaptively adjust the weight values ​​in the random forest prediction model based on actual network conditions to achieve a quality assessment process for similar structures. It can also automatically determine the structural differences between the model and the real image. If significant distortion exists, the target coding module for encoding audio and video data is readjusted, and then the data frames to be processed are encoded through the target coding module. This allows for the prediction of audio and video quality trends before encoding the audio and video data, and the adjustment of the coding module to restore quality when audio and video data is distorted and key parameters need to be adjusted in time. Thus, while improving audio and video coding quality, the bandwidth occupied by streaming media at the same frame rate and resolution is significantly reduced, solving the problems of low efficiency and poor stability in audio and video data processing.

[0064] The above steps are explained in detail below:

[0065] In step 410, in one possible embodiment, the target data in this application embodiment includes at least one of the following: bit rate, quadtree partitioning depth, frame data coding block, quantization parameter, and rate distortion value.

[0066] Here, the bitrate, QP value, rate-distortion value, and SSIM evaluation value provided in the embodiments of this application are used to describe the structured related parameters of audio and video evaluation quality, and to guide the quality evaluation process of HEVC protocol encoding by utilizing the changes and coupling relationships between the data.

[0067] Regarding step 420, in one possible embodiment, step 420 may specifically include:

[0068] Step 4201: In the case that the random forest prediction model includes a prediction regression tree, and the leaf nodes in the prediction regression tree are used to determine the structural similarity prediction evaluation value of the data frame to be processed, the training samples are input into the initial random forest prediction model, and the encoder randomly selects the training sample set in the audio and video dataset. The training samples include audio and video data and target data.

[0069] Step 4202: Based on the training sample set, calculate the key feature set corresponding to the training sample set;

[0070] Step 3203: Based on the key feature set, construct a regression tree, and sort the key features in the key feature set according to the preset feature priority information to obtain the sorting result;

[0071] Step 4204: Based on the ranking results, the regression tree is divided using the decision tree features with the minimum mean square error to obtain the predicted regression tree.

[0072] For example, the CU block corresponding to the original audio and video data is used as input. The encoder randomly selects a set of data blocks and extracts a set of key features from the image data to construct a regression tree. The key features, such as image brightness and contrast, are prioritized and sorted. The regression tree is divided based on the decision tree features with the minimum mean square error.

[0073] Therefore, the extracted target data is used as the reference value for weight calculation in the training of the random forest model to construct a regression tree so as to calculate the predicted value of each leaf node feature.

[0074] Furthermore, step 4203 may specifically include:

[0075] In the initial random forest prediction model, a quadtree partitioning module is included. The quadtree in the quadtree partitioning module includes four sub-coding units at the same position and the parent coding unit corresponding to the four sub-coding units. In the case that the key feature set includes the rate-distortion values ​​of the sub-coding units, the sum of the second rate-distortion values ​​of each of the four sub-coding units is determined as the third rate-distortion value of the four sub-coding units.

[0076] By comparing the third rate distortion value with the fourth rate distortion value of the parent coding unit, a second comparison result is obtained;

[0077] If the second comparison result indicates that the third rate distortion value is less than or equal to the fourth rate distortion value, the quadtree partitioning depth of the quadtree in the quadtree partitioning module is increased to obtain the regression tree.

[0078] In order to improve the quality of monitoring and evaluation, this application introduces the concealment effect into the encoding process of monitoring audio and video data streams. Here, the concealment effect in this application refers to the fact that when multiple core stimuli act on the human eye at the same time, the human eye will shield the loss effect of multiple stimuli and reduce the distortion rate. This principle can be integrated into the quality evaluation process to improve the recognition and measurement of time and space complexity.

[0079] It should be noted that, as Figure 5 As shown, step 4203 can incorporate the concealment effect into intra-frame prediction. That is, during the quadtree partitioning process, each CTU continuously splits downwards into four CUs of the same size based on spatial depth. The traditional method for setting the CU partitioning depth is to perform traversal rate-distortion value calculation for each parent CU and determine whether to continue classifying sub-CUs by comparing subjective and objective empirical thresholds. Generally, the finer the CU partitioning depth, the higher the audio and video coding quality. Since the four sub-CUs are in the same position, the concealment effect is introduced. The rate-distortion values ​​of the four sub-CUs are calculated separately, then added together and compared with the rate-distortion value of the parent CU. When it is less than the rate-distortion value of the parent CU, it means that the current picture quality meets the user's needs, and CU partitioning can continue; otherwise, when it is greater than the rate-distortion value of the parent CU, partitioning needs to be stopped to avoid causing distortion. Then, the distortion rate values ​​of the four sub-CUs can be calculated, which is equivalent to applying the same stimulus to the human eye. During the joint calculation of the distortion rate value, some low-quality points will be partially masked. This can maximize the division depth of the CU, which is beneficial to improve the prediction effect of intra-frame distortion. At the same time, the bit rate is adjusted in real time, which adapts to the bandwidth and maximizes the improvement of image quality.

[0080] Therefore, the method provided in this application embodiment may include incorporating the concealment effect into intra-frame prediction. That is, for each coding unit (CTU), it is divided into four CU units at equal positions using a four-part tree method. The concealment effect is introduced, and the rate-distortion values ​​of the four sub-CUs are calculated respectively. Then, they are added together and compared with the rate-distortion value of the parent CU. If the sum is less than the rate-distortion value of the parent CU, the division of CUs continues; otherwise, if the sum is greater than the rate-distortion value of the parent CU, the division needs to be stopped to avoid causing distortion.

[0081] Before step 430, in one possible embodiment, the audio and video processing method for streaming media in this application embodiment may further include:

[0082] Based on each leaf node in the predictive regression tree, the structural similarity prediction evaluation value corresponding to each leaf node is calculated in a round-robin fashion.

[0083] For example, the structural similarity prediction evaluation value of each leaf node is calculated in a round-robin fashion, and the average value of the structural similarity prediction evaluation values ​​of the child nodes is finally calculated to obtain the output of the random forest prediction model. The output result is the target structural similarity prediction evaluation value of the audio and video data. When the evaluated feature indicators exceed the distortion threshold, it indicates that the weights need to be further adjusted, and the recoding process is completed.

[0084] Regarding step 430, in one possible embodiment, step 430 may specifically include:

[0085] Step 4301: For the structural similarity prediction evaluation value corresponding to each leaf node, Gaussian weighting is used under the intervention of the concealment effect to calculate the target value for each polling process. The target value includes the mean, variance and covariance.

[0086] Step 4302: The average of multiple target values ​​is determined as the target structure similarity prediction evaluation value of the audio and video data.

[0087] Furthermore, step 4301 may specifically include:

[0088] In the case where the random forest prediction model includes a quadtree, and the quadtree includes at least two adjacent sub-coding units on the time axis, the first rate-distortion value of each sub-coding unit in the at least two adjacent sub-coding units is calculated, and the average rate-distortion value of the at least two adjacent sub-coding units is calculated based on the first rate-distortion value of each sub-coding unit.

[0089] The first comparison result is obtained by comparing the average ratio distortion value with the preset ratio distortion value;

[0090] If the average rate distortion of the first comparison result is less than or equal to the preset rate distortion value, increase the number of adjacent sub-coding units on the time axis;

[0091] The target value for each polling cycle is calculated by adding a quadtree to the sub-coding unit.

[0092] For example, still refer to Figure 5The concealment effect is incorporated into inter-frame prediction. Audio and video data can be divided into a flowing set of images, and each adjacent coding unit (CU) in the image also changes over time. Therefore, the data capacity of inter-frame prediction can be understood as its complexity in the time domain. To minimize the rate-distortion cost, a balance between distortion rate and coding rate needs to be calculated. By introducing the concealment effect, several adjacent CUs on the time axis move at almost a constant bit rate. Their distortion rates are calculated and averaged, then compared with the optimal distortion rate. If the average is less than the optimal distortion rate, it means that the number of adjacent CU blocks can be increased; if it is greater, the original number of CU blocks is retained.

[0093] It should be noted that the number of CU blocks is similar to a dynamic sliding window, which can determine the time complexity of inter-frame prediction in audio and video coding. During the SSIM evaluation process, the target values ​​(mean, variance, and covariance) can be adjusted appropriately based on the complexity values, and the real-time bitrate can be adjusted accordingly.

[0094] Therefore, the above-mentioned intra-frame and inter-frame prediction concealment effect algorithms can be applied in the modeling, weight adjustment, and SSIM quality assessment processes, enhancing the efficiency and quality of encoding processing. Using the target data mentioned above as reference values ​​for weight calculation in the training of the random forest prediction model, a regression tree is constructed. The structural similarity prediction evaluation value of each leaf node feature is calculated. With the intervention of the concealment effect, the mean, variance, and covariance of the predicted S-structure similarity prediction evaluation value are calculated using Gaussian weighting for each iteration. Finally, the target structural similarity prediction evaluation value is used as the SSIM of the audio and video, i.e., the average structural similarity, to complete the output.

[0095] Regarding step 440, in one possible embodiment, step 440 may specifically include:

[0096] If the target structure similarity prediction evaluation value is greater than or equal to the preset threshold, adjust the weight values ​​of the initial random forest model to obtain the target random forest model;

[0097] Based on the objective random forest model, an objective encoding module is generated.

[0098] Here, SSIM evaluates the coding quality of audio and video data by comparing the similarity of image structure, brightness, and contrast metrics. This ensures that the computational structure is largely consistent with the actual results. The audio and video stream data being evaluated is input into the model, which, combined with key parameters affecting video quality, extracts target data, establishes a random forest prediction model, and adaptively adjusts the weights according to the actual network conditions to complete the SSIM quality evaluation process. Furthermore, it automatically judges the structural differences between the model and the real image. If there is significant distortion, it readjusts the system coding coefficients to avoid continued distortion.

[0099] In summary, the method provided in this application embodiment has a superior automatic quality assessment effect, with higher accuracy and practicality. Specifically, the module executing the monitoring audio and video data stream encoding process can predict the changing trend of audio and video quality in advance. When the S-target structural similarity prediction evaluation value is lower than a certain specified threshold, it indicates that the audio and video data has been distorted, and key parameters need to be adjusted in time to restore its quality. In addition, in order to improve the SSIM evaluation quality, this application embodiment integrates the audio and video concealment effect algorithm into the intra-frame and inter-frame prediction process, which can maximize the partitioning depth of CU in each CTU. The number of CUs in different frames determines the data complexity, optimizes the SSIM evaluation process, and effectively improves the intra-frame and inter-frame prediction effect. This results in a significant reduction in the bandwidth occupied by the media stream at the same frame rate and resolution while improving the audio and video encoding quality.

[0100] It should be noted that the streaming media audio and video processing method provided in this application embodiment can be executed by a streaming media audio and video processing device, or a control module in the streaming media audio and video processing device for executing the streaming media audio and video processing method. This application embodiment uses the execution of the streaming media audio and video processing method by a streaming media audio and video processing device as an example to illustrate the streaming media audio and video processing device provided in this application embodiment.

[0101] Based on the same inventive concept, this application also provides an audio and video processing device for streaming media. (Specifically combined with...) Figure 6 Please provide a detailed explanation.

[0102] Figure 6 This is a schematic diagram of the structure of an audio and video processing device for streaming media provided in an embodiment of this application.

[0103] like Figure 6 As shown, the audio and video processing device 60 for streaming media is applied to audio and video processing equipment for streaming media, and may specifically include:

[0104] The acquisition module 601 is used to acquire the data frames to be processed in the audio and video data of the streaming media and the target data of the encoding module. The encoding module is a module for encoding the audio and video data, and the target data is the data required for encoding historical data frames in the encoding module. The historical data frames are the data frames that have been encoded in the audio and video data.

[0105] Module 602 is used to construct a random forest prediction model corresponding to the target data based on the target data. The random forest prediction model is used to determine the structural similarity prediction evaluation value of the data frame to be processed.

[0106] The determination module 603 is used to determine the target structural similarity prediction evaluation value of audio and video data based on the structural similarity prediction evaluation value and through the image quality evaluation algorithm of the concealment effect.

[0107] The adjustment module 604 is used to adjust the random forest prediction model when the target structure similarity prediction evaluation value meets the preset conditions, so as to obtain the target encoding module and encode the data frame to be processed through the target encoding module.

[0108] Therefore, key parameters affecting audio and video quality, such as bitrate, quadtree partitioning depth, frame data coding blocks, quantization parameters, and rate distortion values, can be extracted to establish a random forest prediction model. This model can adaptively adjust the weight values ​​in the random forest prediction model based on actual network conditions to achieve the quality assessment process of Structure Similarity in Simulation (SSIM). It can also automatically determine the structural differences between the model and the real image. If significant distortion exists, the target coding module for encoding audio and video data is readjusted, and then the data frame to be processed is encoded through the target coding module. This allows for the prediction of audio and video quality trends before encoding the audio and video data, and the adjustment of the coding module to restore quality when audio and video data is distorted and key parameters need to be adjusted in time. In this way, while improving the quality of audio and video encoding, the bandwidth occupied by streaming media at the same frame rate and resolution is significantly reduced, thus solving the problems of low efficiency and poor stability in audio and video data processing.

[0109] The audio and video processing device 60 for this streaming media is described in detail below:

[0110] In one or more possible embodiments, the construction module 602 may be specifically used to input training samples into the initial random forest prediction model, and to randomly select a training sample set from the audio and video dataset by the encoder, wherein the training samples include audio and video data and target data, in the case where the random forest prediction model includes a prediction regression tree and the leaf nodes in the prediction regression tree are used to determine the structural similarity prediction evaluation value of the data frame to be processed.

[0111] Based on the training sample set, calculate the set of key features corresponding to the training sample set;

[0112] Based on the key feature set, a regression tree is constructed, and the key features in the key feature set are prioritized according to the preset feature priority information to obtain the ranking result;

[0113] Based on the ranking results, the regression tree is divided using the decision tree features with the minimum mean square error, and the predicted regression tree is obtained.

[0114] In another or more possible embodiments, the audio and video processing apparatus 60 for streaming media may further include a calculation module for polling and calculating the structural similarity prediction evaluation value corresponding to each leaf node based on each leaf node in the prediction regression tree.

[0115] In one or more possible embodiments, the determining module 603 may be specifically used to calculate the target value for each polling process by applying Gaussian weighting to the structural similarity prediction evaluation value corresponding to each leaf node under the intervention of the concealment effect. The target value includes the mean, variance and covariance.

[0116] The average of multiple target values ​​is used as the target structural similarity prediction evaluation value for audio and video data.

[0117] In one or more possible embodiments, the determining module 603 may specifically be used to, in the case that the random forest prediction model includes a quadtree, and the quadtree includes at least two adjacent sub-coding units on the time axis, calculate the first rate-distortion value of each of the at least two adjacent sub-coding units, and calculate the average rate-distortion value of the at least two adjacent sub-coding units based on the first rate-distortion value of each sub-coding unit.

[0118] The first comparison result is obtained by comparing the average ratio distortion value with the preset ratio distortion value;

[0119] If the average rate distortion of the first comparison result is less than or equal to the preset rate distortion value, increase the number of adjacent sub-coding units on the time axis;

[0120] The target value for each polling cycle is calculated by adding a quadtree to the sub-coding unit.

[0121] In one or more possible embodiments, the construction module 602 may be specifically used to determine the sum of the second rate-distortion values ​​of each of the four sub-coding units as the third rate-distortion value of the four sub-coding units, in the case that the initial random forest prediction model includes a quadtree partitioning module, the quadtree in the quadtree partitioning module includes four sub-coding units at the same position and the parent coding unit corresponding to the four sub-coding units, and the key feature set includes the rate-distortion values ​​of the sub-coding units;

[0122] By comparing the third rate distortion value with the fourth rate distortion value of the parent coding unit, a second comparison result is obtained;

[0123] If the second comparison result indicates that the third rate distortion value is less than or equal to the fourth rate distortion value, the quadtree partitioning depth of the quadtree in the quadtree partitioning module is increased to obtain the regression tree.

[0124] In one or more possible embodiments, the adjustment module 604 may be specifically used to adjust the weight values ​​of the initial random forest model to obtain the target random forest model when the target structure similarity prediction evaluation value is greater than or equal to a preset threshold.

[0125] Based on the objective random forest model, an objective encoding module is generated.

[0126] In one or more possible embodiments, the target data includes at least one of the following: bit rate, quadtree partitioning depth, frame data coding block, quantization parameter, rate distortion value.

[0127] The audio and video processing device for streaming media in this application embodiment can be a device, or a component, integrated circuit, or chip in an electronic device. The device can be a mobile electronic device or a non-mobile electronic device. For example, mobile electronic devices can be mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs), etc., while non-mobile electronic devices can be servers, network-attached storage (NAS), personal computers (PCs), televisions (TVs), ATMs, or self-service machines, etc. This application embodiment does not impose specific limitations.

[0128] The audio and video processing device for streaming media in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.

[0129] The audio and video processing apparatus for streaming media provided in this application embodiment can achieve... Figures 1 to 6 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.

[0130] In this embodiment, the process involves acquiring the data frames to be processed from the audio and video data of the streaming media, as well as the target data of the encoding module. The encoding module is used to encode the audio and video data, and the target data is the data required for encoding historical data frames in the encoding module. The historical data frames are the encoded data frames in the audio and video data. Based on the target data, a random forest prediction model corresponding to the target data is constructed. The random forest prediction model is used to determine the structural similarity prediction evaluation value of the data frames to be processed. Based on the structural similarity prediction evaluation value, an image quality assessment algorithm with camouflage effect is used to determine the target structural similarity prediction evaluation value of the audio and video data. When the target structural similarity prediction evaluation value meets preset conditions, the random forest prediction model is adjusted to obtain the target encoding module, which is then used to encode the data frames to be processed. Therefore, key parameters affecting audio and video quality, such as bitrate, quadtree partitioning depth, frame data coding blocks, quantization parameters, and rate distortion values, can be extracted to establish a random forest prediction model. This model can adaptively adjust the weight values ​​in the random forest prediction model based on actual network conditions to achieve the quality assessment process of Structure Similarity in Simulation (SSIM). It can also automatically determine the structural differences between the model and the real image. If significant distortion exists, the target coding module for encoding audio and video data is readjusted, and then the data frame to be processed is encoded through the target coding module. This allows for the prediction of audio and video quality trends before encoding the audio and video data, and the adjustment of the coding module to restore quality when audio and video data is distorted and key parameters need to be adjusted in time. In this way, while improving the quality of audio and video encoding, the bandwidth occupied by streaming media at the same frame rate and resolution is significantly reduced, thus solving the problems of low efficiency and poor stability in audio and video data processing.

[0131] Optional, such as Figure 7 As shown, this application embodiment also provides a streaming media audio and video processing device 70, including a processor 701, a memory 702, and a program or instructions stored in the memory 702 and executable on the processor 701. When the program or instructions are executed by the processor 701, they implement the various processes of the above-described streaming media audio and video processing method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0132] Figure 8 This is a schematic diagram of the hardware structure of an audio and video processing device for streaming media provided in an embodiment of this application.

[0133] The audio and video processing device 800 for this streaming media includes, but is not limited to, components such as: radio frequency unit 801, network module 802, audio output unit 803, input unit 804, sensor 805, display unit 806, user input unit 807, interface unit 808, memory 809, processor 810, and microphone 88.

[0134] Those skilled in the art will understand that the streaming media audio and video processing device 800 may also include a power supply (such as a battery) for powering various components. The power supply may be logically connected to the processor 810 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 8 The audio and video processing device structure shown in the figure does not constitute a limitation on the audio and video processing device for streaming media. The audio and video processing device for streaming media may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0135] It should be understood that the input unit 804 may include a graphics processing unit (GPU) 8041 and a microphone 8042. The GPU 8041 processes image data of still images or videos acquired by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 806 may include a display panel 8061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 807 includes a touch panel 8071 and other input devices 8072. The touch panel 8071 is also called a touch screen. The touch panel 8071 may include a touch detection device and a touch controller. Other input devices 8072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, joysticks, etc., which will not be described in detail here. The memory 809 can be used to store software programs and various data, including but not limited to applications and operating systems. The processor 810 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understandable that the aforementioned modem processor may not be integrated into the processor 810.

[0136] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described audio and video processing method embodiments for streaming media and achieve the same technical effects. To avoid repetition, these will not be described again here.

[0137] The processor is the processor in the streaming media audio and video processing device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0138] In addition, this application embodiment provides another chip, which includes a processor and a communication interface. The communication interface and the processor are coupled. The processor is used to run programs or instructions to implement the various processes of the above-described audio and video processing method embodiment for streaming media, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0139] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0140] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0141] Furthermore, it should be noted that the scope of the methods and apparatus in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. In addition, features described with reference to certain examples may be combined in other examples.

[0142] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.

[0143] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A method for processing audio and video streams, characterized in that, include: The method acquires the data frames to be processed and the target data of the encoding module from the audio and video data of the streaming media. The encoding module is a module for encoding the audio and video data. The target data is the data required by the encoding module to encode historical data frames. The historical data frames are the data frames that have been encoded in the audio and video data. The target data includes at least one of the following: bitrate, quadtree partitioning depth, frame data coding block, quantization parameters, and rate distortion value. Based on the target data, a random forest prediction model corresponding to the target data is constructed. The random forest prediction model is used to determine the structural similarity prediction evaluation value of the data frame to be processed. Based on the structural similarity prediction evaluation value, the target structural similarity prediction evaluation value of the audio and video data is determined through an image quality evaluation algorithm with concealment effect. If the target structure similarity prediction evaluation value meets the preset conditions, the random forest prediction model is adjusted to obtain the target encoding module, so as to encode the data frame to be processed through the target encoding module.

2. The method according to claim 1, characterized in that, The random forest prediction model includes a prediction regression tree, and the leaf nodes in the prediction regression tree are used to determine the structural similarity prediction evaluation value of the data frame to be processed. The step of constructing a random forest prediction model corresponding to the target data includes: The training samples are input into the initial random forest prediction model, and the encoder randomly selects a training sample set from the audio and video dataset. The training samples include the audio and video data and the target data. Based on the training sample set, calculate the key feature set corresponding to the training sample set; Based on the set of key features, a regression tree is constructed, and the key features in the set of key features are prioritized according to preset feature priority information to obtain the ranking result; Based on the ranking results, the regression tree is divided using the decision tree feature with the minimum mean square error to obtain the predicted regression tree.

3. The method according to claim 2, characterized in that, Before determining the target structural similarity prediction value of the audio and video data using an image quality assessment algorithm based on the structural similarity prediction evaluation value and through concealment effects, the method further includes: Based on each leaf node in the predicted regression tree, the structural similarity prediction evaluation value corresponding to each leaf node is calculated in a round-robin fashion.

4. The method according to claim 3, characterized in that, The step of determining the target structural similarity prediction value of the audio and video data based on the structural similarity prediction evaluation value and through an image quality assessment algorithm for concealment effects includes: For the structural similarity prediction evaluation value corresponding to each leaf node, Gaussian weighting is used under the intervention of the concealment effect to calculate the target value for each polling process. The target value includes the mean, variance and covariance. The average of multiple target values ​​is determined as the target structural similarity prediction evaluation value of the audio and video data.

5. The method according to claim 4, characterized in that, The random forest prediction model includes a quadtree, which includes at least two adjacent sub-coding units on the time axis. The structural similarity prediction evaluation value corresponding to each leaf node is calculated using Gaussian weighting under the intervention of the concealment effect, and the target value is calculated for each polling process, including: Calculate the first rate-distortion value of each sub-coding unit in the at least two adjacent sub-coding units, and calculate the average rate-distortion value of the at least two adjacent sub-coding units based on the first rate-distortion value of each sub-coding unit. The first comparison result is obtained by comparing the average rate distortion value with the preset rate distortion value; If the first comparison result indicates that the average rate distortion is less than or equal to the preset rate distortion value, the number of adjacent sub-coding units on the time axis is increased. The target value for each polling cycle is calculated by adding a quadtree to the sub-coding unit.

6. The method according to claim 2, characterized in that, The initial random forest prediction model includes a quadtree partitioning module, wherein the quadtree in the quadtree partitioning module includes four sub-coding units at the same position and the parent coding unit corresponding to the four sub-coding units; the key feature set includes the rate-distortion values ​​of the sub-coding units; The construction of the regression tree based on the set of key features includes: The sum of the second rate-distortion values ​​of each of the four sub-coding units is determined as the third rate-distortion value of the four sub-coding units; By comparing the third rate-distortion value with the fourth rate-distortion value of the parent coding unit, a second comparison result is obtained; If the second comparison result indicates that the third rate distortion value is less than or equal to the fourth rate distortion value, the quadtree partitioning depth of the quadtree in the quadtree partitioning module is increased to obtain a regression tree.

7. The method according to claim 2, characterized in that, When the target structure similarity prediction evaluation value meets preset conditions, the random forest prediction model is adjusted to obtain the target encoding module, including: If the target structure similarity prediction evaluation value is greater than or equal to a preset threshold, the weight values ​​of the initial random forest model are adjusted to obtain the target random forest model; Based on the aforementioned target random forest model, a target encoding module is generated.

8. An audio and video processing device for streaming media, characterized in that, include: An acquisition module is used to acquire data frames to be processed from the audio and video data of the streaming media and target data from the encoding module. The encoding module is a module for encoding the audio and video data. The target data is the data required by the encoding module to encode historical data frames. The historical data frames are data frames that have already been encoded in the audio and video data. The target data includes at least one of the following: bitrate, quadtree partitioning depth, frame data coding block, quantization parameters, and rate distortion value. A construction module is used to construct a random forest prediction model corresponding to the target data based on the target data. The random forest prediction model is used to determine the structural similarity prediction evaluation value of the data frame to be processed. The determination module is used to determine the target structural similarity prediction evaluation value of the audio and video data based on the structural similarity prediction evaluation value and through an image quality evaluation algorithm with concealment effect. An adjustment module is used to adjust the random forest prediction model when the target structure similarity prediction evaluation value meets preset conditions, so as to obtain a target encoding module to encode the data frame to be processed.

9. A computer device, characterized in that, include: Memory and processor The memory is used to store computer programs; The processor is configured to execute a computer program stored in the memory, wherein when the computer program is executed, the processor performs the steps of the audio and video processing method for streaming media as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Video image component prediction method and device, and computer storage medium

    CA3109008A1

  • Structural similarity-based efficient video code perceiving code rate control optimizing method

    CN103634601A