Video perception model coefficient calculation method, video perception evaluation method, and device
By constructing the cost function and using the gradient descent algorithm to solve it, the problem of the calculation deviation of video perception model coefficients in VoNR video perception evaluation is solved, and the accuracy of the evaluation results is improved.
Patent Information
- Application Number
- CN202310558104.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-17
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-05-17
AI Technical Summary
In the existing VoNR video perception evaluation method, there is a deviation in the calculation of the video perception model coefficient, which leads to the inconsistent evaluation results with the actual user perception results.
By obtaining the video-aware label dataset, constructing a cost function, and solving it using the gradient descent algorithm, the target video-aware model coefficient is obtained.
The calculation accuracy of the video perception model coefficients is improved, making the evaluation results more in line with the actual perception of the user.
Smart Images

Figure CN116523896B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of wireless communication and terminal technologies, in particular to a method for calculating video perception model coefficients, a method for evaluating video perception, and a device. Background Art
[0002] After entering the 5G era, with the continuous layout of operators in 5G network coverage, VoNR (Voice over New Radio, voice service based on 5G new air interface) voice service has become the solution for audio and video in the 5G SA network architecture, and among them, VoNR video service will be one of the key development services of 5G services. The ultra-clear video service of VoNR will bring a new experience to users. However, how to evaluate the picture quality and ensure user perception during VoNR video calls will be one of the key concerns of VoNR services.
[0003] In related technologies, the method for evaluating VoNR video perception mainly focuses on the key indicators related to user transmission, and directly performs fitting evaluation and analysis through the algorithm model of video VMOS. Among them, the key coefficients of the video perception evaluation model are mainly obtained by the least squares method through simulating the relationship between subjective video perception and key indicators such as frame rate, bit rate, and packet loss rate. However, the model coefficients calculated by this method may lead to a deviation between the evaluation result and the actual user perception result. In summary, the technical problems existing in related technologies need to be solved urgently. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a method for calculating video perception model coefficients, a method for evaluating video perception, and a device to improve the calculation accuracy of video perception model coefficients.
[0005] On the one hand, the present invention provides a method for calculating video perception model coefficients, and the method includes:
[0006] Obtain a video perception label data set, where the video perception label data set includes multiple groups of video perception label data, and each group of video perception label data includes a bit rate, a frame rate, and a video quality robustness parameter;
[0007] Construct a cost function according to the video perception label data set;
[0008] Solve the cost function according to the gradient descent algorithm to obtain the target video perception model coefficients.
[0009] Optionally, the obtaining the video perception label data set includes:
[0010] Obtain several groups of initial data, and each group of initial data includes a packet loss rate, a bit rate, a frame rate, and a video perception score;
[0011] Based on the several groups of initial data, several corresponding video quality robustness parameters are calculated;
[0012] Perform tagging processing on the bitrate, the frame rate, and the video quality robustness parameters to obtain a video perception label data set.
[0013] Optionally, the expression adopted by the cost function is:
[0014]
[0015] In the formula, J represents the cost function, M represents the number of groups of the video perception label data set, i represents a variable, v8, v9, v 10 、v 11 and v 12 represent the coefficients of the video perception model to be calculated, f i represents the i-th bitrate, b i represents the i-th frame rate, y i represents the i-th video quality robustness parameter.
[0016] Optionally, solving the cost function according to the gradient descent algorithm to obtain the target video perception model coefficients includes:
[0017] Initialize the algorithm parameters, where the algorithm parameters include the coefficients of the video perception model to be calculated, the learning rate, the number of iterations, and the loop termination condition;
[0018] Perform gradient descent iterative calculation on the coefficients of the video perception model to be calculated according to the learning rate, and update the coefficients of the video perception model to be calculated;
[0019] Judge whether to stop the gradient descent iterative calculation according to the number of iterations and the loop termination condition, and determine the updated coefficients of the video perception model to be calculated at the stop as the target video perception model coefficients.
[0020] Optionally, performing gradient descent iterative calculation on the coefficients of the video perception model to be calculated according to the learning rate and updating the coefficients of the video perception model to be calculated includes:
[0021] Perform gradient descent iterative calculation on the coefficients of the video perception model to be calculated according to the learning rate, where the formula adopted for the gradient descent iterative calculation is:
[0022]
[0023] In the formula, represents the updated coefficients of the video perception model to be calculated, Denote the video perception model coefficients to be calculated as \(\alpha\), the learning rate as \(\alpha\), and the number of iterations as iter.
[0024] Optionally, the step of stopping the gradient descent iterative calculation according to the number of iterations and the loop termination condition, and determining the updated video perception model coefficients corresponding to the stop as the target video perception model coefficients includes:
[0025] Substitute the updated video perception model coefficients to be calculated into the cost function for calculation to obtain the updated cost function value;
[0026] Return to the step of performing gradient descent iterative calculation on the video perception model coefficients to be calculated according to the learning rate and updating the video perception model coefficients to be calculated until the number of iterative calculations is greater than the number of iterations or the updated cost function value is less than the loop termination condition, stop the iterative calculation, and determine the updated video perception model coefficients corresponding to the stop of the iterative calculation as the target video perception model coefficients.
[0027] On the other hand, an embodiment of the present invention further provides a video perception evaluation method, including:
[0028] Obtain the VoNR video call media plane data;
[0029] Input the VoNR video call media plane data into the video perception model constructed by the video perception model coefficients obtained by the video perception model coefficient calculation method described above to obtain the video perception evaluation result.
[0030] On the other hand, an embodiment of the present invention further provides a video perception model coefficient calculation device, including:
[0031] The first module is used to obtain a video perception label data set, where the video perception label data set includes multiple groups of video perception label data, and each group of video perception label data includes a bit rate, a frame rate, and a video quality robustness parameter;
[0032] The second module is used to construct a cost function according to the video perception label data set;
[0033] The third module is used to solve the cost function according to the gradient descent algorithm to obtain the target video perception model coefficients.
[0034] Optionally, the first module is used to obtain a video perception label data set, where the video perception label data set includes multiple groups of video perception label data, and each group of video perception label data includes a bit rate, a frame rate, and a video quality robustness parameter, including:
[0035] Obtain several groups of initial data, each group of initial data including packet loss rate, bit rate, frame rate, and video perception score;
[0036] According to the several groups of initial data, calculate the corresponding several video quality robustness parameters;
[0037] Perform labeling processing on the bit rate, the frame rate, and the video quality robustness parameters to obtain a video perception label data set.
[0038] Optionally, the expression of the cost function is:
[0039]
[0040] In the formula, J represents the cost function, M represents the number of groups of the video perception label data set, i represents a variable, v8, v9, v 10 , v 11 and v 12 represent the video perception model coefficients to be calculated, f i represents the i-th bit rate, b i represents the i-th frame rate, y i represents the i-th video quality robustness parameter.
[0041] Optionally, the third module is used to solve the cost function according to the gradient descent algorithm to obtain the target video perception model coefficients, including:
[0042] The first unit is used to initialize the algorithm parameters, and the algorithm parameters include the video perception model coefficients to be calculated, learning rate, number of iterations, and loop termination condition;
[0043] The second unit is used to perform gradient descent iterative calculation on the video perception model coefficients to be calculated according to the learning rate, and update the video perception model coefficients to be calculated;
[0044] The third unit is used to judge whether to stop the gradient descent iterative calculation according to the number of iterations and the loop termination condition, and determine the updated video perception model coefficients corresponding to the stop as the target video perception model coefficients.
[0045] Optionally, the second unit is used to perform gradient descent iterative calculation on the video perception model coefficients to be calculated according to the learning rate, and update the video perception model coefficients to be calculated, including:
[0046] Perform gradient descent iterative calculation on the video perception model coefficients to be calculated according to the learning rate, where the formula used for gradient descent iterative calculation is:
[0047]
[0048] In the formula, represents the updated video perception model coefficients to be calculated, represents the video perception model coefficients to be calculated, α represents the learning rate, and iter represents the number of iterations.
[0049] Optionally, the third unit is configured to stop the gradient descent iterative calculation according to the number of iterations and the loop termination condition, and determine the updated video perception model coefficients corresponding to the stop as the target video perception model coefficients, including:
[0050] Substitute the updated video perception model coefficients to be calculated into the cost function for calculation to obtain the updated cost function value;
[0051] Return to the step of performing gradient descent iterative calculation on the video perception model coefficients to be calculated according to the learning rate and updating the video perception model coefficients to be calculated, until the number of iterative calculations is greater than the number of iterations or the updated cost function value is less than the loop termination condition, stop the iterative calculation, and determine the updated video perception model coefficients corresponding to the stop of the iterative calculation as the target video perception model coefficients.
[0052] On the other hand, an embodiment of the present invention also discloses an electronic device, including a processor and a memory;
[0053] The memory is used to store a program;
[0054] The processor executes the program to implement the video perception model coefficient calculation method as described above.
[0055] On the other hand, an embodiment of the present invention also discloses a computer-readable storage medium, where the storage medium stores a program, and the program is executed by a processor to implement the video perception model coefficient calculation method as described above.
[0056] On the other hand, an embodiment of the present invention also discloses a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the video perception model coefficient calculation method as described above.
[0057] Compared with the prior art, the present invention adopting the above technical solutions has the following technical effects: In the embodiments of the present invention, a video perception label data set is obtained to construct a cost function, and the target video perception model coefficients are obtained through gradient descent iterative calculation of the cost function, which can calculate the video perception model coefficients according to the machine learning method, improving the accuracy of coefficient calculation. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0059] Figure 1 is a flowchart of a method for calculating video perception model coefficients provided by an embodiment of the present invention;
[0060] Figure 2 is a flowchart of step S101 provided by an embodiment of the invention;
[0061] Figure 3 is a flowchart of step S103 provided by an embodiment of the invention;
[0062] Figure 4 is a schematic structural diagram of a video perception system provided by an embodiment of the present invention;
[0063] Figure 5 is an effect diagram of reducing complaints about VoNR video call quality provided by an embodiment of the invention;
[0064] Figure 6 is an effect diagram of reducing revenue loss provided by an embodiment of the invention;
[0065] Figure 7 is an effect diagram of reducing the disposal time provided by an embodiment of the invention;
[0066] Figure 8 is an effect diagram of saving operation and maintenance costs provided by an embodiment of the invention;
[0067] Figure 9 is a schematic structural diagram of a device for calculating video perception model coefficients provided by an embodiment of the present invention;
[0068] Figure 10 is a schematic structural diagram of a device for calculating video perception model coefficients provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0069] To make the objectives, technical solutions, and advantages of this application clearer and more understandable, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for explaining this application and are not intended to limit this application.
[0070] First, analyze several terms involved in this application:
[0071] VoNR (Voice over New Radio, voice service based on 5G new air interface) is a call service that can be completed only with a 5G network and can provide ultra-high-definition video call services.
[0072] The packet loss rate refers to the ratio of the number of lost data packets to the number of transmitted data packets in a test.
[0073] The bit rate refers to the number of bits (bit) transmitted per second, also known as the bit rate. The higher the bit rate, the faster the data transmission speed.
[0074] The frame rate is a measure used to measure the number of displayed frames.
[0075] The Video Mean Opinion Score (VMOS) is a video quality evaluation method that judges the video quality by normalizing the evaluations of observers.
[0076] Deep Packet Inspection (DPI) technology adds application protocol identification, data packet content detection, and deep decoding of application layer data on top of traditional IP data packet detection technology. Current in-network DPI collection devices have both traditional IP data packet detection technology and deep packet detection technology.
[0077] The ITU-T G.1070 specification is a standard for no-reference evaluation of audio and video quality issued by the ITU. It supports not only no-reference evaluation of videos alone, but also no-reference evaluation of audio alone and integration of video and audio evaluation results.
[0078] In related technologies, the VoNR video perception evaluation methods are mainly divided into two types: plug-in and non-plug-in.
[0079] The plug-in method mainly conducts comparative analysis by relying on reference sample files. Currently, the more traditional method for evaluating video call quality is to analyze based on drive test software. The plug-in method evaluates by relying on reference video samples, but this evaluation test process is cumbersome, time-consuming, laborious, and inefficient. At the same time, it cannot reflect the situation of the entire network. In addition, this method can only reflect the video quality during the test and requires a huge investment in labor costs.
[0080] The non-insertion type mainly focuses on the key indicators related to user transmission, and directly conducts fitting evaluation and analysis through the algorithm model of video VMOS. Compared with the traditional insertion type of video evaluation, the non-insertion type of direct evaluation and scoring method based on DPI data does not require test sample data, which is more in line with the actual situation of users. At the same time, it can conduct network-wide evaluation, saving a large amount of labor costs.
[0081] Regarding the non-insertion type of user perception evaluation method for video services, currently the industry mainly refers to the ITU-T G.1070 specification. The ITU-T G.1070 video perception evaluation model mainly refers to key indicators such as packet loss rate, frame rate, bit rate, resolution, and video level for evaluation. Calculating the video perception VMOS involves 12 key coefficients from V1 to V12. The 12 key coefficients of the existing video perception evaluation model are mainly obtained by the least squares method through simulating the relationship between the subjective video perception VMOS and key indicators such as frame rate, bit rate, and packet loss rate. The coefficients calculated in this way have the following drawbacks:
[0082] 1. The 12 key coefficients involved in the existing video perception evaluation model are not fitted based on the VoNR video call scenario. If these 12 coefficients are directly used for the VoNR video call service, it may lead to a deviation between the evaluation result and the actual user perception result. For example, for the ofr optimal frame rate, this indicator is fitted through two coefficients, V1 and V2, combined with Brv (bit rate). Taking the H264 encoding method, 480P resolution, and BP video level as an example, if the optimal frame rate is 30 frames, the corresponding Brv (bit rate) needs to reach more than 2300 kbps. However, according to the actual VoNR video call packet capture verification, such a large Brv (bit rate) is not required during the actual transmission process.
[0083] 2. Currently, the coefficients only fit some encoding methods, resolutions, and video level scenarios, and do not cover all scenarios, especially the content related to VoNR video calls. In addition, as the VoNR video service gradually develops, the H265 encoding format video service will gradually increase. Currently, the ITU-T G.1070 also does not have an explanation of the relevant coefficients for the H265 video encoding format, which will also lead to the inability to conduct video perception evaluation for the H265 encoding format video call scenario.
[0084] 3. The 12 key coefficients for which the coefficient fitting depends on the video perception VMOS are obtained through subjective evaluation by humans, but the subjective evaluation is greatly affected by humans.
[0085] 4. The least squares method will have a certain deviation for non-linear relationship content.
[0086] Therefore, for the 12 coefficients in the VoNR video perception evaluation model, a more intelligent, efficient, and comprehensive method needs to be constructed for fitting, so that the VoNR video perception evaluation VMOS is more in line with the actual perception of users.
[0087] Referring to Figure 1 , an embodiment of the present invention provides a method for calculating video perception model coefficients, including:
[0088] S101. Obtain a video perception label data set, where the video perception label data set includes multiple groups of video perception label data, and each group of video perception label data includes a bit rate, a frame rate, and a video quality robustness parameter;
[0089] S102. Construct a cost function according to the video perception label data set;
[0090] S103. Solve the cost function according to the gradient descent algorithm to obtain target video perception model coefficients.
[0091] In the embodiment of the present invention, first, a video perception label data set is obtained. The video perception label data set includes multiple groups of video perception label data, and each group of video perception label data includes a bit rate, a frame rate, and a video quality robustness parameter. Among them, the bit rate and the frame rate can be obtained through a DPI acquisition and parsing device, and the video quality robustness parameter represents the degree of video quality robustness caused by packet loss. In the embodiment of the present invention, the bit rate, the frame rate, and the video quality robustness parameter are a set of data. Then, a cost function is constructed according to the obtained video perception label data set, and the cost function is a mean square error cost function. Finally, the cost function is iteratively calculated based on the gradient descent algorithm to obtain target video perception model coefficients. In the embodiment of the present invention, the target video perception model coefficients include v8 to v 12 . Since different screen sizes, resolutions, and other factors correspond to different 12 video perception model coefficients, the related technology algorithms divide the solution of these coefficients into two stages. First, the coefficients v1 to v7 are solved, and then v8 to v 12 are solved. In both stages, these coefficients are estimated by the least squares method. However, the coefficients in the second stage have a non-linear relationship with the input, and there are errors in the fitting result of the least squares method. To solve this problem, the embodiment of the present invention proposes a method for solving the coefficients in the second stage based on the gradient descent algorithm, which can adaptively find the global optimal solution and achieve good results in both linear models and non-linear models, making the final solution result close to the expected value and improving the calculation accuracy of the video perception model coefficients.
[0092] Further as an optional implementation manner, referring to Figure 2 , in the above step S101, the obtaining of the video perception label data set includes:
[0093] S201. Obtain several groups of initial data, where each group of initial data includes packet loss rate, bit rate, frame rate, and video perception score;
[0094] S202. Calculate corresponding video quality robustness parameters based on the several groups of initial data;
[0095] S203. Perform marking processing on the bit rate, the frame rate, and the video quality robustness parameters to obtain a video perception label dataset.
[0096] In the embodiment of the present invention, obtaining several groups of initial data, where each group of initial data includes packet loss rate, bit rate, frame rate, and video perception score, can be obtained by parsing the media plane information of VoNR video calls through a DPI acquisition and parsing device. The parsed content includes information such as the coding method, frame rate, and resolution of VoNR video calls. The embodiment of the present invention can construct sample data using different coding methods, resolutions, and video levels according to the packet loss rate, frame rate, and coding rate indicators collected from VoNR video calls to obtain several groups of initial data. Then, based on the packet loss rate P plv , bit rate Fr V , frame rate Br V and video perception score vMOS, calculate the video quality robustness parameters through the algorithm provided based on the G.1070 specification, and calculate the corresponding video quality robustness parameter D Pplv for each group of initial data. The number of video quality robustness parameters is the same as the number of groups of initial data. Finally, perform marking processing on the bit rate, frame rate, and video quality robustness parameters to obtain a video perception label dataset. That is, extract and mark the corresponding bit rate, frame rate, and video quality robustness parameters of the same group to obtain sample label data, and integrate multiple groups of sample label data to obtain a video perception label dataset, as shown in Table 1 below. In the table, Fr V represents the bit rate, Br V represents the frame rate, D Pplv represents the video quality robustness parameter, and fi, bi, and yi respectively represent the values of the bit rate, frame rate, and video quality robustness parameter in the i-th group of sample label data.
[0097]
[0098] Table 1 Video perception label dataset
[0099] Based on the constructed sample label data, the embodiment of the present invention can automatically output video perception model coefficients and apply them to the evaluation of video VMOS. Compared with traditional manual simulation and calculation, the embodiment of the present invention can greatly reduce labor costs and time costs.
[0100] As a further optional implementation manner, in the above step S102, the expression adopted by the cost function is as follows:
[0101]
[0102] In the formula, J represents the cost function, M represents the number of groups of the video perception label dataset, i represents a variable, v8, v9, v 10 , v 11 and v 12 represent the video perception model coefficients to be calculated, f i represents the i-th bit rate, b i represents the i-th frame rate, and y i represents the i-th video quality robustness parameter.
[0103] In the embodiment of the present invention, a cost function is constructed through the video perception label dataset. The cost function is defined on the entire training set and is the average of all sample errors, that is, the average of all loss function values. The video perception label dataset includes multiple groups of sample label data, and each group of sample label data includes a bit rate, a frame rate, and a video quality robustness parameter. Substituting the values in Table 1 above into the cost function can perform cost calculation, where the video perception model coefficients to be calculated are unknowns. The purpose of the embodiment of the present invention is to calculate the video perception model coefficients, and the optimal video perception model coefficients v8 to v 12 will make the value of the cost function J the smallest, that is, statistically make the M D 12 estimated based on v8 to v Pplv closest to the y in the table, that is:
[0104]
[0105] In the formula, M represents the number of groups of the video perception label dataset, i represents a variable, v8, v9, v 10 , v 11 and v 12 represent the video perception model coefficients to be calculated, f i represents the i-th bit rate, b i represents the i-th frame rate, and y i represents the i-th video quality robustness parameter.
[0106] As a further optional implementation manner, referring to Figure 3 , in the above step S103, the solving of the cost function according to the gradient descent algorithm to obtain the target video perception model coefficients includes:
[0107] S301. Initialize the algorithm parameters, where the algorithm parameters include the video perception model coefficients to be calculated, a learning rate, the number of iterations, and a loop termination condition;
[0108] S302. Perform gradient descent iterative calculation on the video perception model coefficients to be calculated according to the learning rate, and update the video perception model coefficients to be calculated.
[0109] S303. Determine whether to stop the gradient descent iterative calculation according to the number of iterations and the loop termination condition, and determine the updated video perception model coefficients corresponding to the stop as the target video perception model coefficients.
[0110] In the embodiment of the present invention, the gradient descent algorithm is used to perform iterative calculation on the cost function to obtain the target video perception model coefficients. First, the algorithm parameters are initialized. The algorithm parameters include the video perception model coefficients v8 to v 12 ., the learning rate α, the number of iterations iter, and the loop termination condition ε. Among them, the learning rate is used to update the video perception model coefficients, and both the number of iterations and the loop termination condition are used to control the iterative calculation. The loop termination condition is used to compare with the cost function value. When the cost function value is less than the loop termination condition, the iterative calculation is stopped. Then, perform gradient descent iterative calculation on the video perception model coefficients to be calculated according to the learning rate, and update the video perception model coefficients to be calculated. Finally, determine whether to stop the gradient descent iterative calculation according to the number of iterations and the loop termination condition, and determine the updated video perception model coefficients corresponding to the stop as the target video perception model coefficients. That is, when the above gradient descent iterative calculation reaches the corresponding number of iterations, or the loop termination condition is satisfied during the calculation process, the iterative calculation will stop, and the video perception model coefficients calculated at this time are used as the target video perception model coefficients. Generally, the iterative calculation will stop satisfying the loop termination condition before reaching the number of iterations. If the above gradient descent iterative calculation stops when the corresponding number of iterations is reached, the embodiment of the present invention can also record the video perception model coefficients to be calculated updated during each iterative calculation process, and select the optimal video perception model coefficients as the target video perception model coefficients. In a feasible implementation manner, the initialized algorithm parameters are respectively assigned values of v8 = 0.1, v9 = 513.8, v 10 = 0.7, v 11 = 6.5, v 12 = 13.7, α = 0.1, iter = 200, ε = 0.001.
[0111] Further as an optional implementation manner, the performing gradient descent iterative calculation on the video perception model coefficients to be calculated according to the learning rate, and updating the video perception model coefficients to be calculated includes:
[0112] Perform gradient descent iterative calculation on the video perception model coefficients to be calculated according to the learning rate. The formula used for gradient descent iterative calculation is as follows:
[0113]
[0114] In the formula, represents the updated video perception model coefficients to be calculated, represents the video perception model coefficients to be calculated, α represents the learning rate, and iter represents the number of iterations.
[0115] In the embodiment of the present invention, the video perception model coefficients v8 to v are updated respectively according to the gradient descent algorithm 12 . Assume that the k-th iterative calculation is in progress. The gradient descent iterative calculation formula is as follows:
[0116]
[0117] The embodiment of the present invention defines the intermediate calculation result of the i-th input in the k-th iteration
[0118]
[0119] Then:
[0120]
[0121]
[0122]
[0123]
[0124]
[0125] In the formula, J represents the cost function. Substitute the above formula into the gradient descent iterative calculation formula respectively to iteratively update the video perception model coefficients v8 to v 12 for iterative update.
[0126] Further as an optional implementation manner, the stopping judgment of the gradient descent iterative calculation is performed according to the number of iterations and the loop termination condition, and the updated video perception model coefficients corresponding to the stop are determined as the target video perception model coefficients, including:
[0127] Substitute the updated video perception model coefficients to be calculated into the cost function for calculation to obtain the updated cost function value;
[0128] Return to the step of performing gradient descent iterative calculation on the video perception model coefficients to be calculated according to the learning rate and updating the video perception model coefficients to be calculated, until the number of iterative calculations is greater than the number of iterations or the updated cost function value is less than the loop termination condition, stop the iterative calculation, and determine the updated video perception model coefficients corresponding to when the iterative calculation stops as the target video perception model coefficients.
[0129] In an embodiment of the present invention, the gradient descent iterative calculation is stopped and judged according to the number of iterations and the loop termination condition, that is, during the above-mentioned gradient descent iterative calculation loop process, if the number of loops reaches the number of iterations or meets the loop termination condition, the iterative calculation can be stopped, and the updated video perception model coefficients corresponding to when it stops are determined as the target video perception model coefficients. In a feasible implementation manner, the above iterative calculation meets the loop termination condition in the 67th calculation, that is, the video perception model coefficients calculated in the 67th iterative calculation are used as the target video perception model coefficients. In each iterative calculation process, the embodiment of the present invention substitutes the updated video perception model coefficients to be calculated into the cost function constructed above for calculation to obtain the updated cost function value, and compares the updated cost function value with the loop termination condition. When the cost function value is less than the loop termination condition, the iterative calculation can be stopped, and the updated video perception model coefficients corresponding at this time are the target video perception model coefficients finally solved.
[0130] On the other hand, the embodiment of the present invention also provides a video perception evaluation method, including:
[0131] Obtain VoNR video call media plane data;
[0132] Input the VoNR video call media plane data into the video perception model constructed by the video perception model coefficients obtained by the video perception model coefficient calculation method described above to obtain a video perception evaluation result.
[0133] In the embodiments of the present invention, first, the VoNR video call media plane data is obtained. This data can be collected and parsed by a DPI acquisition and parsing device, and the parsed content includes information such as the coding method, frame rate, and resolution of the VoNR video call. Then, the VoNR video perception model can be constructed using the video perception model coefficients obtained by the video perception model coefficient calculation method described above. Finally, the VoNR video call media plane data is input into the video perception model for video perception evaluation to obtain the video perception evaluation result. By applying the video perception model coefficients obtained by the above video perception model coefficient calculation method to the video perception model, the VoNR video perception VMOS evaluation scenario in the embodiments of the present invention is more comprehensive, enabling the present invention to perform learning for specific scenarios of VoNR video services. The embodiments of the present invention can better support the VoNR video perception VMOS evaluation under different resolutions and video levels in the H264 and H265 coding methods, and can also solve the problem that the G.1070 specification does not provide relevant parameters for the H265 coding method, resulting in the inability to perform video perception evaluation for video calls using the H265 coding method.
[0134] Referring to Figure 4 , the embodiments of the present invention also provide a video perception system, which includes a DPI acquisition and parsing device, a perception parameter output module, and a VoNR video perception evaluation VMOS output module;
[0135] Among them, the DPI acquisition and parsing device is used for parsing the VoNR video call media plane bitstream information, and the parsed content includes information such as the coding method, frame rate, and resolution of the VoNR video call.
[0136] The perception parameter output module is used to implement the specific content of a video perception model coefficient calculation method provided by the embodiments of the present invention, and output the relevant coefficients involved in the VoNR video perception model, that is, the video perception model coefficients.
[0137] The VoNR video perception evaluation VMOS output module outputs the final VoNR video perception evaluation VMOS result based on the basic index information of the VoNR video call obtained by the DPI acquisition and parsing device and the video perception model coefficients obtained by the machine learning module.
[0138] The embodiments of the present invention output video perception model coefficients through the sample data of VoNR video calls, and the output video perception evaluation results are more in line with the actual perception of users. The traditional video perception model coefficients are not based on the VoNR video call scenario, and there may be certain deviations in the final video VMOS results. The embodiments of the present invention output video perception model coefficients through artificial intelligence, which can be used as a general algorithm module in similar video call quality evaluation scenarios and can be smoothly evolved to other video perception evaluations.
[0139] Referring to Figures 5 - 8 , in an implementation scenario, the video perception system based on the embodiments of the present invention can accurately evaluate the VoNR video call quality, improve the digital service ability of end-to-end user perception, reduce the user complaint volume and churn rate. The VONR user complaint volume has dropped by 20% compared with that before the application, and the user satisfaction has been improved. In this implementation scenario, by running the video perception system of the embodiments of the present invention, 60,000 potential lost customers have been reduced. Calculated according to the annual ARPU value of 500 yuan for ordinary public customers, the annual reduction of customer revenue loss is about 30 million yuan. In addition, the workload of manual analysis has been greatly reduced. The problem handling time has been reduced from 5 days to 2 days, and the efficiency has been improved by 60%. Calculated according to the monthly average savings of 11 vehicles, 11 test devices, and 6 person-months by the provincial company and 11 local networks, the province can save about 918,000 yuan in operation and maintenance costs every year, which has relatively high economic benefits.
[0140] Combined with Figure 1 , the process of the present invention specifically includes: first, obtaining a video perception label data set, where the video perception label data set includes multiple groups of video perception label data, and each group of video perception label data includes a bit rate, a frame rate, and a video quality robustness parameter; then constructing a cost function according to the video perception label data set; and finally solving the cost function according to the gradient descent algorithm to obtain the target video perception model coefficient. In the related art, the G.1070 specification used for video perception evaluation needs to calculate the video perception score under this condition based on the bit rate, frame rate, packet loss rate of the video plus 12 constant coefficients, that is, the video perception model coefficient. However, in the related art, the video perception model coefficient is fitted by the least squares method, which may lead to a deviation between the evaluation result and the result actually perceived by users. Therefore, the embodiments of the present invention introduce an artificial intelligence method, use a large number of key indicators of different resolutions, video coding methods, bit rates, frame rates, and packet loss rates and the inserted video evaluation results as the video perception label data set for automatic learning, construct a cost function and calculate based on the gradient descent algorithm, and automatically summarize and output the video perception model coefficient involved in video perception evaluation. The embodiments of the present invention can make the calculation of the video perception model coefficient more intelligent, more efficient, and more comprehensive, and improve the accuracy of video perception.
[0141] On the other hand, referring to Figure 9 , the embodiments of the present invention also provide a video perception model coefficient calculation device, including:
[0142] A first module 901, configured to obtain a video perception label data set, where the video perception label data set includes multiple groups of video perception label data, and each group of video perception label data includes a bit rate, a frame rate, and a video quality robustness parameter;
[0143] A second module 902, configured to construct a cost function according to the video perception label data set;
[0144] A third module 903, configured to solve the cost function according to the gradient descent algorithm to obtain target video perception model coefficients.
[0145] It can be understood that the content in the embodiments of the above video perception model coefficient calculation method is applicable to the embodiments of this video perception model coefficient calculation device. The functions specifically implemented in the embodiments of this video perception model coefficient calculation device are the same as those in the embodiments of the above video perception model coefficient calculation method, and the beneficial effects achieved are also the same as those in the embodiments of the above video perception model coefficient calculation method.
[0146] Referring to Figure 10 , an embodiment of the present invention further provides an electronic device, including a processor 1001 and a memory 1002; the memory is used to store a program; the processor executes the program to implement the method as described above.
[0147] Similarly, the content in the embodiments of the above video perception model coefficient calculation method is applicable to the embodiments of this video perception model coefficient calculation device. The functions specifically implemented in the embodiments of this video perception model coefficient calculation device are the same as those in the embodiments of the above video perception model coefficient calculation method, and the beneficial effects achieved are also the same as those in the embodiments of the above video perception model coefficient calculation method.
[0148] Corresponding to Figure 1 the method, an embodiment of the present invention further provides a computer-readable storage medium, where the storage medium stores a program, and the program is executed by a processor to implement the method as described above.
[0149] An embodiment of the present invention also discloses a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes Figure 1 the method shown.
[0150] In the related art, the 12 video perception model coefficients provided by the ITU-T G.1070 specification mainly have the following disadvantages:
[0151] 1. The 12 key parameters involved in the existing video perception evaluation model are not fitted based on the VoNR video call scenario. If these 12 coefficients are directly used for the VoNR video call service, it may lead to a deviation between the evaluation result and the result actually perceived by users.
[0152] 2. Currently, the specification coefficient does not cover the H265 encoding method. With the gradual development of the VoNR video service, the video service in the H265 encoding format will gradually increase. Currently, G.1070 does not have the description of the relevant parameters of the H265 video encoding format, which will also lead to the inability to perform video perception evaluation for calls in the H265 encoding format. Among the current video call services in the live network, 5% of the video calls already use the H265 format encoding method. With the gradual development of the VoNR service, the H265 video encoding method will also gradually increase, and there is currently an evaluation blind spot for this part of the calls.
[0153] 3. The 12 key parameters for calculating parameter fitting rely on the video perception VMOS, which is subjectively evaluated by humans. The subjective perception evaluation is greatly interfered by human factors and lacks standardized evaluation data. At the same time, the sample data for human evaluation is limited, and the fitted parameters may have certain deviations.
[0154] 4. The least squares method will have certain deviations for non-linear relationships and lacks intelligent means to improve the video quality evaluation parameter method.
[0155] In summary, the embodiments of the present invention have the following advantages:
[0156] 1. The method for obtaining the VoNR video perception evaluation model coefficient in the embodiments of the present invention is more intelligent. Based on the constructed sample label data of the present invention, the video perception model coefficient can be automatically output and applied to the evaluation of video VMOS. Compared with the traditional manual simulation and calculation, the embodiments of the present invention can greatly reduce the labor cost and time cost.
[0157] 2. The VoNR video perception VMOS evaluation scenario in the embodiments of the present invention is more comprehensive. Through the present invention, specific scenarios of the VoNR video service can be learned. It can better support the VoNR video perception VMOS evaluation under different resolutions and video levels in the H264 and H265 encoding methods. It can solve the problem that the G.1070 specification does not provide relevant parameters for the H265 encoding method, resulting in the inability to perform video perception evaluation for video calls in the H265 encoding method.
[0158] 3. The VoNR video perception VMOS result in the embodiments of the present invention is more accurate. The embodiments of the present invention output the video perception evaluation model coefficient through the sample data of VoNR video calls, and the output VMOS result is more in line with the actual perception of users. The traditional video perception model coefficient is not based on the VoNR video call scenario, and there may be certain deviations in the final video VMOS result.
[0159] 4. The embodiments of the present invention have stronger extensibility for video perception evaluation. The embodiments of the present invention output the coefficients of the video perception evaluation model through artificial intelligence, which can be used as a general algorithm module in similar video call quality evaluation scenarios and can smoothly evolve to other video perception evaluations.
[0160] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order mentioned in the operation diagrams. For example, depending on the functions / operations involved, two consecutive blocks shown may actually be executed substantially simultaneously or the blocks can sometimes be executed in the reverse order. In addition, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated in which the order of various operations is changed and in which sub-operations described as part of a larger operation are executed independently.
[0161] Furthermore, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features described may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It is also understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. Rather, considering the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skill of an engineer. Therefore, those skilled in the art can implement the present invention as set forth in the claims without undue experimentation. It is also understood that the specific concepts disclosed are illustrative only and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.
[0162] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, etc., which can store program codes of various kinds.
[0163] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or used in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in combination with an instruction execution system, apparatus, or device.
[0164] More specific examples (nonexhaustive list) of computer-readable media include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or otherwise processing it as appropriate, and then storing it in a computer memory.
[0165] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0166] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0167] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.
[0168] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.
Claims
1. A method for calculating video perception model coefficients, characterized in that, The method includes: Obtaining a video perception label dataset, where the video perception label dataset includes multiple groups of video perception label data, and each group of video perception label data includes a bit rate, a frame rate, and a video quality robustness parameter, and the video quality robustness parameter represents the degree of video quality robustness caused by packet loss; The obtaining of the video perception label dataset includes: Obtaining several groups of initial data, and each group of initial data includes a packet loss rate, a bit rate, a frame rate, and a video perception score; Calculating corresponding video quality robustness parameters according to the several groups of initial data; Performing a labeling process on the bit rate, the frame rate, and the video quality robustness parameter to obtain a video perception label dataset; Constructing a cost function according to the video perception label dataset; The expression adopted by the cost function is: Wherein, J represents the cost function, M represents the number of groups of the video perception label dataset, i represents a variable, v8, v9, v 10 , v 11 and v 12 represent the coefficients of the video perception model to be calculated, f i represents the i-th bit rate, b i represents the i-th frame rate, y i represents the i-th video quality robustness parameter; Solving the cost function according to the gradient descent algorithm to obtain target video perception model coefficients.
2. The method according to claim 1, wherein The solving of the cost function according to the gradient descent algorithm to obtain target video perception model coefficients includes: Initializing algorithm parameters, where the algorithm parameters include video perception model coefficients to be calculated, a learning rate, the number of iterations, and a loop termination condition; Performing gradient descent iterative calculation on the video perception model coefficients to be calculated according to the learning rate, and updating the video perception model coefficients to be calculated; Judging whether to stop the gradient descent iterative calculation according to the number of iterations and the loop termination condition, and determining the updated video perception model coefficients corresponding to the stop as the target video perception model coefficients.
3. The method according to claim 2, characterized in that, The performing of gradient descent iterative calculation on the video perception model coefficients to be calculated according to the learning rate and updating the video perception model coefficients to be calculated includes: Performing gradient descent iterative calculation on the video perception model coefficients to be calculated according to the learning rate, where the formula adopted during the gradient descent iterative calculation is: In the formula, represents the updated video perception model coefficients to be calculated, represents the video perception model coefficients to be calculated, α represents the learning rate, and iter represents the number of iterations.
4. The method according to claim 2, wherein The judging whether to stop the gradient descent iterative calculation according to the number of iterations and the loop termination condition, and determining the updated video perception model coefficients corresponding to the stop as the target video perception model coefficients includes: Substituting the updated video perception model coefficients to be calculated into the cost function for calculation to obtain an updated cost function value; Returning to the step of performing gradient descent iterative calculation on the video perception model coefficients to be calculated according to the learning rate and updating the video perception model coefficients to be calculated until the number of iterative calculations is greater than the number of iterations or the updated cost function value is less than the loop termination condition, stopping the iterative calculation, and determining the updated video perception model coefficients corresponding to the stop as the target video perception model coefficients.
5. A video perception evaluation method, characterized in that, including: Obtaining VoNR video call media plane data; Inputting the VoNR video call media plane data into a video perception model constructed by the video perception model coefficients obtained by the video perception model coefficient calculation method according to any one of claims 1-4 to obtain a video perception evaluation result.
6. A video perception model coefficient calculation device, characterized in that The device includes: The first module is used to obtain a video perception label dataset, where the video perception label dataset includes multiple groups of video perception label data, and each group of video perception label data includes a bit rate, a frame rate, and a video quality robustness parameter, and the video quality robustness parameter represents the degree of video quality robustness caused by packet loss; The first module, which is used to obtain a video perception label dataset, includes: Obtain several groups of initial data, and each group of initial data includes a packet loss rate, a bit rate, a frame rate, and a video perception score; According to the several groups of initial data, calculate the corresponding several video quality robustness parameters; Perform labeling processing on the bit rate, the frame rate, and the video quality robustness parameter to obtain a video perception label dataset; The second module is used to construct a cost function according to the video perception label dataset; the expression adopted by the cost function is: Where J represents the cost function, M represents the number of groups of the video perception label dataset, i represents a variable, v8, v9, v 10 , v 11 and v 12 represent the coefficients of the video perception model to be calculated, f i represents the i-th bitrate, b i represents the i-th frame rate, y i represents the i-th video quality robustness parameter; The third module is used to solve the cost function according to the gradient descent algorithm to obtain the target video perception model coefficients.
7. An electronic device, characterized in that, The electronic device includes a memory and a processor; The memory is used to store programs; The processor executes the program to implement the video perception model coefficient calculation method according to any one of claims 1 to 4.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the video perception model coefficient calculation method according to any one of claims 1 to 4.
Citation Information
Patent Citations
A video evaluation method and a device
CN109005402A
VoNR quality evaluation method and device based on 5G traffic statistic index
CN113923704A