Video quality evaluation system

The system addresses real-time video quality evaluation by segmenting continuous transmission and using machine learning to assess video quality, enabling immediate adjustments and alerts, thus improving subjective assessment accuracy.

JP2026068644APending Publication Date: 2026-04-22TOYOTA JIDOSHA KK
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
TOYOTA JIDOSHA KK
Filing Date
2024-10-10
Publication Date
2026-04-22

AI Technical Summary

Technical Problem

Conventional video quality evaluation systems are unable to provide real-time assessments during video transmission or playback, limiting the ability to adjust settings or alert users to quality deterioration.

Method used

A system that extracts video segments of predetermined length from continuous transmission and uses a machine learning algorithm to output sequential quality evaluations, incorporating non-video conditions for improved accuracy.

Benefits of technology

Enables real-time video quality evaluation during transmission, allowing for immediate adjustments and alerts, providing a more subjective and timely quality assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026068644000001_ABST
    Figure 2026068644000001_ABST
Patent Text Reader

Abstract

The present invention provides a system for subjectively evaluating video quality, configured to evaluate video quality in real time during video transmission or playback between two points, rather than on a per-video-file basis. [Solution] The system for evaluating video quality includes a video extraction means 31 that sequentially extracts video of a predetermined length from a continuous video being transmitted, and an evaluation value output means 32 that receives the sequentially extracted video and sequentially outputs an evaluation value of the quality of the input video. The evaluation value output means is trained by a machine learning algorithm to output an evaluation value of the video quality by a person who has viewed a training video of a predetermined length when the training video is input.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a system for evaluating the quality of video (moving images), and more particularly to a system capable of evaluating the quality of video in real time.

Background Art

[0002] Various technologies have been proposed for evaluating video quality. For example, Patent Document 1 proposes capturing video displayed on a mobile device's monitor, comparing the captured video signal (degraded video) with the pixels of a pre-prepared reference video to match them, outputting a matched video signal which is a matched pair of the reference video and the degraded video, and determining an objective evaluation value based on the matched video signal and accompanying information which is the frame rate information per unit time of the reference video and the degraded video. Patent Document 2 proposes transmitting camera video from one terminal to another terminal, receiving camera video transmitted back from the other terminal, displaying the transmitted and received video on a display, and acquiring and displaying communication quality information in video transmission. Patent Document 3 proposes a method for evaluating the quality of received video without the original video using artificial intelligence composed of multiple convolutional neural networks and a cyclic neural network with settable learning ranges. Patent Document 4 discloses an attempt to improve the accuracy of estimating the quality of videos with fluctuating bitrates by obtaining feature quantities related to the fluctuation of the video bitrate from the bitrate time series of the video distributed over a network, and deriving an estimated value of the video quality using the video encoding information time series including the bitrate time series and the feature quantities as input. Non-Patent Document 1 proposes an attempt to reduce the amount of computation in a deep video quality assessment method that uses a neural network to evaluate the quality of a video by referring only to the video to be evaluated, by performing calculations while clipping and extracting each image in the time direction, thereby enabling performance evaluation at a short speed regardless of the length of the video. Non-Patent Document 2 proposes a video quality evaluation method that uses two video files, one to be evaluated and one to be referenced, to output a time series score all at once, and is a method widely used in the market as a method for determining encoding parameters. [Prior art documents] [Patent Documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2012-119772 [Patent Document 2] Japanese Patent Publication No. 2016-213784 [Patent Document 3] Japanese Patent Publication No. 2023-152957 [Patent Document 4] International release 2023 / 233631 [Non-patent literature]

[0004] [Non-Patent Document 1] H. Wu et al., “FAST-VQA: Efficient End-to-end Video Quality Assessment with Fragment Sampling”, Proceedings of the European Conference on Computer Vision (ECCV) 2022, https: / / doi.org / 10.1007 / 978-3-031-20068-7_31, arXiv:2207.02595v1 [cs.CV] 6 Jul 2022 [Non-Patent Document 2] vmaf: Perceptual video quality assessment based on multi-method fusion, “https: / / github.com / Netflix / vmaf”, Netflix, Inc., 2017-07-14, retrieved 2017-07-15 [Overview of the Initiative] [Problems that the invention aims to solve]

[0005] In video quality evaluation systems that follow the video quality evaluation techniques proposed in Non-Patent Documents 1 and 2, an attempt is made to score the level of video quality as perceived by a person when they view the video to be evaluated. To this end, in the systems of Non-Patent Documents 1 and 2, generally speaking, various videos and the scores (evaluation values) of the quality of each video as evaluated by people who view each video are prepared as input data and ground truth data in the training data, respectively. A classifier is configured to output a corresponding evaluation value when various videos of the training data are input, according to a machine learning algorithm such as deep learning. When evaluating the quality of any video in such a system, the video to be evaluated is input to the classifier, and the evaluation value output from the classifier is used as an index value for the quality of the video to be evaluated. It should be noted that in such systems, there are known methods for determining the video quality evaluation value by referring only to the video to be evaluated (non-reference method - in the case of Non-Patent Document 1) and methods for determining it by referring to both the video to be evaluated and a video that serves as a standard for evaluation (full-reference method - in the case of Non-Patent Document 2). The evaluation results are referenced, for example, when adjusting various settings and parameters such as bitrate during video preparation.

[0006] Incidentally, if the video quality evaluation values ​​described above could be obtained in real time while the video was being used, that is, while it was being played back or transmitted, it would be convenient to be able to control various settings and parameters for adjusting or preparing the video while it was being used. For example, when transmitting video between any two points (such as between a vehicle and a server on a network) via a wireless or wired network, if the video quality at the destination could be evaluated in real time during the transmission or while the video was being played back while being transmitted, it would be convenient to be able to adjust various settings at the source of the video transmission or during transmission according to the evaluation, or to immediately send alerts when the quality deteriorates. In this regard, conventional video quality evaluation systems (Non-Patent Documents 1 and 2) output video quality evaluation values ​​on a per-video file basis and are not configured to obtain real-time quality evaluations while the video is being transmitted or played back.

[0007] In view of the above circumstances, the main object of the present invention is to provide a system for evaluating video quality, configured such that video quality is evaluated in real time during video transmission or playback, rather than for each video file. [Means for solving the problem]

[0008] According to one aspect of the present invention, the above problem is addressed by a system for evaluating the quality of video, A video extraction means configured to sequentially extract video of a predetermined length from a continuous video in transmission, An evaluation value output means configured to sequentially input the extracted video and sequentially output an evaluation value of the quality of the input video, wherein when a training video of a predetermined length is input, the evaluation value output means is trained by a machine learning algorithm to output the evaluation value given by a person who has watched the training video. This is achieved by a system that includes this.

[0009] In the above configuration, "video" may refer to any moving image. Generally, video is converted into a signal sequence representing the brightness of each pixel on the screen and transmitted sequentially through a line (the signal sequence recording video information is called a "video signal sequence"). In this invention, "video quality" refers to the quality of the video as perceived by a person when they view it, and the quality is expressed as an evaluation value scored (quantified) by human judgment. "Continuous video in transmission" refers to continuous video transmitted between two points, specifically, video that is transmitted sequentially as a video signal sequence over a line when video captured or generated by a camera or video generator is converted into a signal sequence and sent to a video recording device or video display or playback device, or when video is sent from one video recording device to another recording device or video display or playback device. The line through which the video signal sequence is transmitted may be a wireless line or a wired line. The shortest connection may be, for example, a connection from a camera or image generating device to a display, while a longer connection may be, for example, a connection between a vehicle and a server on a network. The connection may be any connection on a communication network such as the Internet. The "image extraction means" is configured to sequentially extract a signal sequence corresponding to an image of a predetermined length from a continuous image in transmission, that is, a continuous sequence of image signals being transmitted sequentially over the connection. Here, the "predetermined length" is the length of time that a person can view the image and determine its quality at one time, and is typically, for example, 4 seconds. The extracted images may or may not be duplicates.

[0010] The "evaluation value output means" is configured to output an evaluation value representing the quality evaluation of a video of a predetermined length extracted by the video extraction means. More specifically, as described above, the evaluation value output means is configured to output an evaluation value of the quality of a video as determined by a person who has viewed the training video, when a training video of a predetermined length is input, using a machine learning algorithm. Here, the training video is any various arbitrary video of a predetermined length, and the score obtained by actually viewing and perceiving the quality of each such video is used as the evaluation value for each training video. The evaluation value output means is a classifier configured to assign evaluation values ​​to videos according to an arbitrary machine learning algorithm, as described above, using the training video as input data and the corresponding evaluation value as the ground truth value. As the machine learning algorithm, any machine learning algorithm that assigns classification values ​​to images, such as a convolutional neural network, may be employed.

[0011] In the system configuration of the present invention described above, the video extraction means sequentially extracts video of a predetermined length from a continuous video being transmitted, and the evaluation value output means sequentially obtains evaluation values ​​for the sequentially extracted video. As a result, the quality of the video can be evaluated sequentially in real time during video transmission or playback.

[0012] Incidentally, the quality of the video being evaluated as described above can change due to conditions other than those contained in the video during imaging, generation, or transmission, such as the communication status when transmitting the video signal train, changes in conditions that cause changes in the image within the video (camera movement speed, acceleration, etc.), and conditions such as the time of imaging, weather, and illumination. In this regard, it is expected that the video quality evaluation can be achieved with greater accuracy if the evaluation value output means is trained to output an evaluation value of the video quality using various conditions not contained in the video that can affect the video quality as described above, as input to the classifier which serves as the evaluation value output means.

[0013] Thus, in the configuration of the present invention described above, when the extracted video and non-video conditions are input to the evaluation value output means, and the evaluation value output means is trained by a machine learning algorithm to output an evaluation value of the video quality by a person who has viewed the training video acquired under the non-video conditions, when the training video of a predetermined length and non-training video conditions at the time of acquisition of the training video are input to the evaluation value output means. Here, "non-video conditions" may be, as described above, the communication status when transmitting the signal train, changes in circumstances that cause changes in the image within the video, such as the camera's movement speed and acceleration, time, weather, and illumination. For example, if the transmitted video is a video taken by an in-vehicle camera, the non-video conditions may be arbitrarily selected from the following: vehicle position (GPS information), vehicle speed, acceleration, yaw rate, steering angle, time, communication quality, video information (information on whether each frame of the video is a key frame (P-frame / I-frame, etc.)), weather information, wind speed, traffic information, etc. In the learning process, various videos acquired at the transmission destination under various external video conditions, along with the external video conditions at the time of acquisition of each video at the transmission destination, are used as input data. Training data is prepared with evaluation values, which are scores obtained by visually evaluating each video, as the correct values. Using this training data, the settings of the evaluation value output means (parameters for calculation, etc.) are adjusted to output the correct values ​​corresponding to the input data.

[0014] As already mentioned, the system of the present invention described above can be used, for example, to evaluate the quality of video in real time during video transmission when transmitting video between a moving object such as a vehicle and a server on a network. More specifically, in the case of a system that transmits video captured by an in-vehicle camera to a server, a predetermined length of video may be successively extracted from the video actually being transmitted, and an evaluation value for the extracted video may be output sequentially in real time. With such a configuration, based on the sequentially output evaluation values, various setting conditions from the capture by the in-vehicle camera to the transmission of the video to the server, such as appropriately controlling the bitrate or outputting an alert when deterioration of video quality is detected, can be executed in real time during video transmission. Similarly, in the case of a system that transmits arbitrary video, such as arbitrary video content, from a server to a vehicle, various setting conditions, such as appropriately controlling the bitrate or outputting an alert when deterioration of video quality is detected, can be executed in real time during video transmission. The above processing may be applied when transmitting video between any two points, and such transmission may be via any network. Therefore, in the system of the present invention, the continuous video during transmission from which the video extraction means extracts a predetermined length of video may be video transmitted from the destination via the network.

[0015] The evaluation of video quality in the system of the present invention described above may be a no-reference method, as already mentioned, that is, a method in which the video quality evaluation value is determined by referring only to the video to be evaluated, or a full-reference method, in which the video is determined by referring to the video to be evaluated and a video that serves as the basis for evaluation. In either case, external video conditions may also be used as input to the evaluation value output means. [Effects of the Invention]

[0016] Thus, according to the above configuration, it becomes possible to evaluate the video quality in real time during video transmission. By using the system of the present invention, it is also possible to evaluate the received video quality, display a graph of the video quality, a video degradation alert, and perform bitrate control in real time during video transmission. In addition, compared to video quality evaluation in file units that is not in real time, it is expected that a higher-quality service can be provided by a quality evaluation that is closer to human subjectivity and has high immediacy.

[0017] Other objects and advantages of the present invention will become apparent from the following description of the preferred embodiments of the present invention.

Brief Description of the Drawings

[0018] [Figure 1] FIG. 1 is a diagram for explaining an outline of an example of a system in which a video to which a video quality evaluation system according to the present embodiment is applied is transmitted. [Figure 2] FIGS. 2(A) to (D) are diagrams showing in block diagram form the configuration of a system in which a video to which a video quality evaluation system according to the present embodiment is incorporated is transmitted. (A) shows a case where one aspect of the video quality evaluation system according to the present embodiment is incorporated in a system in which video is transmitted from a mobile body to a server. (B) shows a case where another aspect of the video quality evaluation system according to the present embodiment is incorporated in a system in which video is transmitted from a mobile body to a server. (C) shows a case where one aspect of the video quality evaluation system according to the present embodiment is incorporated in a system in which video is transmitted from a server to a mobile body. (D) shows a case where another aspect of the video quality evaluation system according to the present embodiment is incorporated in a system in which video is transmitted from a server to a mobile body. [Figure 3]Figs. 3(A) to (B) are diagrams showing the configurations of some aspects of the video quality evaluation system according to the present embodiment in the form of block diagrams. (A) shows the case where an evaluation value is calculated only from a video in a non-reference method. (B) shows the case where an evaluation value is calculated from a video and external video conditions in a non-reference method. (C) shows the case where an evaluation value is calculated only from a video in a full-reference method. (D) shows the case where an evaluation value is calculated from a video and external video conditions in a full-reference method. [Figure 4] Figs. 4(A) to (C) are diagrams for explaining the aspect of sequentially cutting out a video of a predetermined length from consecutive videos during transmission in the video cutting section in the video quality evaluation system according to the present embodiment. (A) shows the case where a video of a predetermined length is cut out without duplication. (B) shows the case where the length of the video to be cut out is extended while sequentially duplicating the cut-out videos. (C) shows the case where a video signal sequence of a predetermined length is cut out with duplication. [Figure 5] Figs. 5(A) and (B) show examples of evaluation values of the quality of a video determined over time along with the transmission of the video in the video quality evaluation system according to the present embodiment. (A) is an example according to the non-reference method, and (B) is an example according to the full-reference method.

Description of Reference Numerals

[0019] 10... Moving body such as a vehicle, 12... Camera, 12a... Encoding section, 13... Communicator, 14... Monitor, 14a... Decoding section, 14b... Video quality evaluation section, 15... Vehicle state sensor, 20... Server room, 21... Server, 21a... Server-side communicator, 21b... Decoding section, 21c... Video quality evaluation section, 22... Server-side monitor, 23... Video supply section, 23a... Encoding section, 30a, b... Communication lines, 31... Video division section (video cutting means), 31o... Original video division section, 32... Evaluation value calculation section (evaluation value output means), NW... Communication network (Internet, etc.)

Best Mode for Carrying Out the Invention

[0020] The present invention will be described in detail below with reference to the attached figures, with reference to several preferred embodiments. In the figures, the same reference numerals indicate the same parts.

[0021] Configuration of the video transmission system The video quality evaluation system according to this embodiment may be applied to a system that transmits video between any two points (video transmission system). The video transmission system may be configured to transmit a signal sequence (video signal sequence) that constitutes video from a camera, video generation device, video playback device, or any other device that outputs video, to a display that shows video or a recording device or recorder that records video, via any type of wired or wireless communication line. Specifically, the configuration of the video transmission system may be from a camera to a display, or, as schematically depicted in Figure 1, it may be a system that transmits and receives video between a mobile body 10 such as a vehicle and a server 21 installed in a server room 20 located away from it. For example, in the system illustrated in Figure 1, in one embodiment, video from a camera 12 mounted to capture images around the mobile body 10 may be converted into a signal sequence and transmitted from a communication device 13 to a server 21 in the server room 20 via a wireless line 30a, and displayed on a monitor 22. With this configuration, it becomes possible for operator P to control the mobile unit 10 using a remote control system (not shown) while referring to the video on monitor 22, or to monitor the surroundings of the mobile unit in autonomous driving mode. In addition, in the configuration of Figure 1, arbitrary content video may be converted into a signal sequence from server 21 and transmitted to communication device 13 via wireless line 30b, and displayed on monitor 14 mounted on the mobile unit. This makes it possible for the occupants of the mobile unit to acquire various information or enjoy various video content.

[0022] Specific configuration of a video transmission system incorporating a video quality evaluation system The video quality evaluation system of this embodiment is a system that, when video is transmitted between two points in a video transmission system, for example between a mobile device 10 and a server 21 as shown in Figure 1 (i.e., when a signal sequence representing the brightness of pixels in multiple images is sequentially transmitted and received over a line), extracts the video at the video receiving side into predetermined length segments and outputs an evaluation value of the quality of the extracted video segments sequentially or in real time.

[0023] In one embodiment of a video transmission system incorporating the video quality evaluation system of this embodiment as described above, as illustrated in Figure 2(A), first, video s1 is captured or generated by a video generating device such as a camera 12 on the transmitting side of the video to be transmitted (for example, a mobile device 10) and sent to the encoding unit 12a, where it is converted into an encoded signal sequence s2. This signal sequence s2 is converted into a transmission signal sequence s3 by a communicator 13 and transmitted to a communicator 21a on the receiving side of the video (for example, a server 21) via a communication line, which may be a communication network NW. At the receiving side of the video, the received signal sequence s3 is converted into a signal sequence s4 for video restoration processing and input to a decoding unit 21b, where it is restored into video s5 and may be displayed on a monitor 22. Furthermore, the restored video s5 is input to the video quality evaluation unit 21c according to this embodiment, where, as described later, the video quality evaluation value can be calculated sequentially or in real time, that is, while the video is being transmitted. This makes it possible to sequentially evaluate the restored video after transmission. The real-time evaluation value calculated by the video quality evaluation unit 21c may be used to control the processing of the video being transmitted from the video transmission side to the encoding unit 12a, for example, to control the bitrate during encoding. In this case, a control command c1, which is configured based on the evaluation value from the video quality evaluation unit 21c, is converted into a transmission signal c2 by the communication device 21a and transmitted to the communication device 13 of the video transmission side 10 via a communication line, which may be a communication network NW, where it is converted into a control signal c3 and provided to the encoding unit 12a.

[0024] By the way, in the video quality evaluation system of this embodiment, which is incorporated into the video transmission system as described above, as already mentioned, when evaluating video quality, conditions other than the video itself during video shooting or generation or video transmission may be taken into consideration. In such a configuration, in addition to the configuration in Figure 2(A), as shown in Figure 2(B), various conditions of the video transmission side 10, for example, in the case of a moving object 10, state quantities detected by various sensors 15 on the moving object, such as the position of the moving object (GPS information), speed, acceleration, yaw rate, steering angle, time, communication quality, and video information (information on whether each frame of the video is a key frame (P-frame / I-frame, etc.)), may be input to the communicator 13 as non-video conditions ix, sent together with the video signal train s3 to the communicator 21a on the video receiving side 20, and input to the video quality evaluation unit 21c. Furthermore, various conditions obtained from the communication network NW, such as weather information, wind speed, and traffic information, may also be used for video quality evaluation. In this case, such information obtained from the communication network NW may be sent as an external video condition is to the communication device 21a on the video receiving side 20 and input to the video quality evaluation unit 21c. This configuration is expected to improve the accuracy of video quality evaluation.

[0025] Furthermore, the direction of video transmission between the video transmitting side, such as the mobile device 10, and the video receiving side, such as the server room 20, may be reversed. Specifically, as shown in Figure 2(C), the server room 20 acts as the video transmitting side, and the video s1 output from the video generator or video player 23 located there is converted into an encoded signal sequence s2 by the encoding unit 23a, converted into a transmission signal sequence s3 by the communicator 21a, and transmitted to the communicator 13 of the video receiving side (for example, the mobile device 10) via a communication line, which may be a communication network NW. The communication device 13 then decodes the transmission signal sequence s3 into a signal sequence s4 for video restoration processing, converts it into video s5 by the decoding unit 14a, and displays it on the monitor 14. This is also input to the video quality evaluation unit 14b according to this embodiment, where the video quality evaluation value can be calculated sequentially or in real time, that is, while the video is being transmitted. In this case as well, the real-time evaluation value calculated by the video quality evaluation unit 14b may be used to control the processing of the video being transmitted from the video transmitter to the encoding unit 23a, for example, to control the bitrate during encoding. In this case, a control command c1 configured based on the evaluation value from the video quality evaluation unit 14b is converted into a transmission signal c2 by the communicator 13 and transmitted to the communicator 21a of the video transmitter 20 via a communication line, which may be a communication network NW, where it is converted into a control signal c3 and provided to the encoding unit 23a.

[0026] Furthermore, as illustrated in Figure 2(D), when video is transmitted from the server room 20 to the mobile device 10 and video quality evaluation is performed simultaneously, non-video conditions during video generation and video signal transmission may also be considered. This is expected to improve the accuracy of video quality evaluation. In such a configuration, in addition to the configuration in Figure 2(C), various conditions obtained from the communication network NW, such as weather information, wind speed, and traffic information, may also be used. In this case, such information obtained from the communication network NW may be sent as non-video conditions is to the communication device 13 on the video receiving side 10 and input to the video quality evaluation unit 14b.

[0027] The evaluation value calculated by the video quality evaluation unit 21c or 14b in the above system may be used for control in the conversion and transmission of the video signal train on the video transmission side as described above, or the system may be configured to issue an alert when the evaluation value falls below a predetermined threshold.

[0028] Configuration and operation of the video quality evaluation unit As previously mentioned, the video quality evaluation units 21c and 14b in this embodiment are configured to sequentially extract video of a predetermined duration from the video transmitted from the video transmitter in the form of a signal train, and to calculate an evaluation value for the extracted video. The calculation of the evaluation value for the extracted video may be achieved by a classifier configured using training data according to a machine learning algorithm.

[0029] More specifically, in the first embodiment, as shown in Figure 3(A), the video quality evaluation units 21c and 14b consist of a video splitting unit 31 that sequentially extracts video m(i) of a predetermined time length from continuously transmitted video M(t), and an evaluation value calculation unit 32 that sequentially calculates an evaluation value R(i) for video m(i). This configuration of the first embodiment is applied to the system configurations shown in Figures 2(A) and (C).

[0030] More specifically, the video extraction in the video splitting unit 31 may be performed in one of the following ways. In the first way, as shown in Figure 4(A), video segments mi of a predetermined duration may be sequentially extracted from the transmitted video M without overlap. The predetermined duration may be set to any duration, for example, 4 seconds. The evaluation value Ri from the evaluation value calculation unit 32 is calculated each time a video of a predetermined duration is collected and input to the evaluation value calculation unit 32, so in the case of Figure 4(A), it will be output at predetermined duration intervals. In the second way, as shown in Figure 4(B), the video may be extracted so that only the portion of the video that has elapsed up to the end of the video segment mi of a predetermined duration is extracted and input to the evaluation value calculation unit 32, so that the evaluation value Ri from the evaluation value calculation unit 32 is output at shorter time intervals than the predetermined duration in Figure 4(A), and the evaluation value Rij can be output at shorter intervals during the elapsed duration of the predetermined duration. Furthermore, in the third embodiment, as shown in Figure 4(C), video segments mi of a predetermined duration are sequentially extracted from the transmitted video M, and the evaluation value Ri from the evaluation value calculation unit 32 may be calculated at the end of each predetermined duration of video mi. In this case as well, as shown in the figure, the evaluation value Ri will be calculated at intervals shorter than the predetermined duration.

[0031] As already mentioned, the evaluation value calculation unit 32 is configured to output an evaluation value that scores the level of video quality as perceived by a person when they view the video being evaluated, according to a machine learning algorithm. In this regard, the evaluation value calculation unit 32 of the first embodiment shown in Figure 3(A) is configured to calculate an evaluation value from video according to the procedure of the non-reference method of video quality evaluation (VQA) (Non-Patent Literature 1). In the case of the non-reference method of VQA, in the preparation of the classifier which will be the evaluation value calculation unit, generally speaking, first, various videos and the scores (evaluation values) of the quality of each video as evaluated by people who have viewed each video are prepared as input data and ground truth data in the training data, respectively. Then, in the training of the classifier, when the prepared training input data is input, the calculation parameters within the classifier are adjusted according to the machine learning algorithm so that the corresponding evaluation value (ground truth data) is output. As the machine learning algorithm, deep learning algorithms such as convolutional neural networks may be used. Once the classifier is adjusted, an arbitrary video m(i) is input to the classifier, and its evaluation value R(i) is calculated. In the configuration of this embodiment, as described above, each video m(i) extracted by the video splitting unit 31 is input to the evaluation value calculation unit 32, and an evaluation value R(i) is calculated for each such video m, thereby sequentially obtaining an evaluation value of the video quality during the transmission of the video M.

[0032] In the second embodiment of the video quality evaluation units 21c and 14b, as shown in Figure 3(B), the video external conditions ix and is, as previously described, are further input to the evaluation value calculation unit 32 along with the video extracted by the video splitting unit 31, and the evaluation value R(i) is calculated based on the video and the video external conditions according to the procedure of the non-reference VQA method. This configuration of the second embodiment is applied to the system configuration shown in Figures 2(B) and (D). In this embodiment, various videos acquired under various video external conditions and the quality evaluation scores (evaluation values) of each video by viewers are used as input data and ground truth data in the training data, respectively, and machine learning of the classifier is performed. In calculating the evaluation value, when an arbitrary video m(i) and the acquired video external conditions ix and is are input to the classifier, the evaluation value R(i) is calculated. As described above, by using a configuration that calculates evaluation values ​​by taking into account not only the video itself but also external conditions, it is expected that a more accurate evaluation of video quality can be achieved.

[0033] Furthermore, the evaluation value calculation unit 32 in the video quality evaluation units 21c and 14b may be configured to calculate the evaluation value of the extracted video according to the procedure of the full-reference VQA (Non-Patent Literature 2). In one embodiment of this, as shown in Figure 3(C), a reference video (reference video) Mo(t) for the video to be evaluated is prepared in the video quality evaluation unit. Such reference video Mo(t) may be the original video of the transmitted video M(t), that is, the video before transmission. The reference video Mo(t) is then sequentially extracted into video mo(i) of a predetermined time length by the original video splitting unit 31o, similar to the video M(t) to be evaluated, and this video mo(i) is input to the evaluation value calculation unit 32 together with the video m(i) extracted by the video splitting unit 31. In the learning process of the evaluation value calculation unit 32 in this case, first, various transmitted video footages are prepared as training data, along with pre-transmission video footage as a reference video. Then, scores (evaluation values) of the quality evaluation of each video footage by people who have viewed the transmitted and pre-transmission video footage are prepared. The classifier, which is the evaluation value calculation unit 32, adjusts its calculation parameters according to a machine learning algorithm so that when the transmitted and pre-transmission video footage, which are the input data for learning, is input, it outputs the corresponding evaluation value (correct answer data). Thus, when arbitrary transmitted video footage m(i) and pre-transmission video footage mo(i) are input to the adjusted evaluation value calculation unit 32, the evaluation value R(i) is calculated. This configuration can be applied to the system configurations shown in Figures 2(A) and (C).

[0034] Furthermore, even in a configuration that calculates the video evaluation value according to the procedure of the full-reference VQA method, the evaluation value may be calculated by referring to the external video conditions ix and is. Accordingly, as shown in Figure 3(D), the external video conditions ix and is may be input to the evaluation value calculation unit 32. In this case, the classifier's machine learning is performed using various post-transmission and pre-transmission videos acquired under various external video conditions, and the quality evaluation scores (evaluation values) of each video by viewers, respectively, as input data and ground truth data in the training data. In calculating the evaluation value, the evaluation value is calculated by inputting an arbitrary post-transmission video m(i), a pre-transmission video mo(i), and the acquired external video conditions ix and is for that video into the classifier.

[0035] In configurations where video quality is evaluated using a full reference method, both pre-transmission and post-transmission video are required to calculate the evaluation value. Therefore, it is not possible to calculate the evaluation value in real time during video transmission between two separated points. Such configurations are used in situations where both pre-transmission and post-transmission video are available at the same location, for example, in environments where direct video from a camera or image generator and video after it has been passed through a transmission system can be viewed simultaneously.

[0036] Examples of sequentially calculated evaluation values According to the video quality evaluation system of this embodiment described above, as illustrated in Figures 5(A) and (B), the evaluation value R is calculated sequentially as time T elapses during video transmission or playback. The system may also be configured to issue a quality degradation alert when, for example, the quality deteriorates and the evaluation value R falls below a predetermined threshold th that is set as appropriate. In Figure 5(A), ▼ represents the case where the video is extracted without overlap, as in Figure 4(A), and ○ represents the case where the video is extracted with overlap, as in Figure 4(C), and the evaluation value is calculated at intervals finer than the extraction length. As can be seen from the figure, it is clear that the latter method can capture fluctuations in the evaluation value more precisely, and thus provide a more accurate evaluation of video quality.

[0037] Thus, according to this embodiment, it is possible to obtain real-time video quality evaluation values ​​when transmitting video between two points. Using this system, it is also possible to evaluate the quality of the received video during transmission and, based on the results, to display video quality graphs, video degradation alerts, and control the bitrate in real time. Furthermore, since the evaluation in this embodiment is an evaluation of video quality, rather than communication quality, it has the advantage of providing an evaluation that is closer to human subjectivity in real time or sequentially.

[0038] While the above description is made in relation to embodiments of the present invention, many modifications and changes are readily possible for those skilled in the art, and it will be clear that the present invention is not limited to the embodiments illustrated above, but can be applied to various devices without departing from the concept of the present invention.

Claims

1. It is a system for evaluating video quality, A video extraction means configured to sequentially extract video of a predetermined length from a continuous video being transmitted, An evaluation value output means configured to sequentially input the extracted video and sequentially output an evaluation value of the video quality of the input video, wherein when a training video of a predetermined length is input, the evaluation value output means is trained by a machine learning algorithm to output an evaluation value of the video quality by a person who has watched the training video. A system that includes this.

2. A system according to claim 1, wherein the evaluation value output means receives external video conditions along with the extracted video, and the evaluation value output means is trained by a machine learning algorithm to output an evaluation value of video quality by a person who has viewed the training video acquired under the external video conditions, when it receives external video conditions at the time the training video was acquired along with a predetermined length of training video.

3. The system according to claim 1, wherein the continuous video being transmitted, from which the video extraction means extracts a predetermined length of video, is video transmitted from a destination via a network.

4. A system according to claims 1 to 3, wherein the evaluation value output means is configured to output an evaluation value corresponding to the extracted video in a no-reference manner.

5. A system according to claims 1 to 3, wherein the evaluation value output means is configured to output an evaluation value corresponding to the extracted video in a full reference manner.

Citation Information

Patent Citations

  • Video quality objective evaluation device and program

    JP2012119772A

  • Real-time video communication quality evaluation method and system

    JP2016213784A

  • Method for evaluating video quality based on non-reference video

    JP2023152957A

  • Video quality estimation device, video quality estimation method, and program

    WO2023233631A1