Video quality assessment method and apparatus
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2021-06-15
- Publication Date
- 2026-06-02
Smart Images

Figure CN115606174B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to a video quality assessment method and apparatus, and more specifically, to a video quality assessment method and apparatus for assessing the quality of frames in a video based on the fuzziness level of the frame. Background Technology
[0002] Distortion occurs in video images during the processes of generation, compression, storage, transmission, and reproduction. Distorted images must be reproduced within the limits of human perception. Therefore, it is necessary to quantify and assess the quality before an image is reproduced in order to understand how the distortion affects the quality perceived by humans.
[0003] Image quality can be assessed using both subjective and objective quality assessment methods. Subjective quality assessment methods involve evaluators directly viewing the video and evaluating its quality, accurately reflecting human quality perception characteristics. However, subjective quality assessment methods have drawbacks: evaluation values differ for each individual, require significant time and cost, and are difficult to implement in real-time.
[0004] Objective quality assessment is an algorithm that quantifies the quality perceived by the human optic nerve and uses this algorithm to assess the degree of quality degradation of compressed images.
[0005] Objective quality assessment methods include full-reference quality assessment methods that compare a reference image with a distorted image, simplified reference quality assessment methods that use partial information about the reference image (e.g., a watermark or auxiliary channel) instead of the reference image itself to perform quality assessment, and no-reference quality assessment methods that use only the distorted image without using any information about the reference image to perform quality assessment.
[0006] The no-reference quality assessment method does not require reference image information and can therefore be used in any application that requires quality assessment. Summary of the Invention
[0007] Technical solution
[0008] A video quality assessment method and apparatus are provided for identifying whether a frame is a fully blurred frame or a partially blurred frame based on the blur level of the frames included in the video, and for evaluating the quality of the frame in a different manner in each case. Attached Figure Description
[0009] The above and other aspects, features and advantages of specific embodiments of the present disclosure will become more apparent from the following description taken in conjunction with the accompanying drawings, in which:
[0010] Figure 1This is a diagram illustrating a video quality assessment device that generates a quality score for a video image according to an embodiment, and an image display device that outputs an image with processed quality on a screen.
[0011] Figure 2 This is a diagram illustrating the power spectrum and spectral envelope obtained from the input frame according to an embodiment;
[0012] Figure 3 This is a diagram illustrating the power spectrum and spectral envelope from the input frame according to an embodiment;
[0013] Figure 4 This is a diagram illustrating a fully blurred frame and a partially blurred frame according to an embodiment;
[0014] Figure 5 This is a diagram illustrating importance information for each sub-region included in a frame according to an embodiment;
[0015] Figure 6 This is a diagram illustrating information obtained for each sub-region included in a frame according to an embodiment;
[0016] Figure 7 This is a block diagram of the internal structure of a computing device according to an embodiment;
[0017] Figure 8 This is a block diagram illustrating the configuration of a processor according to an embodiment;
[0018] Figure 9 This is a block diagram illustrating the configuration of a processor according to an embodiment;
[0019] Figure 10 This is a block diagram of the internal structure of the processor according to an embodiment;
[0020] Figure 11 This is a block diagram of the internal structure of a computing device according to another embodiment;
[0021] Figure 12 This is a block diagram of the internal structure of a computing device according to an embodiment;
[0022] Figure 13 This is a block diagram of the internal structure of an image display device according to an embodiment;
[0023] Figure 14 This is a block diagram of the internal structure of an image display device according to an embodiment;
[0024] Figure 15 This is a flowchart illustrating a frame recognition method according to an embodiment;
[0025] Figure 16 This is a flowchart illustrating a method for obtaining a quality score according to an embodiment;
[0026] Figure 17 This is a flowchart illustrating a method for obtaining a model-based quality score for partially blurred frames according to an embodiment;
[0027] Figure 18 This is a flowchart illustrating a method for obtaining a final quality score according to an embodiment; and
[0028] Figure 19 This is a flowchart illustrating a method for evaluating video quality according to another embodiment. Detailed Implementation
[0029] This application is based on and claims priority to Korean Patent Application No. 10-2020-0076767, filed on June 23, 2020, with the Korean Intellectual Property Office, the entire disclosure of which is incorporated herein by reference.
[0030] A video quality assessment method and apparatus are provided for obtaining different features from fully blurred frames and partially blurred frames and using different features to evaluate the quality of each frame.
[0031] A video quality assessment method and apparatus are provided, which can assess the quality of the frame by obtaining importance information and spectral envelope distribution characteristics of each sub-region of a frame included in the video, and generating and using a weight matrix from the importance information and spectral envelope distribution characteristics of each sub-region, taking into account the characteristics of each sub-region.
[0032] A video quality assessment method and apparatus are provided for evaluating the quality of frames by using a neural network that has been trained for human subjective evaluation.
[0033] According to one aspect of this disclosure, a video quality assessment method includes: receiving frames of a video; identifying whether the frame is a fully blurred frame or a partially blurred frame based on the blur level of the frame; obtaining an analysis-based quality score for the fully blurred frame in response to the frame being the fully blurred frame; obtaining a model-based quality score for the partially blurred frame in response to the frame being the partially blurred frame; and processing the video based on at least one of the analysis-based quality score or the model-based quality score to obtain a processed video.
[0034] The video quality assessment method may further include: obtaining the spectral envelope for each of the plurality of sub-regions included in the frame; and estimating the blur level for each of the plurality of sub-regions based on the spectral envelope. The step of identifying whether the frame is a fully blurred frame or a partially blurred frame may include: identifying the number of plurality of sub-regions whose blur level estimated based on the spectral envelope exceeds a threshold.
[0035] The step of obtaining the spectral envelope for each of the plurality of sub-regions may include: obtaining the signal in the frequency domain for the corresponding sub-region; obtaining the power spectrum of the signal in the frequency domain; and obtaining the spectral envelope based on the power spectrum.
[0036] The step of obtaining the analysis-based quality score may include: obtaining spectral envelope distribution features based on the spectral envelope of each of the plurality of sub-regions included in the fully blurred frame; and obtaining an analysis-based quality score for the fully blurred frame by analyzing the spectral envelope distribution features for each of the plurality of sub-regions.
[0037] The video quality assessment method may further include: obtaining spectral envelope distribution features based on the spectral envelope of each sub-region among the plurality of sub-regions included in the partially blurred frame; obtaining importance information of each sub-region among the plurality of sub-regions included in the partially blurred frame using a first neural network; and generating a weight matrix indicating the weight values of each sub-region among the plurality of sub-regions based on the spectral envelope distribution features and the importance information obtained for each sub-region among the plurality of sub-regions included in the partially blurred frame. The step of obtaining the model-based quality score may include: obtaining a model-based quality score for the partially blurred frame based on the weight matrix and the partially blurred frame.
[0038] The step of obtaining the importance information may include: for each of the plurality of sub-regions, obtaining one or more pieces of importance information related to factors that can affect the quality score by using a first neural network, and the importance information may include at least one of information on whether the object is a foreground object or a background object, semantic information, location information, or content information.
[0039] The plurality of sub-regions may include a first sub-region and a second sub-region adjacent to the first sub-region, and the video quality assessment method may further include: correcting the first spectral envelope distribution features of the first sub-region by using the second spectral envelope distribution features of the second sub-region; correcting the first importance information of the first sub-region by using the second importance information of the second sub-region; and obtaining a corrected weight value for the first sub-region by using the corrected first spectral envelope distribution features and the corrected first importance information.
[0040] The step of obtaining the model-based quality score can be performed using a second neural network trained on the correlation between the feature vector and the mean opinion score (MOS).
[0041] The step of obtaining the model-based quality score may include: extracting features from the partially blurred frames using a second neural network; and obtaining a quality score for the partially blurred frames based on the features and the weight matrix, wherein the features may include at least one of blur-related features, motion-related features, content-related features, depth features, statistical features, perceptual features, spatial features, or correction domain features.
[0042] The video quality assessment method may further include: accumulating the analysis-based quality score and the model-based quality score for multiple frames over a period of time to obtain time series data; and smoothing the time series data to obtain a final quality score.
[0043] The step of smoothing the time series data to obtain a final quality score can be performed using a third neural network model, and the third neural network model may include Long Short-Term Memory (LSTM).
[0044] The steps for processing the video may include: processing the frame based on the final quality score, and the steps for processing the video may be performed by at least one of the following operations: processing the video according to a quality processing model selected based on the final quality score; identifying the number of times the quality processing model is applied based on the final quality score and processing the video by repeatedly applying the quality processing model to the frame according to the number of times; identifying a filter based on the final quality score and processing the video by applying the filter to the frame; or processing the video by using a neural network with hyperparameter values corrected according to the final quality score.
[0045] According to one aspect of this disclosure, a video quality assessment method includes: receiving frames of a video; identifying whether the frame is a fully blurred frame or a partially blurred frame based on the blur level of the frame; obtaining a spectral envelope distribution feature of the fully blurred frame in response to the frame being the fully blurred frame; obtaining a model-based feature of the partially blurred frame in response to the frame being the partially blurred frame; obtaining a final quality score for the frame based on at least one of the spectral envelope distribution feature or the model-based feature; and processing the video based on the final quality score to obtain a processed video.
[0046] The video quality assessment method may further include: obtaining importance information for each sub-region among multiple sub-regions included in the frame; obtaining spectral envelope distribution features for each sub-region among the multiple sub-regions included in the frame; and generating a weight matrix for the frame by obtaining weights for each sub-region among the multiple sub-regions based on the importance information and spectral envelope distribution features. The step of obtaining a final quality score for the frame may include: obtaining a final quality score for the frame by using the spectral envelope distribution features of each sub-region among the multiple sub-regions of the fully blurred frame, model-based features of the partially blurred frame, and the weight matrix.
[0047] The step of obtaining the model-based features may include: extracting at least one of the following from the partially blurred frames using at least one neural network: blur-related features, motion-related features, content-related features, depth features, statistical features, perceptual features, spatial features, or correction domain features.
[0048] According to one aspect of this disclosure, a video quality assessment device includes: a memory storing one or more instructions; and a processor configured to execute the one or more instructions stored in the memory to: identify whether a frame is a fully blurred frame or a partially blurred frame based on the blur level of a frame included in the video; in response to the frame being the fully blurred frame, obtain an analysis-based quality score for the fully blurred frame; in response to the frame being the partially blurred frame, obtain a model-based quality score for the partially blurred frame; and process the video based on at least one of the analysis-based quality score or the model-based quality score to obtain a processed video.
[0049] The processor may also be configured to execute one or more instructions to perform the following operations: obtain spectral envelope distribution features from the spectral envelope of each sub-region among a plurality of sub-regions included in the fully blurred frame, and obtain an analysis-based quality score for the fully blurred frame by analyzing the spectral envelope distribution features for each sub-region among the plurality of sub-regions.
[0050] The processor may also be configured to execute the one or more instructions to perform the following operations: obtaining spectral envelope distribution features based on the spectral envelope of each sub-region among the plurality of sub-regions included in the partially blurred frame; obtaining importance information for each sub-region among the plurality of sub-regions included in the partially blurred frame; generating a weight matrix indicating weight values for each sub-region based on the spectral envelope distribution features and the importance information obtained for each sub-region among the plurality of sub-regions included in the partially blurred frame; and obtaining a model-based quality score for the partially blurred frame based on the weight matrix and the partially blurred frame.
[0051] The video quality assessment device may also include an output interface, and the processor is further configured to control the output interface to provide the processed video to a display panel.
[0052] The output interface may include the display panel, and the output interface is configured to provide a wired connection between the video quality assessment device and the display panel.
[0053] Embodiments will now be described with reference to the accompanying drawings. However, this disclosure may be implemented in many different forms and should not be construed as limited to the examples set forth herein.
[0054] Although these general terms were chosen to describe embodiments for their function, they may vary depending on the intent of those skilled in the art, precedents, the emergence of new technologies, etc. Therefore, these terms must be defined based on their meaning and the entirety of the specification, rather than simply by stating the terms.
[0055] The terminology used in this specification is used to describe particular embodiments and is not intended to limit the scope of this disclosure.
[0056] Throughout the specification, when an element is referred to as “connected” or “coupled” to another element, it may be directly connected or coupled to said other element, or it may be electrically connected or coupled to said other element through an intermediate element inserted therebetween.
[0057] Throughout the disclosure, expressions such as “at least one of a, b, or c” indicate only a, only b, only c, both a and b, both a and c, both b and c, all of a, b, and c, or variations thereof.
[0058] As used herein, the terms “first” or “first” and “second” or “second” may be used with respect to the corresponding component regardless of importance or order, and are used to distinguish one component from another without limiting the components. The terms “a”, “an”, and “the”, and similar indications, should be interpreted to cover both the singular and plural. Furthermore, the steps of all methods described herein may be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by the context. The embodiments are not limited to the described order of operations.
[0059] Therefore, the phrase "according to an embodiment" does not necessarily refer to the same embodiment.
[0060] Embodiments may be described in terms of functional block components and various processing steps. Some or all of these functional blocks may be implemented by any number of hardware and / or software components configured to perform specified functions. For example, functional blocks may be implemented by one or more microprocessors or by circuit components for a specific function. Furthermore, functional blocks may be implemented using any programming or scripting language, for example. Functional blocks may be implemented in algorithms that execute on one or more processors. In addition, the embodiments described herein may employ any number of techniques for electronic construction, signal processing and / or control, data processing, etc. The terms “mechanism,” “element,” “means,” and “configuration” are used broadly and are not limited to mechanical or physical embodiments.
[0061] Furthermore, the connecting lines or connectors between components shown in the various figures are intended to represent exemplary functional relationships and / or physical or logical couplings between components. Connections between components can be represented by numerous alternative or additional functional relationships, physical connections, or logical connections in an actual device.
[0062] The terms “unit,” “device,” “component,” and “module” refer to a unit that performs at least one function or operation and can be implemented as hardware (i.e., a processor or circuit), software, or a combination of hardware and software.
[0063] As used herein, the term "user" refers to a person who controls the functionality or operation of an image display device by using that device. Examples of users may include viewers, administrators, or installation engineers.
[0064] The embodiments will now be described more fully with reference to the accompanying drawings.
[0065] Figure 1 This is a diagram illustrating a video quality assessment device 130 that generates a quality score for a video image according to an embodiment, and an image display device 110 that outputs an image with processed quality on a screen.
[0066] Reference Figure 1The image display device 110 can communicate with the video quality assessment device 130 via the communication network 120.
[0067] The image display device 110 can be an electronic device capable of processing and outputting images. The image display device 110 can be fixed or mobile, and can be a digital television (TV) capable of receiving digital broadcasts, but is not limited thereto, and can be implemented as various types of electronic devices including displays.
[0068] The image display device 110 may include at least one of a desktop personal computer (PC), smartphone, tablet PC, mobile phone, video phone, e-book reader, laptop PC, netbook computer, digital camera, personal digital assistant (PDA), portable multimedia player (PMP), camcorder, navigation device, wearable device, smartwatch, home network system, security system, or medical device.
[0069] The image display device 110 can be implemented not only as a flat panel display device, but also as a curved display device as a screen with curvature, or a flexible display device capable of adjusting curvature. The output resolution of the image display device 110 may include, for example, high definition (HD), full HD, ultra HD, or a resolution clearer than ultra HD.
[0070] Image display device 110 can output video. The video can be configured as multiple frames. The video may include items such as TV programs provided by a content provider or various movies or dramas provided through a video-on-demand (VOD) service. A content provider may refer to a terrestrial broadcasting station or a cable broadcasting station, or an over-the-top (OTT) service provider or an Internet Protocol Television (IPTV) service provider that provides consumers with various content including video.
[0071] Video is captured, compressed, and sent to image display device 110, where it is reconstructed and output. Due to limited bandwidth and physical limitations of the means of capturing the video, information is lost, resulting in image distortion. Distorted video may have degraded quality. Image distortion may include blurring. For example, the edges of objects included in a frame may appear blurry due to loss.
[0072] In one embodiment, the image display device 110 may send the input frame 140, which is included in the video, to the video quality assessment device 130 via the communication network 120 before outputting the input frame 140 to the screen.
[0073] In one embodiment, the video quality assessment device 130 can receive a video consisting of multiple frames from the image display device 110 via the communication network 120 and perform a quality assessment on the video.
[0074] In an embodiment, the video quality assessment device 130 may be a computing device for assessing the quality of a video.
[0075] The video quality assessment device 130 according to the embodiment can objectively assess the quality of a video by using a no-reference quality assessment method. For example, the video quality assessment device 130 can be provided as at least one hardware chip embedded in an electronic device, or it can be included in a server as a chip or electronic device. For example, the video quality assessment device 130 can be implemented as a software module in an electronic device or server.
[0076] The video quality assessment device 130 can receive video and estimate the blur level for each of the multiple frames included in the video.
[0077] exist Figure 1 In this process, the video quality assessment device 130 can divide the input frame 140 received from the image display device 110 into multiple sub-regions and obtain the signal in the frequency domain for each sub-region. The video quality assessment device 130 can obtain the spectral envelope from the signal in the frequency domain. The video quality assessment device 130 can estimate the blur level of each sub-region from the spectral envelope. This will refer to... Figure 2 and Figure 3 Provide a detailed description.
[0078] The video quality assessment device 130 can identify frames as unblurred, fully blurred, or partially blurred based on the blur level.
[0079] A fully blurred frame indicates a frame in which the blur level equal to or greater than a specific reference value is included in a specific region of the entire frame or a larger region. Conversely, a partially blurred frame indicates a frame in which the blur level equal to or greater than the specific reference value is included in a region smaller than the specific region of the frame.
[0080] In one embodiment, the video quality assessment device 130 can obtain the graphical distribution characteristics of the spectral envelope for that frame.
[0081] In an embodiment, when the input frame is determined to be a completely blurred frame, the video quality assessment device 130 can obtain the spectral envelope distribution characteristics of the sub-regions included in the completely blurred frame.
[0082] In an embodiment, when the input frame is determined to be a completely blurred frame, the video quality assessment device 130 can analyze the spectral envelope distribution characteristics of the sub-regions included in the completely blurred frame to obtain an analysis-based quality score for the completely blurred frame.
[0083] In an embodiment, when the input frame is determined to be a partially blurred frame, the video quality assessment device 130 can use artificial intelligence (AI) technology to obtain model-based features for the partially blurred frame.
[0084] In an embodiment, when the input frame is determined to be a partially blurred frame, the video quality assessment device 130 can use AI technology to obtain a model-based quality score for the partially blurred frame.
[0085] AI technology can be configured through machine learning (deep learning). AI technology can be implemented using neural networks that include algorithms or sets of algorithms. Neural networks can receive input data, perform operations for analysis and classification, and output the resulting data.
[0086] Neural networks can have multiple internal layers that perform operations. A neural network generates multiple data points representing features of the image from each of the multiple layers; these are called feature maps. In lower layers, the output is a feature map that is almost identical to the input image, and as the layers deepen, pixel-level information disappears, and the output becomes a detailed feature map that preserves the semantic information of the image.
[0087] However, when the blur level is high throughout the frame, the feature map obtained from the frame using a neural network may include noise. Furthermore, obtaining a model-based quality score using a neural network is more complex and computationally intensive than obtaining an analysis-based quality score.
[0088] Therefore, in this embodiment, the video quality assessment device 130 can evaluate the quality of a frame in different ways depending on the blur level of the input frame. That is, when the frame is a fully blurred frame, the video quality assessment device 130 can use a simple analysis-based method rather than a complex model-based method to evaluate the frame quality. Furthermore, when the frame is a partially blurred frame, the video quality assessment device 130 can use at least one neural network to evaluate the frame quality using a model-based method.
[0089] The video quality assessment device 130 can accumulate analysis-based quality scores and / or model-based quality scores obtained for frames (or multiple frames) (or a specific number of frames) within a specific time period to obtain a final quality score 150 for the video. The video quality assessment device 130 can send the final quality score 150 to the image display device 110 via the communication network 120.
[0090] In another embodiment, the video quality assessment device 130 can assess the quality of a frame from different features based on the blur level of the input frame. That is, when the frame is a fully blurred frame, the video quality assessment device 130 can obtain frequency features from the fully blurred frame, and when the frame is a partially blurred frame, model-based features can be obtained using a model utilizing at least one neural network. The video quality assessment device 130 can more accurately assess the quality of each frame based on the different features obtained according to the frame type.
[0091] The video quality assessment device 130 can accumulate frequency features and / or model-based features obtained for frames (or multiple frames) over a specific time period, input the frequency features and / or model-based features into at least one neural network, and obtain a final quality score 150 for the video.
[0092] In an embodiment, the image display device 110 may perform quality processing on frames included in the video based on a final quality score 150. As described above, the image display device 110 can improve the quality of frames by performing quality processing on them using the final quality score 150. Figure 1 In this process, the image display device 110 can improve the resolution of the input frame 140 as the output frame 160 based on the final quality score 150 obtained from the video quality evaluation device 130. The image display device 110 can then output the output frame 160 on the display.
[0093] In another embodiment, the video quality assessment device 130 may perform quality processing on the frames included in the video directly based on the final quality score 150 and send the quality-processed output frame 160 to the image display device 110, instead of sending the obtained final quality score 150 to the image display device 110.
[0094] In another embodiment, the video quality assessment device 130 may not be separate from the image display device 110, but may be included in the image display device 110 to perform the above-described functions. The video quality assessment device 130 may be provided as at least one hardware chip or implemented as a software module, or may be implemented as a combination of hardware and software embedded in the image display device 110. In this case, the image display device 110 may first perform a quality assessment on the video including the input frame 140 before outputting the input frame 140 to the screen. The image display device 110 may perform quality processing such as adjusting the distortion of the input frame 140 according to the quality assessment score, and may output the quality-processed output frame 160 to the screen.
[0095] As described above, according to the embodiment, the video quality assessment device 130 can identify whether a frame is a completely blurred frame or a partially blurred frame by using the blur level of the frame, and therefore, different methods can be used to assess the quality of the frame.
[0096] According to an embodiment, the video quality assessment device 130 can obtain different features from the frame according to the type of the frame in order to assess the quality of the frame.
[0097] According to an embodiment, the video quality assessment device 130 can obtain a final quality score for the video, process the quality of the video based on the final quality score, and send the video to the image display device 110.
[0098] According to an embodiment, the video quality assessment device 130 can obtain a final quality score for the video and send the final quality score to the image display device 110, and the image display device 110 can process the video based on the final quality score to improve its quality and output the video.
[0099] Figure 2 The diagram shows power spectra 210 and 220, as well as spectral envelopes 211 and 221, obtained from different input frames according to an embodiment.
[0100] The video quality assessment device 130 performs a domain transformation by dividing the input frame into multiple sub-regions and performing a Fast Fourier Transform (FFT) on each sub-region. The video quality assessment device 130 can obtain the signal in the frequency domain for each sub-region.
[0101] The image of each sub-region is a discrete signal rather than a continuous signal, and is defined by a finite period. The video quality assessment device 130 can perform a Fourier transform on the image of each sub-region to decompose and represent the image of each sub-region as a sum of various 2D sine waves. Assuming that the image of the sub-region is a signal f(x,y) of size W×H, the video quality assessment device 130 can obtain the frequency domain signal F(u,v) by performing a discrete Fourier transform on the image of the sub-region. Here, F(u,v) represents the coefficients of a periodic function component with frequency u in the x-axis direction and frequency v in the y-axis direction.
[0102] F(u,v) is a complex number and therefore includes both real and imaginary parts. The video quality assessment device 130 can obtain the power spectrum from the magnitude |F(u,v)| of the complex number F(u,v). The power spectrum represents the intensity of the corresponding frequency components included in the original image. Because the low-frequency region of the power spectrum has very large values, while most of its other regions have values close to 0, the power spectrum is usually represented as a logarithmic value when it is represented as an image. Furthermore, because the original power spectrum image has larger values towards the edges, it is difficult to discern the shape of the power spectrum, and therefore the spectrum is shifted so that an image with its origin centered can be generated as the power spectrum image.
[0103] exist Figure 2 Two power spectra, 210 and 220, are shown on the left. Power spectrum 210, obtained from the clear first image, is shown in... Figure 2 The upper part, and the power spectrum 220 obtained for the blurry second image which is less sharp than the first image, is shown in Figure 2 The lower end.
[0104] It can be seen that the power spectrum 210 obtained for the first image is different from the power spectrum 220 obtained for the second image. In other words, it can be seen that in the power spectrum 210 obtained for the first image, the power values do not change abruptly and are smoothly filled throughout the region, while in the power spectrum 220 obtained for the second image, the power components are concentrated in a specific region relative to the center of the power spectrum 220.
[0105] The video quality assessment device 130 can obtain a spectral envelope from the power spectrum. The video quality assessment device 130 can select one or more columns or rows that pass through the center of the power spectrum or are located within a specific distance from the center of the power spectrum, and can obtain a spectral envelope for the selected columns or rows. The spectral envelope is a line that connects power spectrum values from lower frequencies to higher frequencies and displays the frequency characteristics of the region of interest.
[0106] exist Figure 2 The right side shows the spectral envelopes 211 and 221 obtained from power spectra 210 and 220. Figure 2 In the right figure, the x-axis indicates the bin index value, and the y-axis indicates the value of the spectral envelope of the frequency obtained by taking the logarithm after normalizing the absolute value of the signal.
[0107] exist Figure 2 On the right side, the first spectral envelope 211 is a graph obtained from the power spectrum 210 of the first image, and the second spectral envelope 221 is a graph obtained from the power spectrum 220 of the second image.
[0108] like Figure 2 As shown, the first spectral envelope 211 and the second spectral envelope 221 have different slopes in the graph. In other words, the first spectral envelope 211 has a gentle slope, while the second spectral envelope 221 includes phases with a flat slope and phases with abruptly changed slopes. Spectral envelope values are typically peaks when the bin index values are in the middle. For example, when the value of the spectral envelope changes from 70% of the peak value to the peak value when the bin index value changes by 100, the slope can be abruptly changed slope. When the same spectral envelope value changes when the bin index value changes by 200, the slope can be gentle.
[0109] In one embodiment, the video quality assessment device 130 can estimate the blur level of the image for each sub-region by using the spectral envelope.
[0110] Statistically, when the tilt of the spectral envelope has a sudden change, a blurred region exists in the image. Therefore, the video quality assessment device 130 can estimate the blur level by using the tilt of the spectral envelope.
[0111] In one embodiment, the video quality assessment device 130 may estimate the blur level of the first image to be 0 based on the fact that there are no points where the tilt of the first spectral envelope 211 changes rapidly.
[0112] In an embodiment, the video quality assessment device 130 can identify locations where the tilt of the spectral envelope exceeds a threshold. For example, the video quality assessment device 130 can obtain a bin index value 223 for a point based on the fact that there is a point where the tilt of the second spectral envelope 221 suddenly changes beyond a specific reference value, and can estimate the blur level of the second image from the obtained bin index value 223.
[0113] For example, the video quality assessment device 130 can estimate the blur level of the second image based on the tilt value of the graph during a phase in which the tilt of the second spectral envelope 221 suddenly changes.
[0114] For example, the video quality assessment device 130 can estimate the blur level based on the ratio of the phases in which the tilt of the second spectral envelope 221 is relatively flat to the phases in which the tilt of the second spectral envelope 221 changes abruptly.
[0115] In one embodiment, the video quality assessment device 130 can divide the input frame into multiple sub-regions, obtain the spectral envelope for each sub-region, and estimate the blur level of each sub-region by using the graphical shape of the spectral envelope corresponding to each sub-region.
[0116] Figure 3 The diagram shows power spectra 210 and 230, as well as spectral envelopes 211 and 231, obtained from different input frames according to an embodiment.
[0117] exist Figure 3 In the diagram, the power spectrum 210 of the first image and the power spectrum 230 of the third image are shown on the left, and the graph on the right shows the spectral envelopes 211 and 231 obtained from the power spectrum 210 and the power spectrum 230, respectively.
[0118] exist Figure 3 In the power spectrum 210 and power spectrum 230 shown on the left, power spectrum 210 is for... Figure 2 The first image shown is clear, and the power spectrum 230 is obtained for a third image that is less clear than the second image.
[0119] With Figure 2In the same manner, it can be seen that the power spectrum 210 obtained for the first image differs from the power spectrum 230 obtained for the third image. It can be observed that in the power spectrum 210 obtained for the first image, the power values do not change abruptly and are smoothly filled throughout the region. In contrast, in the power spectrum 230 obtained for the third image, the power components are concentrated in a specific region at the center of the power spectrum 230. Furthermore, it can be seen that the power spectrum 230 obtained for the third image has a higher center-based concentration than the power spectrum 220 obtained for the second image. This high concentration of the power spectrum can indicate abrupt changes in power values on the image.
[0120] Spectral envelopes 211 and 231 are shown in Figure 3 On the right side. Figure 3 In the right-hand image, the x-axis indicates the bin index value, and the y-axis indicates the frequency. The first spectral envelope 211 is obtained from the power spectrum 210 of the first image, and the third spectral envelope 231 is obtained from the power spectrum 230 of the third image. It can be seen that the first spectral envelope 211 has a gentle slope, while the third spectral envelope 231 includes stages with a flat slope and stages with abruptly changed slope. Furthermore, it can be seen that, compared to the first spectral envelope 211, the graph of the third spectral envelope 231 has a very abruptly changed slope, and... Figure 2 Compared to the second spectral envelope 221 shown, it has a further abrupt change in slope. For example, when the bin index value changes by 50, when the value of the spectral envelope changes from 70% of the peak value to the peak value, the slope of the graph can be a very abrupt change in slope.
[0121] In one embodiment, the video quality assessment device 130 can estimate the blur level of the image for each sub-region by using the steepness of the spectral envelope's tilt. The video quality assessment device 130 can determine that the steeper the spectral envelope's tilt, the greater the corresponding blur level of the image.
[0122] In an embodiment, the video quality assessment device 130 can obtain the bin index value 233 of the point where the tilt of the spectral envelope begins to change abruptly, and can estimate the blur level from the obtained bin index value 233. For example, the video quality assessment device 130 can determine that the larger the bin index value of the point where the tilt of the spectral envelope begins to change abruptly, the greater the blur level of the corresponding image. The video quality assessment device 130 can determine that the blur level of the third image is higher than that of the second image by using the fact that: Figure 3 The bin index value 233 of the point where the tilt of the third spectral envelope 221 suddenly changes is greater than 233. Figure 2 The bin index value 223 of the point where the tilt of the second spectral envelope 231 suddenly changes.
[0123] Thus, the video quality assessment device 130 can estimate the blur level by using the spectral envelope in the frequency domain. The video quality assessment device 130 can estimate the blur level of the image by using the tilt value of the spectral envelope, the ratio of the phases of the spectral envelope with a gentle tilt to the phases of the spectral envelope with a sudden change in tilt, or the bin index value at the point where the tilt of the spectral envelope begins to change abruptly.
[0124] Thus, because the video quality assessment device 130 calculates the blur level of an image based on rules or statistics in the frequency domain, the video quality assessment device 130 can estimate the blur level of an image with minimal computation and at high speed.
[0125] Figure 4 This is a diagram illustrating a fully blurred frame and a partially blurred frame according to an embodiment.
[0126] Reference Figure 4 The video quality assessment device 130 can receive multiple frames included in a video and divide each frame into multiple sub-regions.
[0127] Each sub-region can be a region comprising a specific number of pixels. The number of sub-regions or the size of each sub-region can be preset by the user or the video quality assessment device 130, or can be changed by the user or the video quality assessment device 130 based on frames. The user or the video quality assessment device 130 can adjust the number of sub-regions or the size of each sub-region for each frame, such that the frame is divided into more or fewer sub-regions.
[0128] The video quality assessment device 130 can estimate the blur level for each of multiple sub-regions. Based on the estimated blur level for each sub-region, the video quality assessment device 130 can identify the type of frame.
[0129] In an embodiment, when the number of sub-regions whose blur level is equal to or greater than a specific value is equal to or less than a first specific number, the video quality assessment device 130 can determine that there is no blur in the corresponding frame. In this case, the video quality assessment device 130 can highly evaluate the quality score for the corresponding frame.
[0130] In an embodiment, when the number of sub-regions whose blur level is equal to or greater than the specific value exceeds the first specific number, the video quality assessment device 130 can identify whether the corresponding frame is a completely blurred frame or a partially blurred frame.
[0131] In an embodiment, a fully blurred frame can indicate a frame in which a blur level equal to or greater than a specific reference value is included in a specific region or a larger region of the entire frame. The video quality assessment device 130 can identify frames in which a second specific number of sub-regions within a frame have a blur level equal to or greater than the specific reference value as fully blurred frames.
[0132] Partially blurred frames can indicate that the number of sub-regions in a frame whose blur level is equal to or higher than the specific reference value is greater than a first specific number and less than a second specific number of frames.
[0133] exist Figure 4 In this context, when the first frame 410 is input, the video quality assessment device 130 can divide the first frame 410 into a specific number (e.g., nine) sub-regions and estimate the blur level for each sub-region. Figure 4 In the first table 411, the blur level is indicated by the video quality assessment device 130 for each sub-region of the first frame 410.
[0134] The video quality assessment device 130 can identify the number of sub-regions whose blur level value exceeds a specific reference value. For example, the blur level can be a value between 0 and 1. For example, the video quality assessment device 130 can identify sub-regions whose blur level is equal to or greater than a specific value (e.g., 0.3). In this example, the blur level of all nine sub-regions is equal to or greater than 0.3. The video quality assessment device 130 can identify the first frame 410 as a fully blurred frame based on the fact that the number of sub-regions with a blur level equal to or greater than 0.3 exceeds a specific number (e.g., 80% of the total number of sub-regions).
[0135] Similarly, when the second frame 420 is input, the video quality assessment device 130 can divide the second frame 420 into sub-regions and estimate the blur level for each sub-region. Figure 4 In the second table 421, the blur level obtained by the video quality assessment device 130 for each sub-region of the second frame 420 is indicated. The video quality assessment device 130 can identify the number of sub-regions whose blur level value for each sub-region is equal to or higher than the specified value (e.g., 0.3). In this example, the blur level for all four sub-regions is equal to or greater than 0.3. The video quality assessment device 130 can identify the second frame 420 as a partially blurred frame if the number of sub-regions with a blur level equal to or higher than 0.3 exceeds a specific percentage (e.g., 30%) and does not exceed 80%.
[0136] As described above, according to the embodiment, the video quality assessment device 130 can determine whether the input frame is an unblurred frame, a completely blurred frame, or a partially blurred frame based on the blur level of multiple sub-regions included in the input frame.
[0137] Figure 5 This is a diagram illustrating importance information for each sub-region included in a frame according to an embodiment.
[0138] In this embodiment, the video quality assessment device 130 does not uniformly assess the quality of the entire frame. Instead, it uses the features of each sub-region of the frame to assess the quality of the entire frame. That is, the video quality assessment device 130 can more accurately assess the quality of the frame by reflecting the features of each sub-region. Therefore, in this embodiment, the video quality assessment device 130 can obtain importance information for each sub-region of the frame.
[0139] Importance information is information related to factors that may affect the quality score, and can be obtained from the entire frame, from each sub-region included in the frame, or by considering both the entire frame and sub-regions.
[0140] Typically, when a person identifies blur in an image, even when the blur level is the same, the person will not perceive blur equally in all images, and tends to perceive blur differently based on various elements of each image. In this embodiment, the video quality assessment device 130 may consider elements that affect the identification of blur as important information.
[0141] Importance information may include information related to whether a blurred object in the frame is foreground or background.
[0142] Reference Figure 5 ,exist Figure 5 In frame 510, the forest is the background and located behind the image, while the cyclist is the foreground and located in front of the image. The forest as the background has high blur, while the cyclist as the foreground has low blur.
[0143] exist Figure 5 In frame 520, the forest is the background and located behind the image, while the moving person is the foreground and located in front of the image. Unlike frame 510, it can be seen that in frame 520, the forest as the background has low blur, while the moving person has higher blur.
[0144] Humans tend to perceive the blurriness of objects in the foreground as greater than that of objects in the background. In other words, in... Figure 5 In the image, frame 520, where the foreground object has a large blur, is perceived as more distorted than frame 510, where the background object has a large blur.
[0145] In one embodiment, the video quality assessment device 130 can take into account the different characteristics of each sub-region by assigning more important information to the foreground and less important information to the background.
[0146] Figure 6 This is a diagram illustrating importance information for each sub-region included in a frame according to an embodiment.
[0147] Humans perceive blurriness differently based on various elements of an image, and these elements can include information about the type of frame. In other words, humans tend to perceive blurriness differently depending on the type of frame it belongs to.
[0148] Reference Figure 6 Frame 610 shows an image with a lot of movement that requires focusing on, and may correspond to, for example, a sports genre. Frame 620 shows an image with a little movement that does not give significant meaning to the movement, and may correspond to, for example, a genre such as a drama or lecture. Compared to frame 620 with a little movement, a person perceives greater blur in frame 610 with a lot of movement that requires focusing on. This perception may affect the quality rating.
[0149] Furthermore, depending on the object's position within a frame, people may perceive blur to varying degrees. For example, because people tend to see more of the center of the screen than the edges when watching video, they perceive blur differently in frames that are blurred in the center versus frames that are blurred at the edges.
[0150] Furthermore, considering the semantic information of objects included in a frame, people tend to view images more ambiguously. This suggests that the degree of fuzzy perception of objects can vary depending on what objects are included in the frame (i.e., the meaning of the objects in the frame). For example, when a video is such as Figure 6 When viewing a sports event image in frame 610, the importance information that a viewer of the video can perceive varies depending on whether the object included in the frame is an athlete or a spectator. For example, a person perceives blurriness in a frame where the athlete's face is blurred compared to a frame where the spectator's face is blurred.
[0151] Thus, important information related to factors that may affect the quality score may include at least one of the following: information about whether the objects included in the frame are foreground or background, information about the type of frame, semantic information, or positional information.
[0152] As described above, the video quality assessment device 130 according to the embodiment can more accurately assess the quality of a frame by reflecting the perceptual characteristics of a person with poor recognition.
[0153] Furthermore, the video quality assessment device 130 according to the embodiment does not uniformly assess the quality of the entire frame, but can obtain importance information for each sub-region and use that importance information to perform quality assessment, thereby taking into account the characteristics of each sub-region to assess the quality of the frame.
[0154] Figure 7 This is a block diagram of the internal structure of a computing device 700 according to an embodiment.
[0155] Reference Figure 7 The computing device 700 may include a processor 710 and a memory 720. Figure 7 The computing device 700 can be Figure 1 An example of the video quality assessment device 130 shown.
[0156] The memory 720 according to an embodiment may store one or more instructions. The one or more instructions may comprise code generated by a compiler or code executable by an interpreter. The memory 720 may store one or more programs executed by the processor 710. The memory 720 may store at least one neural network and / or predefined operating rules or AI models. The memory 720 may also store data input to or output by the computing device 700.
[0157] The memory 720 may be non-transitory and may include at least one type of storage medium selected from flash memory, hard disk memory, multimedia card micro-memory, card-type memory (e.g., Secure Digital (SD) or Extreme Speed Digital (XD) memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable ROM (EEPROM), programmable ROM (PROM), magnetic memory, magnetic disk, or optical disk.
[0158] The processor 710 controls the overall operation of the computing device 700. The processor 710 can control the operation of the computing device 700 by executing one or more instructions stored in the memory 720.
[0159] In one embodiment, the processor 710 may divide each frame included in the input video into multiple sub-regions and obtain the spectral envelope for each sub-region. The processor 710 may use the spectral envelope to estimate the blur level for each sub-region and, based on this, identify whether the frame is an unblurred frame, a fully blurred frame, or a partially blurred frame.
[0160] In this embodiment, the processor 710 may obtain spectral envelope distribution features from the spectral envelope obtained for each sub-region. The spectral envelope distribution features may be statistical properties of the distribution of the spectral envelope pattern.
[0161] In an embodiment, the quality score based on analysis can be a score of statistical attributes that reflect the frequency characteristics of a frame.
[0162] In an embodiment, when a frame is determined to be a fully blurred frame, the processor 710 can analyze the spectral envelope distribution features obtained for each of the multiple sub-regions included in the fully blurred frame, and obtain an analysis-based quality score for the fully blurred frame.
[0163] In an embodiment, when it is determined that the input frame is a partially blurred frame, the processor 710 can obtain a model-based quality score for the partially blurred frame.
[0164] In an embodiment, the model-based quality score may indicate the quality score obtained using at least one neural network for a partially blurred frame.
[0165] To obtain a model-based quality score, the processor 710 may use at least one neural network.
[0166] A neural network receives data, performs operations for analysis and classification, and outputs result data. To ensure that the neural network accurately outputs result data corresponding to the input data, it needs to be trained. Here, "training" can mean training the neural network so that it can discover or master methods for inputting various types of data into the network and analyzing the input data, methods for classifying the input data, and / or methods for extracting the features needed to generate result data from the input data. Training the neural network means generating an AI model with the desired characteristics by applying a learning algorithm to multiple training datasets. This learning can be performed within the computing device 700 that performs the AI, or via a separate server / system.
[0167] Here, a learning algorithm is a method of training a specific target device (e.g., a robot) using multiple training data sets and allowing the target device to make decisions or predictions on its own. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, and the learning algorithms in the embodiments are not limited to the examples above, unless otherwise stated.
[0168] An algorithm set that outputs data corresponding to the input data through a neural network, the software that executes the algorithm set, and / or the hardware that executes the algorithm set can be referred to as an "AI model" (or "artificial intelligence model").
[0169] The processor 710 can process input data according to predefined operating rules or AI models. Specific algorithms can be used to create the predefined operating rules or AI models. Furthermore, the specific algorithms can be used to train the AI models. The processor 710 can then generate output data corresponding to the input data using the AI models.
[0170] The processor 710 can store at least one AI model. The processor 710 can generate output data from the input image by using multiple AI models.
[0171] In one embodiment, the processor 710 may use at least one neural network to obtain importance information from partially blurred frames.
[0172] In one embodiment, the processor 710 can calculate a weight for each sub-region by using importance information and a spectral envelope distribution feature that represents the frequency characteristics of each sub-region. The weights can indicate the characteristics of each sub-region in the quality assessment.
[0173] The processor 710 can generate a weight matrix for the entire partially blurred frame by using weights for each sub-region.
[0174] In one embodiment, the processor 710 can obtain a model-based quality score for a partially blurred frame from a weight matrix and the partially blurred frame using at least one neural network.
[0175] As described above, in the embodiments, the processor 710 can obtain an analysis-based quality score for a fully blurred frame by analyzing the frequency characteristics of the fully blurred frame, and can obtain a model-based quality score for a partially blurred frame by using at least one neural network.
[0176] Figure 8 The embodiments are shown in more detail. Figure 7 A block diagram of the configuration of the 710 processor.
[0177] Reference Figure 8 The processor 710 may include an identifier 810, an analysis-based quality score obtainr 820, and a model-based quality score obtainr 830, depending on the functions to be performed.
[0178] Recognizer 810 can divide the input frame into sub-regions and convert each sub-region into a signal in the frequency domain. Recognizer 810 can obtain the power spectrum from the signal in the frequency domain and obtain the spectral envelope from the power spectrum.
[0179] Recognizer 810 can estimate the blur level for each sub-region by using the spectral envelope. Recognizer 810 can identify whether the input frame is an unblurred frame, a partially blurred frame, or a completely blurred frame based on the number of sub-regions whose blur level is equal to or greater than a specific reference value.
[0180] Furthermore, the recognizer 810 can obtain spectral envelope distribution features representing statistical properties of the spectral envelope graph from the spectral envelope of each sub-region of the input frame.
[0181] When the frame identified by the recognizer 810 is an unblurred frame, the unblurred frame can bypass both the analysis-based quality score obtainr 820 and the model-based quality score obtainr 830. Furthermore, image quality processing can be omitted from the unblurred frame. That is, an unblurred frame can be output without any image quality enhancement.
[0182] When the frame identified by the recognizer 810 is a completely blurred frame, the completely blurred frame is input into the analysis-based quality score obtainr 820.
[0183] The analysis-based quality score obtainr 820 can receive the spectral envelope distribution features of a completely blurred frame from the recognizer 810 and analyze the spectral envelope distribution features. The analysis-based quality score obtainr 820 can analyze the spectral envelope distribution features to obtain a quality score for the completely blurred frame.
[0184] When the frame identified by the recognizer 810 is a partially blurred frame, the partially blurred frame is input to the model-based quality score obtainr 830.
[0185] The model-based quality score obtainr 830 can obtain a model-based quality score from the input partially blurred frame.
[0186] The model-based quality scorer 830 can use at least one neural network to obtain importance information for each sub-region included in the partially blurred frame. The model-based quality scorer 830 can use the spectral envelope distribution features of each sub-region of the partially blurred frame obtained by the recognizer 810 and the importance information for each sub-region of the partially blurred frame obtained using the neural network to obtain weights for each sub-region. The model-based quality scorer 830 can use the weights for each sub-region to generate a weight matrix for the entire partially blurred frame.
[0187] The model-based quality scorer 830 can use at least one neural network to obtain the final score of the partially blurred frame from the weight matrix and the partially blurred frame. That is, the model-based quality scorer 830 can use at least one neural network to obtain feature vectors from the partially blurred frame and obtain weighted features by considering the feature vectors and the weight matrix together. The model-based quality scorer 830 can obtain a quality assessment score for the entire partially blurred frame from the weighted features.
[0188] When reconstructing quality on a frame-by-frame basis or when processing is required, the quality score for fully blurred frames obtained by the analysis-based quality score obtainr 820 and the quality score for partially blurred frames obtained by the model-based quality score obtainr 830 can be used.
[0189] In an embodiment, the processor 710 may process the quality of a video comprising frames by using at least one of an analysis-based quality score or a model-based quality score.
[0190] For example, processor 710 can improve the quality of each frame by using the score obtained for each frame. For example, processor 710 can send the score obtained for each frame to image display device 110, so that image display device 110 processes the quality for each frame.
[0191] Figure 9 The embodiments are shown in more detail. Figure 7 A block diagram of the configuration of the 710 processor. Figure 9 The processor 710 may include Figure 8 The processor 710 includes a recognizer 810, an analysis-based quality score obtainr 820, and a model-based quality score obtainr 830. In the following text, terms related to... Figure 8 The descriptions given are repeated.
[0192] Reference Figure 9 The recognizer 810 may include a spectral envelope acquirer 911 and a spectral envelope distribution feature acquirer 913.
[0193] The spectral envelope acquirer 911 can divide the input frame into sub-regions and perform a Fourier transform on each sub-region to obtain the signal in the frequency domain. The spectral envelope acquirer 911 can obtain the power spectrum of the signal in the frequency domain and obtain the spectral envelope from the power spectrum.
[0194] The spectral envelope acquirer 911 can estimate the ambiguity level for each sub-region by using the spectral envelope. For example, the spectral envelope acquirer 911 can estimate the ambiguity level for each sub-region by using the slope value of the spectral envelope. For example, the spectral envelope acquirer 911 can estimate the ambiguity level based on the ratio of the phases of the spectral envelope with gentle slope to the phases of rapidly changing slope. For example, the spectral envelope acquirer 911 can estimate the ambiguity level from the bin index value at the point where the slope of the spectral envelope changes rapidly.
[0195] The spectral envelope distribution feature obtainr 913 can obtain the spectral envelope distribution features from the spectral envelope obtained by the spectral envelope obtainr 911 for each sub-region.
[0196] The spectral envelope distribution characteristics represent the statistical properties of the distribution of the spectral envelope graph, and the statistical properties of the graph distribution may include at least one of the following: mode, median, arithmetic mean, harmonic mean, geometric mean, global minimum, global maximum, range, variance, or bias.
[0197] The spectral envelope distribution feature acquirer 913 can acquire the spectral envelope distribution features for each frame in both cases where the frame is a fully blurred frame and a partially blurred frame.
[0198] The spectral envelope distribution feature obtainr 913 can obtain the spectral envelope distribution features of a completely blurred frame and send these features to the analysis-based quality score obtainr 820. Furthermore, the spectral envelope distribution feature obtainr 913 can obtain the spectral envelope distribution features of a partially blurred frame and send these features to the model-based quality score obtainr 830.
[0199] The analysis-based quality score obtainr 820 can receive spectral envelope distribution features obtained for each sub-region included in the fully blurred frame from the spectral envelope distribution feature obtainr 913, and analyze the spectral envelope distribution features. The spectral envelope distribution features for the fully blurred frame can be statistical properties of the spectral envelope graphs obtained for each sub-region included in the fully blurred frame.
[0200] The analysis-based quality scorer 820 can analyze the statistical properties of the spectral envelope graph and obtain a quality score for a fully blurred frame from these statistical properties.
[0201] The model-based quality score generator 830 may include a first neural network 931, a weight matrix generator 933, and a second neural network 935.
[0202] In an embodiment, the first neural network 931 may be a model trained to analyze and classify input data to extract importance information from the input data, wherein the importance information is a feature that affects video quality assessment.
[0203] The first neural network 931 may be an algorithm or a set of algorithms for extracting features from input data, software for executing the set of algorithms, and / or hardware for executing the set of algorithms.
[0204] The first neural network 931 may be a deep neural network (DNN) comprising two or more hidden layers. The first neural network 931 may include a structure that processes input data through hidden layers to output processed data. Each layer of the first neural network 931 is represented by one or more nodes, and nodes between layers are connected by edges.
[0205] The first neural network 931 can obtain importance information for each sub-region included in the partially blurred input frame.
[0206] Importance information may include at least one of the following: information about whether the objects included in the frame are foreground or background, information about the type of frame, semantic information, or location information of a sub-region.
[0207] The first neural network can obtain importance information from the entire partially blurred frame or each sub-region, or by considering the entire partially blurred frame and sub-regions together. For example, the first neural network 931 can obtain information about the type to which the frame belongs from the entire partially blurred frame. For example, the first neural network 931 can obtain coordinate values indicating the location of each sub-region included in the frame. For example, the first neural network 931 can obtain semantic information about the objects included in the frame by considering the entire frame and sub-regions together.
[0208] The weight matrix generator 933 can receive importance information from the first neural network 931 and receive the spectral envelope distribution features of partially blurred frames from the spectral envelope distribution feature obtainr 913.
[0209] The weight matrix generator 933 can obtain the weights for each sub-region by using importance information and spectral envelope distribution characteristics for each sub-region. In an embodiment, the weight matrix generator 933 can obtain the weights for each sub-region by multiplying the importance information by the spectral envelope distribution characteristics.
[0210] Weights can be information indicating the characteristics of each sub-region. Since importance information and spectral envelope distribution features are obtained from the content of the frame, the weights generated using importance information and spectral envelope distribution features can also vary depending on the content included in the frame.
[0211] The weight matrix generator 933 can use weights for each sub-region to generate a weight matrix for the entire partially blurred frame.
[0212] In an embodiment, the weight matrix generator 933 may take into account the importance information and spectral envelope distribution characteristics of one or more neighboring sub-regions to correct the weights of the sub-regions.
[0213] The weight matrix generator 933 can correct the spectral envelope distribution characteristics of a sub-region to be more natural by using the spectral envelope distribution characteristics of one or more neighboring sub-regions. For example, the weight matrix generator 933 can correct the spectral envelope distribution characteristics of the first sub-region by using at least one of the one or more neighboring sub-regions located adjacent to the first sub-region (i.e., the right, left, top, and bottom sides of the first sub-region relative to the first sub-region included in the partially blurred frame).
[0214] The weight matrix generator 933 can correct the spectral envelope distribution characteristics of the first sub-region by considering the adjacent sub-regions and the first sub-region together. For example, the weight matrix generator 933 can correct the spectral envelope distribution characteristics of the first sub-region by using the average of the spectral envelope distribution characteristic values of the first sub-region and the adjacent sub-regions.
[0215] Similarly, the weight matrix generator 933 can correct the importance information of the first sub-region by taking into account the importance information of at least one neighboring sub-region adjacent to the first sub-region.
[0216] The weight matrix generator 933 can obtain the weight values for the first sub-region by using the spectral envelope distribution characteristics and the importance information of the reference neighboring sub-region correction, and thereby generate the weight matrix.
[0217] The weight matrix is input into the second neural network 935.
[0218] Because multidimensional feature vectors are complex, statistical analysis methods alone cannot be used to consider the various characteristics of partially blurred frames. Therefore, in this embodiment, the model-based quality scorer 830 can obtain a final score for partially blurred frames by using a second neural network 935 to consider the various characteristics of each sub-region.
[0219] The second neural network 935 may be an algorithm, a set of algorithms, software that executes the set of algorithms, and / or hardware that executes the set of algorithms, trained to analyze and classify input data to extract features of the input data and obtain quality scores from the features.
[0220] The second neural network 935 can be a regression model. A regression model is an analytical method that obtains a model between continuous variables and then measures the fit. The second neural network 935 can be one of support vector machine (SVM) regression, random forest regression, or deep neural networks, but is not limited to these.
[0221] The second neural network 935 can be a pre-trained model for video quality assessment. The second neural network 935 can learn the Mean Opinion Score (MOS). MOS is obtained through subjective human evaluation and indicates the average value of various parameters for the video quality assessed by a person. The second neural network 935 can be trained by pre-learning the correlation between feature vectors and MOS.
[0222] The trained second neural network 935 can receive partially blurred frames and obtain feature vectors representing various features related to the quality of the partially blurred frames. Quality-related features may include at least one of the following: blur-related features, motion-related features, content-related features, perceptual features, spatial features, depth features extracted from multiple hidden layers for each layer, or features extracted statistically from lower to higher levels. The second neural network 935 can obtain feature vectors representing the above features before the final output operation.
[0223] The second neural network 935 can receive a weight matrix and partially blurred frames. The second neural network 935 can obtain features that reflect the weights by considering the obtained features and weight matrix together.
[0224] The second neural network 935 can obtain a weighted feature vector from the partially blurred input frame and output an objective quality score that closely matches the subjective rating of a person.
[0225] As described above, according to the embodiment, the model-based quality score obtainr 830 can use a second neural network 935 that has learned MOS to obtain features reflecting weights from partially blurred frames, and obtain a quality score from the features that is similar to a human subjective evaluation result.
[0226] Figure 10 According to the embodiments Figure 7 A block diagram of the internal structure of the processor 710.
[0227] Reference Figure 10 ,Apart from Figure 8 In addition to the recognizer 810, the analysis-based quality score obtainr 820, and the model-based quality score obtainr 830 included in the processor 710 shown, the processor 710 may also include a final quality score obtainr 1010. In the following text, terms related to... Figure 8 The descriptions given are repeated.
[0228] The final quality score obtainr 1010 can receive quality scores for fully blurred frames from the analysis-based quality score obtainr 820 and quality scores for partially blurred frames from the model-based quality score obtainr 830.
[0229] The final quality scorer 1010 can accumulate scores for each frame and timestamps for each frame. The final quality scorer 1010 can also accumulate scores for each frame received over a specific time period to obtain time-series data.
[0230] The final quality score obtainr 1010 can take into account the temporal impact or temporal dependence related to the identified video by using the quality scores of frames accumulated over time.
[0231] For example, even if the video quality subsequently improves, people tend to continue to evaluate the video based on its initial poor quality. For instance, when a series of poor-quality frames are output, people tend to perceive the quality of consecutive frames as worse than when evaluating a single frame of poor quality. For example, when videos have the same level of blur, people tend to perceive a higher level of blur in a video at a lower frame rate (fps) compared to a video at a higher frame rate (fps).
[0232] The final quality score obtainr 1010 can take this time effect into account when calculating the final quality score.
[0233] The final quality scorer 1010 obtains a final quality score for the entire video by smoothing the time-series data. For example, the final quality scorer 1010 can use simple heuristics or complex models to smooth the time-series data.
[0234] In an embodiment, when the final quality score obtainr 1010 obtains the final quality score using a model, the final quality score obtainr 1010 may use at least one neural network. For ease of explanation, the neural network used by the final quality score obtainr 1010 is referred to as the third neural network.
[0235] The third neural network can be an algorithm, a set of algorithms, software that executes the set of algorithms, and / or hardware that executes the set of algorithms, trained to analyze and classify accumulated input data to extract time-series features of the input data and obtain a final quality score from the time-series features.
[0236] In this embodiment, the final quality scorer 1010 may use a Long Short-Term Memory (LSTM) model as a third neural network. LSTM is a recurrent neural network (RNN) that learns long-term dependencies between time steps of sequence data. LSTM can receive sequential or time-series data and learn the long-term dependencies between time steps of the sequence data.
[0237] The final quality scorer 1010 can receive features accumulated through a third neural network and can obtain a final quality score for the entire video, taking into account the effects over time.
[0238] As described above, according to the embodiment, the video quality assessment device 130 can accumulate the scores obtained for each of the fully blurred frames and partially blurred frames to obtain time series data and smooth the time series data, and obtain a final quality score for the entire video.
[0239] Figure 11 This is a block diagram of the internal structure of a computing device 1100 according to another embodiment.
[0240] Figure 11 The computing device 1100 can be Figure 1 An example of the video quality assessment device 130 shown. (See reference...) Figure 11 The computing device 1100 may include a processor 1110 and a memory 1120.
[0241] The memory 1120 according to an embodiment may store at least one instruction. The at least one instruction may include code generated by a compiler or code executable by an interpreter. The memory 1120 may store at least one program executed by the processor 1110. At least one neural network and / or predefined operating rules or AI model may be stored in the memory 1120. Furthermore, the memory 1120 may store data input to or output from the computing device 1100.
[0242] Processor 1110 controls the overall operation of computing device 1100. Processor 1110 can control the operation of computing device 1100 by executing one or more instructions stored in memory 1120.
[0243] exist Figure 11In the processor 1110, there may be a recognizer 1111, a first feature obtainr 1112, a second feature obtainr 1113, an importance information obtainr 1114, a weight matrix generator 1115, and a final score obtainr 1116.
[0244] In one embodiment, the recognizer 1111 can divide each frame included in the input video into multiple sub-regions and obtain the spectral envelope for each sub-region. The recognizer 1111 can use the spectral envelope to estimate the blur level for each sub-region, and based on the blur level, can identify whether the frame is an unblurred frame, a fully blurred frame, or a partially blurred frame.
[0245] In an embodiment, the recognizer 1111 can obtain spectral envelope distribution features from the spectral envelope obtained for each sub-region.
[0246] Figure 11 The computing device 1100 can obtain different features from fully blurred frames and partially blurred frames, and obtain the final quality of the entire video from at least one neural network by using different features.
[0247] Instead of using different methods for each frame type to obtain a score for each frame. Figure 7 The computing devices are different from those in the 700 series. Figure 11 The computing device 1100 can obtain different features for each frame type, input the different features as input data into the neural network, and obtain the final score from the neural network.
[0248] In an embodiment, when a frame is determined to be a completely blurred frame, the recognizer 1111 sends the completely blurred frame to the first feature obtainr 1112, and when a frame is determined to be a partially blurred frame, the partially blurred frame is sent to the second feature obtainr 1113.
[0249] The first feature acquirer 1112 can acquire a first feature from a completely blurred frame. In an embodiment, the first feature may be a spectral envelope distribution feature.
[0250] In one embodiment, the first feature acquirer 1112 may receive spectral envelope distribution features for each sub-region of a fully blurred frame from the recognizer 1111. In another embodiment, in addition to the recognizer 1111, the first feature acquirer 1112 may acquire spectral envelope distribution features for each sub-region among a plurality of sub-regions included in the fully blurred frame.
[0251] Spectral envelope distribution features can be statistical properties of the distribution of a spectral envelope graph. Spectral envelope distribution features represent the statistical properties of the distribution of a spectral envelope graph and can include at least one of the following: mode, median, arithmetic mean, harmonic mean, geometric mean, global minimum, global maximum, range, variance, or bias.
[0252] The second feature obtainr 1113 can obtain a second feature from a partially blurred frame. In an embodiment, the second feature may be a feature obtained from the partially blurred frame using at least one neural network. The neural network used by the second feature obtainr 1113 may be an algorithm or set of algorithms for extracting features from input data, software executing the set of algorithms, and / or hardware executing the set of algorithms.
[0253] The second feature extractor 1113 can obtain features related to the quality of the partially blurred frame. The second feature extractor 1113 can use a neural network to extract at least one of the following from the partially blurred frame: blur-related features, motion-related features, content-related features, perceptual features, spatial features, depth features extracted from multiple hidden layers for each layer, or features extracted from lower-level to higher-level statistics.
[0254] Importance information acquirer 1114 can obtain importance information from the input frame. Importance information acquirer 1114 can obtain importance information for each sub-region included in the frame by using a neural network trained to obtain importance information of factors that may affect video quality assessment from the input frame.
[0255] The neural network used by the importance information acquirer 1114 can be a model trained to extract features that influence quality assessment by analyzing and classifying input frames. The neural network used by the importance information acquirer 1114 can be a deep neural network (DNN) comprising two or more hidden layers.
[0256] The neural network used by the importance information obtainr 1114, which obtains importance information from both fully blurred and partially blurred frames, versus obtaining importance information only from partially blurred frames. Figure 9 The first neural network 931 is different. The neural network used by the importance information obtainr 1114, which performs the function of obtaining importance information from the input frame, can be the same as the first neural network 931.
[0257] Importance information acquirer 1114 can acquire importance information for each sub-region from the entire frame or each sub-region, or by considering the entire frame and sub-regions together. Importance information may include at least one of the following: information relating to whether an object included in the frame is foreground or background, information about the type to which the frame belongs, semantic information, or positional information of the sub-region.
[0258] The weight matrix generator 1115 receives importance information from the importance information obtainr 1114. Furthermore, the weight matrix generator 1115 can receive spectral envelope distribution features for each sub-region of the frame input from the recognizer 810.
[0259] The weight matrix generator 1115 can obtain the weights for each sub-region by using the importance information and spectral envelope distribution characteristics for each sub-region. The weight matrix generator 1115 can generate a weight matrix for the entire frame by summing the weights for each sub-region.
[0260] The final score obtainr 1116 can receive a first feature for a fully blurred frame obtained from the first feature obtainr 1112 and a second feature for a partially blurred frame obtained from the second feature obtainr 1113. Because the final score obtainr 1116 receives the first and second features, it is not necessary to obtain the features separately from the input data.
[0261] The final score obtainr 1116 can receive the weight matrix generated by the weight matrix generator 1115 for both fully blurred frames and partially blurred frames. The final score obtainr 1116 can receive the first feature, the second feature, the weight matrix, and the frame, and accumulate the features for each frame and the timestamp of each frame received for a specific time period to obtain time series data.
[0262] For reference Figure 10 The computing device 700 can accumulate scores obtained for each frame received over a specific time period to obtain time-series data, but unlike this, it includes... Figure 11 The final score obtainr 1116 in the computing device 1100 can accumulate features obtained for each frame received over a specific time period. The final score obtainr 1116 can obtain time series data by accumulating first and second features for a specific time period and accumulating each frame and a weight matrix for each frame.
[0263] Taking into account the effects over time, the final scorer 1116 can obtain a final quality score for the entire video from the accumulated time-series data.
[0264] The final scorer 1116 obtains a final quality score for the entire video by smoothing the time-series data. For example, the final scorer 1116 can use simple heuristics or complex models to smooth the time-series data. In an embodiment, the final scorer 1116 can use a Long Short-Term Memory (LSTM) model.
[0265] As described above, according to the embodiment, the computing device 1100 can obtain different features from the frame according to the type of the input frame.
[0266] Furthermore, the computing device 1100 can accumulate different features obtained for each frame received over a specific time period and evaluate the quality of the entire video by using these different features.
[0267] Figure 12This is a block diagram of the internal structure of a computing device 1200 according to an embodiment. (Refer to...) Figure 12 In addition to the processor 1210 and memory 1220, the computing device 1200 may also include a communicator 1230.
[0268] Figure 12 The computing device 1200 may include Figure 7 The computing device 700. As described below, including... Figure 12 The processor 1210 and memory 1220 in the computing device 1200 are respectively connected to the processor 1210 and memory 1220 included in the computing device 1200. Figure 7 The processor 710 and memory 720 in the computing device 700 have the same function, so their repeated description is omitted.
[0269] According to an embodiment, the communicator 1230 can transmit and receive signals by communicating with an external device connected via a wired or wireless network under the control of the processor 1210. The communicator 1230 may include at least one communication module, such as a short-range communication module, a wired communication module, a mobile communication module, a broadcast receiving module, etc. The communication module may include a communication module capable of performing data transmission or reception through a network conforming to communication standards such as tuner, Bluetooth, wireless LAN (WLAN) (Wi-Fi), wireless broadband (Wibro), Global Microwave Access Interoperability (WiMAX), CDMA, or WCDMA.
[0270] In this embodiment, the communicator 1230 can send and receive data with an image display device 110 external to the computer device 1200.
[0271] The communicator 1230 can receive video from the image display device 110. The processor 1210 can determine the blur level by dividing a frame input through the communicator 1230 into sub-regions and identifying the input frame as an unblurred frame, a fully blurred frame, and a partially blurred frame by executing one or more instructions stored in the memory 1220. The one or more instructions may contain code generated by a compiler or code executable by an interpreter. The processor 1210 can obtain an analysis-based quality score for a fully blurred frame corresponding to an input frame, and a model-based quality score for a partially blurred frame corresponding to an input frame.
[0272] The communicator 1230 can send the scores obtained by the processor 1210 for each frame included in the input video to the image display device 110.
[0273] In an embodiment, the processor 1210 can accumulate analysis-based quality scores and model-based quality scores for each frame included in the input video for a specific time period to obtain time series data and smooth the time series data, and obtain a final quality score for the video.
[0274] The communicator 1230 can send the final quality score for the video obtained by the processor 1210 to the image display device 110.
[0275] Figure 13 This is a block diagram of the internal structure of the image display device 1300 according to an embodiment. (Refer to...) Figure 13 The image display device 1300 may include a processor 1310, a memory 1320, a display 1330, and a quality processor 1340.
[0276] Figure 13 Image display device 1300 may include Figure 7 The computing device 700. That is, in the embodiment, the image display device 1300 is not separate from the computing device 700, but may include the computing device 700 that performs functions to evaluate video quality.
[0277] In the following text, including Figure 13 The processor 1310 and memory 1320 in the image display device 1300 are respectively connected to the processor 1310 and memory 1320 included in the image display device 1300. Figure 7 The processor 710 and memory 720 in the computing device 700 have the same function, so their repeated description is omitted.
[0278] The processor 1310 controls the overall operation of the image display device 1300. The processor 1310 can measure the quality of the corresponding video before outputting live broadcast programs or programs received via streaming or downloading VOD services onto the screen.
[0279] The processor 1310 can identify whether an input frame is unblurred, partially blurred, or completely blurred based on its blur level. When the input frame is unblurred, the processor 1310 can output the input frame through the display 1330.
[0280] When the input frame is a fully blurred frame, the processor 1310 can obtain an analysis-based quality score for the fully blurred frame, and when the input frame is a partially blurred frame, the processor 1310 can obtain a model-based quality score for the partially blurred frame.
[0281] The processor 1310 can use time-series data obtained by accumulating analysis-based quality scores and model-based quality scores to obtain a final quality score for the video.
[0282] In one embodiment, the image quality processor 1340 may process frames to improve quality based on at least one of an analysis-based quality score, a model-based quality score, or a final quality score for the video.
[0283] In this embodiment, the image quality processor 1340 may use multiple AI models to perform various image processing operations on each frame or the entire video (or different parts of the video), such as decoding, rendering, scaling, noise filtering, frame rate conversion, and resolution conversion, to improve quality. For example, the image quality processor 1340 may perform image processing operations on each frame separately using different AI models.
[0284] In an embodiment, each of the multiple AI models can be an image reconstruction model, wherein the image reconstruction model is capable of using one or more neural networks to output a result that optimally improves quality based on the score of each frame or the final quality score of the entire video.
[0285] The image quality processor 1340 can select an image reconstruction model from multiple neural network models, or design such a model directly based on a score for each frame or a final score for the entire video. In an embodiment, the model corresponding to each score can be predetermined based on the final score. Each model can be a model trained to process frames with a corresponding score. For example, model A can be a model for improving the quality of frames with quality scores ranging from 0 to 1, and model B can be a model for processing frames with quality scores ranging from 1 to 2. The image quality processor 1340 can improve quality by processing frames using an AI model based on the selected neural network.
[0286] In one embodiment, the image quality processor 1340 may determine the number of times to apply the image reconstruction model to optimally improve the quality of a frame or video. The image quality processor 1340 may optimally improve the quality of a frame or video by repeatedly applying the image reconstruction model to the frame or video according to the determined number of times.
[0287] In an embodiment, the image quality processor 1340 may correct various hyperparameter values used in the neural network based on the score for each frame or the final score of the video. The image quality processor 1340 may correct one or more of various hyperparameter values (such as filter size, filter coefficients, kernel size, and node weights) based on the frame or video score to select the hyperparameter values of the model that will have the best performance when applied to the frame or video. The image quality processor 1340 may then use an AI model with such hyperparameters to optimally improve the quality of the frame or video.
[0288] In one embodiment, the image quality processor 1340 may design filters for performing image reconstruction based on the rating. The image quality processor 1340 may design bandpass filters (BPF) or high-pass filters (HPF) with bandwidth varying according to the frame or video rating, and process the frame or video by using the designed filters to alter the high-frequency bands of the signal.
[0289] The image quality processor 1340 can use the various methods described above to identify AI models that can optimally improve the quality of each frame or the entire video based on frame or video scores. The image quality processor 1340 can optimally improve the quality of frames or videos by using AI models.
[0290] The display 1330 according to the embodiment can output frames and videos processed by the image quality processor 1340.
[0291] When the display 1330 is implemented as a touch screen, the display 1330 can be used as both an input device and an output device. For example, the display 1330 may include at least one of a liquid crystal display (LCD), a thin-film transistor liquid crystal display (TFT-LCD), an organic light-emitting diode (OLED), a flexible display, a three-dimensional (3D) display, or an electrophoretic display. Furthermore, according to embodiments of the image display device 1300, the image display device 1300 may include two or more displays 1330.
[0292] As described above, according to the embodiment, the image display device 1300 can obtain a quality score for each frame and use the quality score to select an image reconstruction model suitable for each frame or the entire video. After improving the quality of each frame or video, the image display device 1300 can output each frame or video through the display 1330.
[0293] Figure 14 This is a block diagram of the internal structure of the image display device 1400 according to an embodiment. Figure 14 Image display device 1400 may include Figure 13 The components of the image display device 1300. Therefore, the components are omitted. Figure 13 The descriptions of processor 1310, memory 1320 and display 1330 are the same as those in the description.
[0294] Reference Figure 14 In addition to the processor 1310, memory 1320 and display 1330, the image display device 800 may also include a tuner 1410, a communicator 1420, a sensor 1430, an input / output interface 1440, a video processor 1450, an audio processor 1460, an audio interface 1470 and a user interface 1480.
[0295] Tuner 1410 can tune to the frequency of a channel selected from a plurality of radio wave components obtained through amplification, mixing, resonance, etc., of wired or wireless broadcast content. Content received via tuner 1410 can be decoded and divided into audio, video, and / or additional information. The audio, video, and / or additional information can be stored in memory 1320 under the control of processor 1310.
[0296] The communicator 1420, under the control of the processor 1310, can connect the image display device 1400 to an external device or server. The image display device 1400 can download or web browse required programs or applications from the external device or server via the communicator 1420. The communicator 1420 can also receive content from external devices.
[0297] The communicator 1420 may include one of a wireless local area network (LAN) 1421, a Bluetooth interface 1422, and a wired Ethernet interface 1423 corresponding to the performance and architecture of the image display device 1400. The communicator 1420 may include a combination of the wireless LAN 1421, the Bluetooth interface 1422, and the wired Ethernet interface 1423. The communicator 1420 may receive control signals via a control device such as a remote controller under the control of the processor 1310. The control signals may be implemented as Bluetooth signals, radio frequency (RF) signals, or Wi-Fi signals. In addition to the Bluetooth interface 1422, the communicator 1420 may also include short-range communication (e.g., near field communication (NFC) or Bluetooth Low Energy (BLE)). According to embodiments, the communicator 1420 may send connection signals to or receive connection signals from external devices via short-range communication such as the Bluetooth interface 1422 or BLE.
[0298] Sensor 1430 can sense a user's voice, a user's image, or interactions with the user, and may include a microphone 1431, a camera 1432, and a light receiver 1433. Microphone 1431 can receive the user's voice, convert the received voice into an electrical signal, and output the electrical signal to processor 1310. Camera 1432 may include a sensor and a lens, and can capture images formed on a screen. Light receiver 1433 can receive light signals (including control signals). Light receiver 1433 can receive light signals corresponding to user input (e.g., touch, press, touch gesture, voice, or movement) from a control device such as a remote controller or mobile phone. Control signals can be extracted from the received light signals under the control of processor 1310.
[0299] Input / output interface 1440, under the control of processor 1310, can receive video (e.g., moving image signals or still image signals), audio (e.g., voice signals or music signals), and additional information (e.g., content description, content title, and content storage location) from devices external to image display device 1400. Input / output interface 1440 may include one of a High Definition Multimedia Interface (HDMI) port 1441, a component jack 1442, a PC port 1443, and a USB port 1444. Input / output interface 1440 may also include a combination of HDMI port 1441, component jack 1442, PC port 1443, and USB port 1444.
[0300] The video processor 1450 can process image data to be displayed by the display 1330 and perform various image processing operations on the image data, such as decoding, rendering, scaling, noise filtering, frame rate conversion, and resolution conversion.
[0301] In this embodiment, the video processor 1450 is executable. Figure 13 The image quality processor 1340 of the image display device 1300 has the function of this processor. That is, the video processor 1450 can improve the quality of the video and / or frames based on the score for each frame or the final quality score of the entire video obtained by the processor 1310.
[0302] For example, the video processor 1450 can select a quality processing model based on the score, thereby improving the quality of the frame / video.
[0303] For example, the video processor 1450 can determine the number of times to apply the quality processing model based on the score, and improve the quality of the frame / video by repeatedly applying the quality processing model to the frame according to the determined number of times.
[0304] For example, the video processor 1450 can design a filter based on the rating and apply the filter to the frame / video to improve the quality of the frame / video.
[0305] For example, the video processor 1450 can correct hyperparameter values based on the score and improve frame quality by using a neural network with corrected hyperparameter values.
[0306] The display 1330 can output content received from a broadcast station or from an external server or external storage medium on its screen. The content is a media signal and therefore may include video signals, images, text signals, etc. The display 1330 can display video signals or images received via HDMI port 1441 on its screen.
[0307] In one embodiment, when the video processor 1450 improves the quality of a video or frame, the display 1330 can output the improved quality video or frame.
[0308] When the display 1330 is implemented as a touch screen, the display 1330 can be used as both an input device and an output device. According to an embodiment of the image display device 1400, the image display device 1400 may include two or more displays 1330.
[0309] The audio processor 1460 processes audio data. The audio processor 1460 can perform various processing operations on the audio data, such as decoding, amplification, or noise filtering.
[0310] The audio interface 1470 can output, under the control of the processor 1310, audio including content received via the tuner 1410, audio input via the communicator 1420 or the input / output interface 1440, and audio stored in the memory 1320. The audio interface 1470 may include at least one of a speaker 1471, a headphone output port 1472, or a Sony / Philips digital interface (S / PDIF) output port 1473.
[0311] User interface 1480 can receive user input for controlling image display device 1400. User interface 1480 may include, but is not limited to, various types of user input devices, including touch panels that sense user touch, buttons that receive user push operations, wheels that receive user rotation operations, keyboards, dome switches, microphones for voice recognition, and motion detection sensors. When image display device 1400 is operated by a remote controller, user interface 1480 can receive control signals from the remote controller.
[0312] Figure 15 This is a flowchart illustrating a frame recognition method according to an embodiment.
[0313] Reference Figure 15 The video quality assessment device 130 can receive frames (operation 1510) and divide the frames into multiple sub-regions (operation 1520). The video quality assessment device 130 can convert each sub-region within the multiple sub-regions to the frequency domain (operation 1530).
[0314] The video quality assessment device 130 can obtain the power spectrum of the signal in the frequency domain for each sub-region and obtain the spectral envelope from the power spectrum. The video quality assessment device 130 can use the tilt value of the spectral envelope or the bin index of the point where the tilt of the spectral envelope changes rapidly to estimate the blur level of each sub-region (operation 1540).
[0315] The video quality assessment device 130 can identify whether a frame is a fully blurred frame or a partially blurred frame by using the estimated blur level for each sub-region (operation 1550). The video quality assessment device 130 can also identify whether a blurred frame is a fully blurred frame or a partially blurred frame based on whether the number of sub-regions whose blur level exceeds a certain reference value is equal to or greater than a certain number.
[0316] Figure 16 This is a flowchart illustrating a method for obtaining a quality score according to an embodiment.
[0317] Reference Figure 16 The video quality assessment device 130 can identify whether the input frame is a completely blurred frame or a partially blurred frame (operation 1610), and when the input frame is a completely blurred frame, it obtains an analysis-based quality score for the completely blurred frame (operation 1620).
[0318] In one embodiment, the quality score based on analysis could be a score obtained based on the frequency characteristics of the frame.
[0319] The video quality assessment device 130 can obtain spectral envelope distribution features from the spectral envelope of each of the multiple sub-regions included in a fully blurred frame. The spectral envelope distribution features may include at least one statistical property of the spectral envelope, such as mode, median, mean, range, bias, variance, global minimum, or global maximum.
[0320] The video quality assessment device 130 can obtain an analysis-based quality score for a completely blurred frame by analyzing the spectral envelope distribution characteristics.
[0321] When the input frame is a partially blurred frame, the video quality assessment device 130 can use at least one neural network to obtain a model-based quality score for the partially blurred frame (operation 1630).
[0322] In an embodiment, the model-based quality score may indicate the quality score obtained using at least one neural network for a partially blurred frame.
[0323] The video quality assessment device 130 can process the video to improve its quality by using at least one of an analysis-based quality score obtained for fully blurred frames or a model-based quality score obtained for partially blurred frames (operation 1640). The video quality assessment device 130 can accumulate scores for frames over a specific time period and obtain a final score for the video from these scores. The video quality assessment device 130 can use the final score to perform image quality processing, such as compensating for video distortion.
[0324] Figure 17 This is a flowchart illustrating a method for obtaining a model-based quality score for partially blurred frames according to an embodiment.
[0325] Reference Figure 17 When the input frame is a partially blurred frame (operation 1710), the video quality assessment device 130 can obtain the spectral envelope for each sub-region of the partially blurred frame and obtain the spectral envelope distribution characteristics from the spectral envelope (operation 1720).
[0326] Furthermore, the video quality assessment device 130 can acquire importance information for each sub-region of a partially blurred frame (operation 1730). The video quality assessment device 130 can use at least one neural network to obtain importance information related to factors that may affect the quality score from the partially blurred frame. The importance information may include at least one of the following: information about whether the objects included in the frame are foreground or background, information about the type of the frame, semantic information, or positional information.
[0327] The video quality assessment device 130 can use importance information for each sub-region and spectral envelope distribution characteristics for each sub-region to obtain weights for each sub-region, and obtain a weight matrix for the entire partially blurred frame from these weights (operation 1740).
[0328] The video quality assessment device 130 can obtain a model-based quality score for partially blurred frames from a weight matrix and partially blurred frames (operation 1750). The video quality assessment device 130 can obtain feature vectors from the partially blurred frames using at least one neural network, and obtain features reflecting the weights by using the feature vectors with the weight matrix. The video quality assessment device 130 can obtain a model-based quality score for the partially blurred frames from the features reflecting the weights.
[0329] Figure 18 This is a flowchart illustrating a method for obtaining a final quality score according to an embodiment. (Refer to...) Figure 18 The video quality assessment device 130 can obtain analysis-based scores and model-based scores (operation 1810), and accumulate analysis-based scores and model-based scores for a specific time period or according to a specific number of frames.
[0330] The video quality assessment device 130 can obtain time-series data for a specific time period or a specific number of frames by accumulating the scores for each frame and the timestamp of each frame (operation 1820).
[0331] The video quality assessment device 130 can obtain a final quality score for the video by smoothing the time-series data (operation 1830).
[0332] The video quality assessment device 130 can take into account the impact of time on video quality to obtain a final score for the entire video. The video quality assessment device 130 can use at least one neural network to obtain a final quality score for the entire video from time-series data.
[0333] Figure 19 This is a flowchart illustrating a method for evaluating video quality according to another embodiment.
[0334] Reference Figure 19 The video quality assessment device 130 can receive blurred frames (operation 1910) and can identify whether the input frame is a completely blurred frame or a partially blurred frame (operation 1920).
[0335] The video quality assessment device 130 can obtain different features depending on the type of input frame.
[0336] When the input frame is a completely blurred frame, the video quality assessment device 130 can obtain the spectral envelope distribution characteristics from the spectral envelope obtained for each sub-region of the completely blurred frame (operation 1930).
[0337] When the input frame is a partially blurred frame, the video quality assessment device 130 can obtain model-based features from the partially blurred frame (operation 1940). The model-based features are related to the quality of the partially blurred frame and may include at least one of the following: blur-related features, motion-related features, content-related features, perceptual features, spatial features, depth features extracted from multiple hidden layers for each layer, or features extracted statistically from lower to higher levels.
[0338] The video quality assessment device 130 can obtain importance information affecting the quality assessment for each sub-region from the input frame (operation 1950). The video quality assessment device 130 can obtain the spectral envelope distribution characteristics for each sub-region of the input frame, and generate a weight matrix from the importance information for each sub-region and the spectral envelope distribution characteristics for each sub-region (operation 1960).
[0339] The video quality assessment device 130 obtains time-series data by receiving and accumulating input frames for a specific time period, spectral envelope distribution features of fully blurred frames, and model-based features and weight matrices of partially blurred frames. The video quality assessment device 130 can obtain a final quality score for the video from the time-series data (Operation 1970).
[0340] Video quality assessment methods and apparatus according to some embodiments may be implemented as storage media including computer-executable instruction code (such as computer-executable program modules). Computer-readable media can be any available medium accessible by a computer and includes all volatile / non-volatile and removable / non-removable media. Computer-readable media can be non-transitory. Furthermore, computer-readable media can include all computer storage and communication media. Computer storage media includes all volatile / non-volatile and removable / non-removable media implemented using specific methods or techniques for storing information (such as computer-readable instruction code, data structures, program modules, or other data). Communication media typically include computer-readable instruction code, data structures, program modules, or other data, or other transmission mechanisms of modulated data signals, and includes any information transmission medium.
[0341] The video quality assessment method and apparatus according to the above embodiments can be implemented as a computer program product, wherein the computer program product includes a recording medium storing a program for performing the video quality assessment method, the video quality assessment method comprising: receiving frames included in a video; identifying whether an input frame is a fully blurred frame or a partially blurred frame based on the blur level of the input frame; obtaining an analysis-based quality score for a fully blurred frame corresponding to the input frame; obtaining a model-based quality score for a partially blurred frame corresponding to the input frame; and processing the video using at least one of the analysis-based quality score or the model-based quality score to improve its quality.
[0342] The video quality assessment method and apparatus according to the embodiments can identify whether a frame is a fully blurred frame or a partially blurred frame based on the blur level of the frames included in the video, and assess the quality in different ways in each case.
[0343] The video quality assessment method and apparatus according to the embodiments can obtain different features from fully blurred frames and partially blurred frames, and use the different features to assess the quality of each frame.
[0344] The video quality assessment method and apparatus according to the embodiments can obtain importance information and spectral envelope distribution characteristics for each sub-region of a frame included in a video, and use a weight matrix generated from the importance information and spectral envelope distribution characteristics to quantify the video quality, thereby taking into account the characteristics of each sub-region of the frame to assess the quality of the frame.
[0345] The video quality assessment method and apparatus according to the embodiments can use a neural network that has learned human subjective evaluation to assess the quality of frames.
[0346] Although embodiments have been described in detail, those skilled in the art will understand that various changes in form and detail may be made in this disclosure without departing from the spirit and scope of the disclosure as defined by the claims.
Claims
1. A video quality assessment method, comprising: Receive video frames; Based on the blur level of the frame, identify whether the frame is a completely blurred frame or a partially blurred frame; In response to the frame being the fully blurred frame, an analysis-based quality score is obtained for the fully blurred frame using an analysis-based method; In response to the frame being the partially blurred frame, a model-based quality score is obtained for the partially blurred frame using a model-based approach; and The video is processed based on at least one of the analysis-based quality score or the model-based quality score to obtain a processed video. The analysis-based method differs from the model-based method.
2. The video quality assessment method as described in claim 1, further comprising: Obtain the spectral envelope for each of the multiple sub-regions included in the frame; and The ambiguity level for each of the plurality of sub-regions is estimated based on the spectral envelope. The step of identifying whether the frame is a fully blurred frame or a partially blurred frame includes: identifying the number of multiple sub-regions whose blur level exceeds a threshold based on the spectral envelope estimation.
3. The video quality assessment method as described in claim 2, wherein, The steps for obtaining the spectral envelope for each of the plurality of sub-regions include: Obtain the signal in the frequency domain for the corresponding sub-region; Obtain the power spectrum of the signal in the frequency domain; and The spectral envelope is obtained based on the power spectrum.
4. The video quality assessment method as described in claim 2, wherein, The steps to obtain the analysis-based quality score include: The spectral envelope distribution features are obtained based on the spectral envelope of each of the multiple sub-regions included in the fully blurred frame; and An analysis-based quality score for the fully blurred frame is obtained by analyzing the spectral envelope distribution characteristics for each of the plurality of sub-regions.
5. The video quality assessment method as described in claim 2, further comprising: The spectral envelope distribution features are obtained based on the spectral envelope of each of the multiple sub-regions included in the partially blurred frame; Importance information for each sub-region among the plurality of sub-regions included in the partially blurred frame is obtained by using a first neural network; and Based on the spectral envelope distribution features and importance information obtained for each of the plurality of sub-regions included in the partially blurred frame, a weight matrix indicating the weight values for each of the plurality of sub-regions is generated. The step of obtaining the model-based quality score includes: obtaining a model-based quality score for the partial blurred frames based on the weight matrix and the partial blurred frames.
6. The video quality assessment method as described in claim 5, wherein, The step of obtaining the importance information includes: for each of the plurality of sub-regions, obtaining one or more pieces of importance information related to factors that can influence the quality score by using a first neural network, and The importance information includes at least one of the following: whether the object is a foreground or background object, semantic information, location information, or content information.
7. The video quality assessment method as described in claim 5, wherein, The plurality of sub-regions includes a first sub-region and a second sub-region adjacent to the first sub-region, and The video quality assessment method further includes: The first spectral envelope distribution characteristics of the first sub-region are corrected by using the second spectral envelope distribution characteristics of the second sub-region; The first importance information of the first sub-region is corrected by using the second importance information of the second sub-region; and The corrected weight values for the first sub-region are obtained by using the corrected first spectral envelope distribution features and the corrected first importance information.
8. The video quality assessment method as described in claim 5, wherein, The step of obtaining the model-based quality score is performed using a second neural network trained on the correlation between the feature vector and the average opinion score.
9. The video quality assessment method as described in claim 8, wherein, The steps to obtain the model-based quality score include: Features are extracted from the partially blurred frames using a second neural network; and The quality score of the partially blurred frame is obtained based on the features and the weight matrix, and The features include at least one of fuzzy correlation features, motion correlation features, content correlation features, depth features, statistical features, perceptual features, spatial features, or modified domain features.
10. The video quality assessment method as described in claim 1, further comprising: Accumulate the analysis-based quality scores and the model-based quality scores for multiple frames over a period of time to obtain time series data; and The time series data is smoothed to obtain a final quality score.
11. The video quality assessment method as described in claim 10, wherein, The step of smoothing the time series data to obtain the final quality score is performed using a third neural network model, and The third neural network model includes Long Short-Term Memory (LSTM).
12. The video quality assessment method as described in claim 10, wherein, The steps for processing the video include: processing the frames based on the final quality score. The video processing steps are performed by at least one of the following operations: The video is processed according to the quality processing model selected based on the final quality score. The video is processed by identifying the number of times the quality processing model is applied based on the final quality score, and by repeatedly applying the quality processing model to the frames according to the number of times. Based on the final quality score identification filter, the video is processed by applying the filter to the frame, or The video is processed using a neural network that corrects for hyperparameter values based on the final quality score.
13. A video quality assessment method, comprising: Receive video frames; The frame is identified as either fully blurred or partially blurred based on its blur level. In response to the frame being the fully blurred frame, the spectral envelope distribution features of the fully blurred frame are obtained, and a final quality score for the frame is obtained based on the spectral envelope distribution features. In response to the frame being the partially blurred frame, model-based features of the partially blurred frame are obtained, and a final quality score for the frame is obtained based on the model-based features. and The video is processed based on the final quality score to obtain a processed video.
14. The video quality assessment method as described in claim 13, further comprising: Obtain importance information for each of the multiple sub-regions included in the frame; Obtain the spectral envelope distribution features for each of the plurality of sub-regions included in the frame; and A weight matrix for the frame is generated by obtaining the weights for each of the multiple sub-regions based on the importance information and spectral envelope distribution characteristics of each sub-region. The step of obtaining the final quality score for the frame includes: obtaining the final quality score for the frame by using the spectral envelope distribution features of each sub-region of the plurality of sub-regions of the fully blurred frame, the model-based features of the partially blurred frame, and the weight matrix.
15. A video quality assessment device, comprising: Memory, which stores one or more instructions; as well as The processor is configured to execute one or more instructions stored in memory to perform the following operations: The frame is identified as either fully blurred or partially blurred based on the blur level of the frames included in the video. In response to the frame being the fully blurred frame, an analysis-based quality score is obtained for the fully blurred frame using an analysis-based method; In response to the frame being the partially blurred frame, a model-based quality score is obtained for the partially blurred frame using a model-based approach; and The video is processed based on at least one of the analysis-based quality score or the model-based quality score to obtain a processed video. The analysis-based method differs from the model-based method.