How to predict the actual resolution of video content
A non-reference video-based AI model using deep learning predicts actual video resolution by measuring quality score differences, addressing the limitations of metadata-based methods and enhancing video service efficiency and quality.
Patent Information
- Application Number
- JP2024112467
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2024-02-27
- Filing Date
- 2024-07-12
- Publication Date
- 2025-10-01
- Estimated Expiration
- 2044-07-12
AI Technical Summary
Existing metadata-based methods for determining video content resolution are unreliable in distinguishing genuine from counterfeit ultra-high definition (UHD) content, as metadata can be manipulated, and manual quality assessment is time-consuming and impractical for real-time use.
A non-reference video-based AI model using deep learning is trained on a dataset of video content with various resolutions to predict actual resolution by measuring quality score differences across downscaled clips, filtering out defective content, and calculating quality improvement factors.
Accurately identifies fake UHD videos and optimizes service operations by prioritizing high-quality content, reducing costs and improving customer experience for OTT providers.
Smart Images

Figure 0007747831000001 
Figure 0007747831000002 
Figure 0007747831000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method for predicting the actual resolution of video content, and more particularly, to a method for predicting the actual resolution of video content using a non-reference video-based AI model constructed by learning a learning dataset consisting of video content with various resolutions using a deep learning method in order to overcome the limitations of existing metadata-based resolution discrimination methods. [Background technology]
[0002] Perceived video quality can be divided into aesthetic and technical scores. Aesthetic scores are information about the content of a video, and generally, a video containing content like a bear doll is perceived as having better quality than a video containing content like scissors. Technical scores are a technical quality score, such as the level of shaking during filming or pixel distortion.
[0003] Recently, there has been an increase in demand for ultra-high definition (UHD) video content, which generally has a resolution of 3840x2160 (2160p, 4K) or 8680x4320 (4320p, 8K) (the following explanation will use 4K as an example). However, taking advantage of the increasing demand for 4K video content, counterfeit 4K video content, in which only the resolution has been changed to 4K, is being widely distributed.
[0004] The most common way to find out the resolution of existing video content is to read the file's metadata. Most video content contains metadata that includes technical information about the video. In other words, users can read the metadata using video processing software or tools to check the resolution of the video content.
[0005] The advantage of this method is that for most video content, the resolution included in the metadata is the same as the actual resolution at the time of shooting, so the resolution of the video content can be obtained easily and quickly without complex calculations.
[0006] The drawback of this method is that the resolution confirmed through the metadata of the video content may differ from the actual resolution at the time of shooting. For example, if a high-resolution video is downscaled to a lower resolution and then upscaled again, the metadata may show a high resolution, but the actual resolution may be lower.
[0007] As such, the problem with the conventional metadata-based quality assessment method is that it is not possible to distinguish between genuine 4K video and fake 4K video, which is obtained by forcibly upscaling 4K video content that has been intentionally downscaled to 1440p or lower resolution back to 4K resolution.
[0008] To solve this problem, fake 4K videos can be identified by comparing the Mean Opinion Score (MOS) given by people after viewing any video content at each resolution, and checking whether there is a difference between the MOS score given after viewing 4K video content and the MOS score given after viewing video content with a lower resolution (assuming that the only difference between the video content is resolution).
[0009] However, the above method requires a great deal of time and money when determining whether video content, which may be several hours long, is authentic or not, and has the drawback of being difficult to perform in real time or automatically. [Prior art documents] [Patent documents]
[0010] [Patent Document 1] Republic of Korea Patent Publication No. 10-2022-0081246 [Patent Document 2] Republic of Korea Patent Publication No. 10-2005-0107488 [Patent Document 3] Republic of Korea Patent Publication No. 10-2023-0102658 Summary of the Invention [Problem to be solved by the invention]
[0011] The main objective of the present invention is to overcome the limitations of existing metadata-based resolution discrimination methods by providing a method for predicting the actual resolution of video content using a no-reference video-based AI model constructed by deep learning using a training dataset consisting of video content with various resolutions. [Means for solving the problem]
[0012] To achieve the above-mentioned object, the method for predicting the actual resolution of video content of the present invention includes: (a) dividing the video content to be predicted into video clips of a predetermined length; (b) downscaling the resolution of each video clip to a predetermined resolution; (c) measuring the quality of the downscaled video clips in step (b) in order from low to high resolution using a non-reference video-based AI model, and then calculating the quality score difference between two video clips of adjacent resolutions in order from low to high resolution; (d) predicting the resolution of each video clip based on the quality score difference; and (e) combining the resolutions of each video clip to predict the actual resolution of the corresponding video content.
[0013] In the above configuration, the resolutions that can be downscaled are 360p, 480p, 720p, 1080p, 1440p and 2160p.
[0014] The maximum resolution that can be predicted when predicting the resolution of video content is min(U, the resolution of the original video content).
[0015] Step (d) is a step of generating a current quality score difference D, which is the difference between the quality scores of two adjacent resolution video clips currently being processed. P is less than the reference value E (d1), and the current quality score difference D P If the reference value E is less than the reference value E, the quality score difference between the adjacent resolution video clips processed immediately before is the previous quality score difference D B is less than the reference value E (d2), and the current quality score difference D P and the previous quality score difference D B If both are less than the standard value E, the previous quality score difference D P and (d3) predicting the low resolution used in calculating (a) as the resolution of the corresponding video clip.
[0016] (d) Stage: Current quality score difference D P Or the previous quality score difference D B If at least one of the above is equal to or greater than the reference value E, the quality score difference is calculated up to the final resolution, and the current quality score difference D based on the final resolution is calculated. P and the previous quality score difference D B The maximum quality difference D max and predicting the generated resolution as the resolution of the corresponding video clip (d4).
[0017] The AI model used in step (b) is constructed by training multiple (M) video training datasets with various resolutions, including 360p, 480p, 720p, 1080p, 1440p, and 2160p, on a deep learning platform.
[0018] The training dataset is the YouTube-UGC (User-Generated Contents) dataset.
[0019] The number of training datasets (M) is 1000 or more.
[0020] During the training dataset preparation process, images with a high rate of quality defects are filtered out by each defect item up to a maximum of floor ((N / 100)*M; N is a configurable natural number).
[0021] The quality defects are removed by filtering video content with a high defect ratio of (1) excessively dark, (2) excessively bright, and (3) excessively blurred for each of the defect items (1), (2), and (3).
[0022] Video clips having quality defects after step (a) and before step (b) are removed through filtering.
[0023] For each video clip, a quality improvement factor A between the lowest and highest resolutions is calculated, and then the average is determined as the final quality improvement factor for the video content.
[0024] In step (e), the resolution of the video clip having the highest resolution among the resolutions of the video clips is predicted as the resolution of the corresponding video content. [Effects of the Invention]
[0025] The method for predicting the actual resolution of video content of the present invention can accurately predict the actual resolution of video content using a non-reference image-based AI model, which can help service providers that provide a large number of videos, such as OTT (Over The Top) providers, to sort out fake 4K videos that have quality unrelated to the resolution included in the metadata. As a result, video service providers can optimize their service operation costs and customers can enjoy reliable 4K video.
[0026] In addition, even for true 4K video content, by providing an auxiliary indicator showing how much the quality has improved compared to when it was provided at a lower resolution, video service providers can select true 4K video that should be stored as a top priority in environments where service capacity is limited, such as when storage capacity is insufficient, thereby reducing service operation costs. [Brief explanation of the drawings]
[0027] [Figure 1A] 1 is a flowchart illustrating a method for predicting the actual resolution of video content according to the present invention. [Figure 1B] 1 is a flowchart illustrating a method for predicting the actual resolution of video content according to the present invention. [Figure 2A] 1 is a diagram illustrating an example of a method for predicting the resolution of each video clip in a method for predicting the actual resolution of video content according to an embodiment of the present invention; [Figure 2B] 1 is a diagram illustrating an example of a method for predicting the resolution of each video clip in a method for predicting the actual resolution of video content according to an embodiment of the present invention; [Figure 2C] 1 is a diagram illustrating an example of a method for predicting the resolution of each video clip in a method for predicting the actual resolution of video content according to an embodiment of the present invention; DETAILED DESCRIPTION OF THE INVENTION
[0028] The main improvement goal of the method for predicting the actual resolution of video content of the present invention is to overcome the limitations of existing resolution discrimination methods based only on the metadata of video content and to propose a method for predicting the actual resolution that affects the quality of video content. The specific improvement goals are as follows:
[0029] First, obtaining resolution based on the quality perceived by humans: We propose a method that can predict resolution based on the actual quality of the video, that is, a method that predicts the resolution of video content based on the quality perceived by humans through the video content, rather than based on metadata information.
[0030] Second, it can be used without metadata: It can be applied to any video as long as only the RGB of the video content can be obtained through a non-reference video-based quality measurement method, regardless of the codec type or the presence or absence of metadata. As a result, we propose a method that can be applied to virtually all video content as long as it is playable, regardless of the presence or absence of metadata.
[0031] Third, separate management of quality improvement factor A by resolution: For each video content, the quality multiplier of the video quality at the lowest resolution (e.g., 144p) and the video quality at the highest resolution (e.g., 4K) is identified for each video based on the degree of improvement in the quality experienced by humans, allowing for the selection and management of videos with a greater degree of quality improvement as the resolution increases.
[0032] For example, a color movie can have a greater improvement in human perceived quality when viewed in 4K video compared to a black and white movie. Video content that has a significant improvement in quality when viewed in 4K video is therefore video content that has a high priority for storage in an environment where physical storage space is insufficient for OTT service providers to provide services, so we provide a function that allows you to select, store, and manage this content.
[0033] Fourth, pre-filtering for quality defects during the learning and prediction process: Video content with a high rate of outliers, such as excessively dark, bright, or blurry screens, cannot correctly learn quality changes when training an AI model based on no-reference video. Therefore, video content with a high rate of outliers is filtered out to a maximum of N% from the training dataset, so that the AI model that measures video quality is less sensitive to outlier data and is trained to measure quality based solely on resolution, and video sections with such outlier characteristics are also excluded from the resolution prediction process.
[0034] This prevents situations such as when some sections of a movie are shot in darkness, and even if the video in question is genuine UHD video content, some sections have abnormal values such as darkness, i.e. low-resolution quality characteristics, and the resolution is predicted to be low, such as 144p.
[0035] Fifth, the aesthetic score and other technical scores are preserved as much as possible during the resolution prediction process: In this invention, the resolution of the video content is forcibly downscaled to a resolution lower than the physical resolution of the video content, for example, the metadata-based resolution, and the quality of the downscaled video content is measured multiple times. During this downscaling process, only the resolution is changed, and other information such as the codec information and color space information is not changed at the same time.
[0036] In other words, the aesthetic score of the video content is kept as constant as possible during the downscaling process, and the quality of the video content is measured while changing only the resolution among the technical elements, thereby eliminating the impact of codecs and other factors on quality and allowing for more precise measurement of only the extent to which resolution affects quality.
[0037] Sixth, the introduction of the concept of durable resolution steps allows for more accurate resolution prediction: Resolution is predicted based on the range of downscaled resolutions from the lowest to the highest, but among the eight commonly used resolutions (144p, 240p, 360p, 480p, 720p, 1080p, 1440p, 2160p), there may be sections where the degree of quality improvement is relatively small compared to other sections where the resolution is increased.
[0038] For example, there may be video content in which the degree of perceived quality improvement is relatively small or almost nonexistent in the 720p->1080p resolution increase section, but is particularly pronounced in the 1080p->1440p resolution increase section. Therefore, by introducing the number of durable resolution steps (W) in the process of determining whether there is a quality improvement while increasing the downscaled resolution, it is possible to prevent the quality of the video content from being mistakenly predicted to be 720p resolution because, in the example above, there is a quality improvement in the 1080p->1440p section, but the quality improvement is very small in the 720p->1080p section (e.g., a difference of less than 0.02 in MOS value).
[0039] Hereinafter, a preferred embodiment of the method for predicting the actual resolution of video content according to the present invention will be described in detail with reference to the accompanying drawings.
[0040] The main technical idea of the method for predicting the actual resolution of video content of the present invention is to measure only the quality change due to resolution as a technical element while leaving the aesthetic elements of the video content intact by downscaling the original video to multiple relatively low resolutions and then measuring and observing the quality change while increasing the resolution.
[0041] In this process, the resolution that increases but does not increase quality is predicted as the actual resolution of the video content, and in this process, sections with abnormal value characteristics such as the screen being excessively bright, excessively dark, or excessively blurred are excluded from the resolution prediction process, thereby improving the accuracy of the resolution prediction.
[0042] 1A and 1B are flowcharts illustrating the method for predicting the actual resolution of video content of the present invention, which mainly includes a process of constructing an AI model for non-reference-based video quality measurement (Process A: steps S110 to S130) and a process of predicting the actual resolution of any video content through the AI model thus constructed (Process B: steps S210 to S275).
[0043] First, in step S110 of the AI model construction process (Process A), a training dataset having various resolutions ranging from 360p to 2160p, for example, the YouTube-UGC (User-Generated Contents) dataset (https: / / media.withyoutube.com), is prepared. In this case, the number of video contents (M) included in the training dataset may be 1381. In addition, the maximum resolution (U) used for training in this process is separately defined, and in the above example, U may be 2160p.
[0044] Next, in step S120, images with quality defects in the training dataset, for example, image content with a high defect rate such as excessively dark, excessively bright, or excessively blurred, are filtered out. For example, up to floor((N / 100)*M) images are filtered out for each defect category. For example, assuming that all image content with a high defect rate such as (1) excessively dark, (2) excessively bright, and (3) excessively blurred is to be removed, the number of removal steps (S) is 3. In the above example, if N=1, floor((N / 100)*M) is 13, and 13 images are removed as defective image content for each of the above-mentioned defect categories (1), (2), and (3).
[0045] Meanwhile, to determine whether or not each defect item has a defect, an algorithm that can numerically measure darkness, an algorithm that can numerically measure brightness, or an algorithm that can numerically measure blur can be used appropriately. The defective video content removal process can be carried out sequentially, but ultimately S*floor((N / 100)*M) videos are removed, so in the case of the YouTube-UGC dataset, 39 (=3*13) video contents are removed, and the final number of video contents that can be learned is 1,342.
[0046] Here, S and N are arbitrarily settable variables, and M and U are variables dependent on the training dataset. U is the maximum resolution that can be predicted in the process of predicting the user's video resolution based on the perceived quality, and a training dataset having such U is appropriately selected. In this case, M is preferably 1000 or more.
[0047] Next, in step S130, a deep learning-based no-reference video quality measurement AI model is constructed that predicts video quality at various resolutions independently while being less affected by outliers through learning using a learning dataset.
[0048] Next, in step S210 of the actual resolution prediction process (process B), the video content to be subjected to resolution prediction is divided into video clips of a specific length, for example, 10 seconds. For example, assuming a video content of one hour (3600 seconds) in length, a total of 360 video clips will be obtained.
[0049] Next, in step S215, among all the video clips, video clips that are inappropriate for use in resolution prediction, for example, video clips that have defects such as being too dark, too bright, or too blurry, are filtered out.
[0050] Meanwhile, the presence or absence of defects in each video clip may be determined by the algorithm used in step S120. For example, if video content with a brightness level of less than 5.12 out of 100 is deemed abnormal in the learning process, video clips with a brightness level of less than 5.12 will also be deemed abnormal in the prediction process. Taking the above case as an example, if two video clips out of a total of 360 clips are overly bright and there are no overly dark or overly blurred video clips, only 358 video clips will be used for resolution prediction.
[0051] Next, in step S220, the resolution of each video clip is downscaled to a predetermined resolution, for example, 360p, 480p, 720p, 1080p, 1440p, and 2160p. The maximum resolution that can be predicted when predicting the resolution of the video content is min(U, the resolution of the original video content).
[0052] Next, in step S225, the quality of the downscaled video clip is measured using the AI model constructed in step S130. For example, the quality of the video clips at each resolution is measured in order from low to high resolution, and then the quality score difference D between two video clips at adjacent resolutions is calculated in order from low to high resolution.
[0053] Next, in step S230, the difference in quality scores (DP (Hereinafter, simply referred to as the "current quality score difference") is determined to be less than a predetermined reference value E.
[0054] At step S230, the current quality score difference is D P If D is less than the reference value E, the process returns to step S240 to calculate the quality score difference (D) between the video clips of adjacent resolutions processed immediately before. B In step S240, it is determined whether the previous quality score difference D (hereinafter simply referred to as "previous quality score difference") is less than the reference value E. B If the difference is less than the reference value E, the video clip has almost no difference in quality across the three resolutions, so the process proceeds to step S250. P The lower resolution used in the calculation is predicted as the resolution of the video clip.
[0055] On the other hand, at stage S230, the current quality score difference is D P If the reference value E is equal to or greater than the reference value E, the process proceeds to step S235 to determine whether the final resolution, for example, 2160p, has been reached. If the final resolution has not been reached in step S235, the process proceeds to step S245 to increase the resolution to be processed by one step, and then returns to step S225.
[0056] At step S240, the previous quality score difference D B If the resolution is equal to or greater than the reference value E, the process proceeds to step S235. If the resolution being processed in step S235 reaches the final resolution, the process proceeds to step S255, where the current quality score difference D based on the final resolution is calculated. P and the previous quality score difference D B The maximum quality difference D max The resolution at which the error occurred is predicted as the resolution of the corresponding video clip.
[0057] 2A to 2C are diagrams illustrating an example of a method for predicting the resolution of each video clip in the method for predicting the actual resolution of video content according to the present invention. Hereinafter, the reference value E is set to 0.02. Assuming that the quality scores of video clips having resolutions of 360p, 480p, 720p, 1080p, 1440p, and 2160p are 2.21, 2.61, 3.01, 3.01, 3.02, and 3.10, respectively, as illustrated in FIG. 2A. The quality score differences D between 360p->480p and 480p->720p are 0.4, which are greater than the reference value (E=0.02). Therefore, the results of the determinations in steps S230 and S240 are both "N."
[0058] Therefore, the process proceeds to step S235 to determine whether the final resolution has been reached. Since the currently processed resolution is 720p, which is not the final resolution, the process proceeds to step S245 to calculate the quality score difference D between 720p and 1080p, which is the next step up in resolution.
[0059] As a result of repeating this process, in the example of Figure 2a, the current quality score difference D P is 0.01, and the difference in quality score just before 720p->1080p is D B 0 and both are less than the reference value E, the process proceeds to step S250 and the previous quality score difference D B The resolution of the video clip is predicted to be 720p, the lower resolution used in the calculation.
[0060] Next, as illustrated in FIG. 2B, if the quality scores of video clips having resolutions of 360p, 480p, 720p, 1080p, 1440p, and 2160p are 2.21, 2.61, 3.01, 3.01, 3.10, and 3.10, respectively, the quality score difference D is calculated based on the final resolution. Since the quality score difference D is not determined as "Y" in steps S230 and S240, the quality score difference D is calculated based on the final resolution. P and the previous quality score difference D B The maximum quality score difference (D max=0.09) is predicted as the resolution of the video clip in question, 1440p.
[0061] Finally, as illustrated in FIG. 2C, if the quality scores of video clips having resolutions of 360p, 480p, 720p, 1080p, 1440p, and 2160p are 2.21, 2.61, 3.01, 3.01, 3.10, and 3.30, respectively, the maximum quality score difference (D) is calculated in step 255 based on the result of calculating the quality score difference D up to the final resolution. max =0.2) is predicted as the resolution of the video clip in question, which is 2160p.
[0062] Returning to FIG. 1B again, in step S260, a quality improvement factor A between the lowest and highest resolutions is calculated for each video clip. In the example of FIG. 2A, the quality score for the lowest resolution, 360p, is 2.21, and the quality score for the highest resolution, 2160p, is 3.10, so the quality improvement factor A is 1.40 (=3.10 / 2.21).
[0063] Next, in step S265, it is determined whether all video clips have been processed. If unprocessed video clips remain, the process returns to step S220. If not, the process proceeds to step S270, where the resolution of the video clip with the highest resolution is finally predicted as the resolution of the corresponding video content. In the above example, if some of the total 358 video clips are predicted to be 360p or 720p and some are predicted to be 1440p, the final resolution of the corresponding video content is predicted to be 1440p. This is because, due to the influence of dynamic encoding, there may be no quality difference between low resolution and high resolution in some sections.
[0064] Finally, in step S275, the average of the quality improvement factors A of each video clip is determined as the final quality improvement factor for the corresponding video content, thereby improving service operation efficiency. For example, assume that the actual resolutions of two video contents predicted by the actual resolution prediction method described in FIG. 1 are all predicted to be 1440p, and the quality improvement factors A of each video content are 1.40 and 1.70. In this state, if only one video needs to be downscaled to 1080p due to a lack of storage capacity, the video content that should be downscaled first may be the video content with a relatively low quality improvement factor A of 1.40, thereby improving service operation efficiency.
[0065] The above description is provided merely to aid in understanding the methods described herein, and the present invention is not limited thereto.
[0066] The present invention can be modified in various ways and can have various embodiments. The scores and reference values mentioned in the above embodiments are merely illustrative and may be changed as appropriate. For example, while the above embodiment describes a maximum resolution of 2160p, this is not limited to this and may be extended to 8K (4320p) or higher. In the above embodiment, the number of durable resolution steps (W) is set to two steps (steps S230 and S240), including the current and previous steps. However, this may be changed to three steps, including the previous step. Furthermore, rather than determining defective video content or defective video clips by defect category, such as being too dark, too bright, or too blurry, defective video content or defective video clips may be determined by a score that collectively quantifies the defect categories.
[0067] Therefore, the scope of the present invention should be determined by the following claims.
Claims
1. (a) dividing a video content to be subjected to resolution prediction into video clips of a predetermined length; (b) downscaling the resolution of each video clip to a predetermined resolution; (c) inputting the downscaled video clips in step (b) into a non-reference video-based AI model that is constructed by learning through deep learning using a learning dataset including a plurality of videos with various resolutions and predicting quality scores of videos with various resolutions, calculating quality scores of the downscaled video clips in step (b) in order from low resolution to high resolution, and then calculating quality score differences between video clips of two adjacent resolutions in order from low resolution to high resolution; (d) predicting the resolution of each video clip based on the quality score difference; (e) predicting the actual resolution of the corresponding video content by combining the resolutions of each video clip; A method for predicting actual resolution of video content, comprising:
2. 2. The method for predicting actual resolution of video content according to claim 1, wherein resolutions to be downscaled are 360p, 480p, 720p, 1080p, 1440p, and 2160p.
3. 3. The method for predicting the actual resolution of video content according to claim 2, wherein the maximum resolution that can be predicted when predicting the resolution of video content is min(U, resolution of original video content).
4. Step (d) includes determining whether a current quality score difference (DP), which is a difference in quality scores between two adjacent resolution video clips currently being processed, is less than a reference value E (d1); If the current quality score difference (DP) is less than the reference value E, it is determined whether the previous quality score difference (DB), which is the quality score difference between the video clips of adjacent resolutions processed immediately before, is less than the reference value E (d2); The method for predicting the actual resolution of video content as described in claim 3, characterized in that it includes a step (d3) of predicting the low resolution used when calculating the previous quality score difference (DP) as the resolution of the corresponding video clip if both the current quality score difference (DP) and the previous quality score difference (DB) are less than the reference value E.
5. The method for predicting the actual resolution of video content as described in claim 4, characterized in that step (d) includes step (d4) of calculating the quality score difference up to the final resolution if at least one of the current quality score difference (DP) or the previous quality score difference (DB) is equal to or greater than the reference value E, and predicting the resolution at which the maximum quality score difference (Dmax) occurs among the current quality score difference (DP) and the previous quality score difference (DB) based on the final resolution as the resolution of the corresponding video clip.
6. The method for predicting the actual resolution of video content described in claim 1, characterized in that the learning dataset is a YouTube-UGC (User-Generated Contents) dataset.
7. A method for predicting the actual resolution of video content as described in Claim 6, characterized in that the number (M) of videos included in the training dataset is 1,000 or more.
8. A method for predicting the actual resolution of video content as described in Claim 7, characterized in that during the preparation process of the learning dataset, videos with a high proportion of quality defects are filtered and removed by each defect item up to a maximum floor ((N / 100) * M; N is a configurable natural number).
9. The method for predicting the actual resolution of video content described in claim 8, characterized in that video content with a high defect rate of (1) excessively dark, (2) excessively bright, and (3) excessively blurred is filtered and removed for each of the defect items (1), (2), and (3).
10. 10. The method of claim 9, wherein video clips having quality defects after step (a) and before step (b) are removed through filtering.
11. The method of claim 10, wherein a quality improvement factor A between the lowest resolution and the highest resolution is calculated for each video clip, and then the average is determined as the final quality improvement factor for the corresponding video content.
12. 12. The method of claim 11, wherein in step (e), a resolution of a video clip having a highest resolution among the resolutions of the video clips is predicted as a resolution of the corresponding video content.
Citation Information
Patent Citations
Deep Learning Medical Systems and Methods for Medical Procedures
JP2020500378A
Inference method, inference device, and program
JP2024021654A
Reducing Bandwidth Consumption with Generative Adversarial Networks
JP2024512476A
Scalable encoding and decoding of interlaced digitalvideo data
KR1020050107488A
Method and apparatus for restoring low resolution of video to high resolution data based on self-supervised learning
KR1020220081246A